SmolVLA (LIBERO) compiled for the RDK BPU β€” S100P and S600

LeRobot's SmolVLA running entirely on a D-Robotics RDK board's BPU, through the BLLM runtime. Two camera frames, a proprioceptive state vector and an English instruction go in; a chunk of 50 end-effector deltas comes out.

Two builds, one per board:

s100p/ (RDK S100P) s600/ (RDK S600)
per 50-action chunk 562 ms 159 ms
↳ vision tower, 2 cameras 261 ms 60 ms (4 BPU cores)
↳ SmolLM2 prefix, 177 tokens 163 ms 43 ms
↳ action expert, 4 denoising steps 121 ms 43 ms
graphs resident 1.43 GB 0.43 GB
libero_spatial 68 / 100 71 / 100
libero_object 76 / 100 75 / 100
libero_goal 79 / 100 81 / 100
libero_10 51 / 100 54 / 100
all four suites / 400 274 281

The same checkpoint through the same harness on an RTX 4090 scores 73 / 72 / 79 / 48 on the four suites β€” 272 / 400, against the boards' 274 and 281. Twelve 100-episode evaluations, same env, same init states, same seeds, same wire, with only the side computing the chunk changed.

The sign flips between suites (the boards are behind on libero_spatial and ahead on the other three), which is what n = 100 noise looks like; a systematic porting loss would move all four the same way. The per-task profile is shared as well β€” libero_goal task 6 is 5 on all three, libero_10 task 0 is 1 on all three.

libero_10 is the checkpoint, not the board. It sits at 48–54 where the other suites are 68–81, and the reference is the lowest of the three. Scored without a reference column, that 51 would read as "the port collapses on long-horizon tasks".

Read these numbers as this checkpoint's, not as a porting loss.

Four denoising steps, not ten

These packages ship 4 steps where the reference uses 10. The action expert runs once per step, so this is the one latency lever that costs no rebuild: 741 β†’ 562 ms on S100P, 285 β†’ 220 ms on S600.

Measured across three paired comparisons rather than one, because a single 100-episode result has an interval of about Β±9 points β€” wider than the effect:

suite / board 10 steps 4 steps
libero_spatial Β· S100P 70 68
libero_spatial Β· S600 71 71
libero_object Β· S100P 76 76
total, 300 episodes 217 215

Two of the three pairs are identical. If you want the reference schedule, rebuild the package with bllm-make-model-dir smolvla --steps 10 and a 10-row cond_table β€” note that truncating this package's 4-row table would give a different schedule, not a shorter one (10 steps visit t = 1.0, 0.9, 0.8 …; 4 steps visit 1.0, 0.75, 0.5, 0.25).

Use it

hf download ruisv/bllm-smolvla-libero-512 --include "s100p/*" --local-dir .
import bllm, numpy as np

policy = bllm.load_policy("s100p")          # or "s600"

actions = policy.act(scene_rgb,             # uint8[h, w, 3], any resolution
                     wrist_rgb,             # uint8[h, w, 3]
                     "pick up the black bowl and place it on the plate",
                     state=eef_state)       # float32[8]
# actions: float32[50, 7] β€” end-effector deltas in the dataset's own units,
# already unnormalised. Feed them to the arm one row at a time.

There is also a C++ entry point, bllm_smolvla_policy, which does the same from the shell.

Runtime requirement. arch: "smolvla" support is newer than the current released bllm package β€” build BLLM from source until a release including it is out. A conda-installed bllm 0.4.2 will report an unknown arch.

Proprioception is not optional

Unlike Ο€0.5's LIBERO configuration, SmolVLA reads the state: embed_prefix projects it into a prefix token. Omitting it is a different observation, not a missing convenience β€” measured at 2.0 in action space, a full gripper unit.

The 8 values are LIBERO's own: end-effector position (3), the end-effector quaternion converted to axis-angle (3), and the gripper joint positions (2).

Pinning the latent

SmolVLA is a flow model integrated from a Gaussian, so two calls on the same observation legitimately return different actions β€” and here by a lot: the reference disagrees with itself at a position max|Ξ”| of 0.408 across latents. Pass noise= when two runs need to be comparable:

z = np.random.default_rng(0).standard_normal((50, 32)).astype(np.float32)
a1 = policy.act(scene, wrist, prompt, state=s, noise=z)
a2 = policy.act(scene, wrist, prompt, state=s, noise=z)   # bit-identical

Comparing an implementation against a reference without pinning it measures flow matching, not the implementation.

Accuracy

Same observation, same pinned latent, against the PyTorch reference:

cosine β€–aβ€–/β€–a_refβ€– max|Ξ”| position
S100P package vs reference 0.999923 0.9971 0.0381
S600 package vs reference 0.999919 0.9970 0.0398
S600 vs S100P 0.9999995 0.9999 0.0017
policy's own spread across latents 0.9960 Β±0.009 0.408

The two boards agree with each other twenty times more closely than either agrees with the reference they were both built from, and both sit an order of magnitude inside the policy's own stochasticity.

Quantisation

int8 per-output-channel weights, int16 static per-tensor activations, on all three graphs. Each stage is calibrated on the previous stage's compiled output β€” the vision .hbm's real tokens set the prefix graph's ranges, the prefix .hbm's real KV set the expert's β€” rather than on float intermediates.

Replaying that exact scheme inside the reference model on a host costs 0.09 % of action magnitude. 16-bit activations are a decision rather than a default: at int8 the same replay flips the gripper (cosine 0.9926, position 0.291).

What is in the package

file why the runtime cannot infer it
visual.hbm, model.hbm, expert.hbm SmolVLM vision tower, SmolLM2 prefix, interleaved action expert
embed_tokens.bin the tied embedding table, 49280 Γ— 960 fp16, pre-scaled by sqrt(960) (see below)
tokenizer.json SmolVLM2's tokenizer
state_proj.f32 the state projection, [960, 32] + bias
cond_table.f32 flow-matching conditioning, precomputed per denoising step
norm_stats (in model.json) MEAN_STD for the 8 state and 7 action dims
smolvla.cameras, .prompt_len, .horizon, .steps compile-time shapes

embed_prescaled is not bookkeeping. embed_prefix multiplies the tied embedding row by sqrt(hidden) after calling embed_language_tokens, and does the same to the image tokens. A package shipping the raw rows makes the prompt 31Γ— too small β€” a policy that can barely read its own instruction β€” and nothing catches it, because cosine is scale-invariant and a golden built the same way carries the same error. The table is therefore stored pre-scaled and the runtime refuses a manifest that does not say so.

Limits

  • One configuration. 2 cameras, 48 language slots, horizon 50, state 8, action 7 β€” all compile-time shapes. A finetune with three cameras or a longer instruction budget needs its own build. All 40 LIBERO instructions fit in 22 of the 48 slots; a longer one is a hard error, never a silent truncation.
  • Mixed BPU core counts on S600. The vision tower is compiled for four cores (59.0 β†’ 29.8 ms per camera), the prefix and expert for one β€” the runtime reads each graph's compiled count independently, so one package can mix them. Output is bit-identical to an all-single-core build and libero_spatial is identical episode for episode (71/100, 355 policy calls, both): multi-core changes scheduling, not arithmetic. The prefix gains only 7 % at two cores and the expert does not compile at four, which is why they stay at one. S100P has a single BPU core, so none of this applies there.
  • Image preprocessing is a plain bilinear resize with aspect-preserving padding on the left and top (SmolVLA's convention, not openpi's centred one). Feed 512Γ—512 directly to remove the question.

Provenance and terms

These are derived weights: lerobot/smolvla_libero compiled for the BPU.

licence
lerobot/smolvla_libero (the checkpoint compiled here) Apache-2.0
lerobot/libero (source of the normalisation statistics) Apache-2.0
HuggingFaceTB/SmolVLM2-500M-Video-Instruct (tokenizer) Apache-2.0
lerobot/smolvla_base (upstream base of the finetune) not declared

The checkpoint these are derived from declares Apache-2.0, which is the operative statement; the base model declaring nothing is recorded here rather than glossed. Check upstream terms before redistributing.

Conversion method and acceptance record: bllm-model-zoo β†’ models/smolvla-libero/.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for ruisv/bllm-smolvla-libero-512