SmolVLA (LIBERO) compiled for the RDK BPU β S100P and S600
LeRobot's SmolVLA running entirely on a D-Robotics RDK board's BPU, through the BLLM runtime. Two camera frames, a proprioceptive state vector and an English instruction go in; a chunk of 50 end-effector deltas comes out.
Two builds, one per board:
s100p/ (RDK S100P) |
s600/ (RDK S600) |
|
|---|---|---|
| per 50-action chunk | 562 ms | 159 ms |
| β³ vision tower, 2 cameras | 261 ms | 60 ms (4 BPU cores) |
| β³ SmolLM2 prefix, 177 tokens | 163 ms | 43 ms |
| β³ action expert, 4 denoising steps | 121 ms | 43 ms |
| graphs resident | 1.43 GB | 0.43 GB |
| libero_spatial | 68 / 100 | 71 / 100 |
| libero_object | 76 / 100 | 75 / 100 |
| libero_goal | 79 / 100 | 81 / 100 |
| libero_10 | 51 / 100 | 54 / 100 |
| all four suites / 400 | 274 | 281 |
The same checkpoint through the same harness on an RTX 4090 scores 73 / 72 / 79 / 48 on the four suites β 272 / 400, against the boards' 274 and 281. Twelve 100-episode evaluations, same env, same init states, same seeds, same wire, with only the side computing the chunk changed.
The sign flips between suites (the boards are behind on libero_spatial and
ahead on the other three), which is what n = 100 noise looks like; a systematic
porting loss would move all four the same way. The per-task profile is shared as
well β libero_goal task 6 is 5 on all three, libero_10 task 0 is 1 on all
three.
libero_10 is the checkpoint, not the board. It sits at 48β54 where the
other suites are 68β81, and the reference is the lowest of the three. Scored
without a reference column, that 51 would read as "the port collapses on
long-horizon tasks".
Read these numbers as this checkpoint's, not as a porting loss.
Four denoising steps, not ten
These packages ship 4 steps where the reference uses 10. The action expert runs once per step, so this is the one latency lever that costs no rebuild: 741 β 562 ms on S100P, 285 β 220 ms on S600.
Measured across three paired comparisons rather than one, because a single 100-episode result has an interval of about Β±9 points β wider than the effect:
| suite / board | 10 steps | 4 steps |
|---|---|---|
| libero_spatial Β· S100P | 70 | 68 |
| libero_spatial Β· S600 | 71 | 71 |
| libero_object Β· S100P | 76 | 76 |
| total, 300 episodes | 217 | 215 |
Two of the three pairs are identical. If you want the reference schedule, rebuild
the package with bllm-make-model-dir smolvla --steps 10 and a 10-row
cond_table β note that truncating this package's 4-row table would give a
different schedule, not a shorter one (10 steps visit t = 1.0, 0.9, 0.8 β¦;
4 steps visit 1.0, 0.75, 0.5, 0.25).
Use it
hf download ruisv/bllm-smolvla-libero-512 --include "s100p/*" --local-dir .
import bllm, numpy as np
policy = bllm.load_policy("s100p") # or "s600"
actions = policy.act(scene_rgb, # uint8[h, w, 3], any resolution
wrist_rgb, # uint8[h, w, 3]
"pick up the black bowl and place it on the plate",
state=eef_state) # float32[8]
# actions: float32[50, 7] β end-effector deltas in the dataset's own units,
# already unnormalised. Feed them to the arm one row at a time.
There is also a C++ entry point, bllm_smolvla_policy, which does the same from
the shell.
Runtime requirement.
arch: "smolvla"support is newer than the current releasedbllmpackage β build BLLM from source until a release including it is out. A conda-installedbllm0.4.2 will report an unknown arch.
Proprioception is not optional
Unlike Ο0.5's LIBERO configuration, SmolVLA reads the state: embed_prefix
projects it into a prefix token. Omitting it is a different observation, not a
missing convenience β measured at 2.0 in action space, a full gripper unit.
The 8 values are LIBERO's own: end-effector position (3), the end-effector quaternion converted to axis-angle (3), and the gripper joint positions (2).
Pinning the latent
SmolVLA is a flow model integrated from a Gaussian, so two calls on the same
observation legitimately return different actions β and here by a lot: the
reference disagrees with itself at a position max|Ξ| of 0.408 across
latents. Pass noise= when two runs need to be comparable:
z = np.random.default_rng(0).standard_normal((50, 32)).astype(np.float32)
a1 = policy.act(scene, wrist, prompt, state=s, noise=z)
a2 = policy.act(scene, wrist, prompt, state=s, noise=z) # bit-identical
Comparing an implementation against a reference without pinning it measures flow matching, not the implementation.
Accuracy
Same observation, same pinned latent, against the PyTorch reference:
| cosine | βaβ/βa_refβ | max|Ξ| position | |
|---|---|---|---|
| S100P package vs reference | 0.999923 | 0.9971 | 0.0381 |
| S600 package vs reference | 0.999919 | 0.9970 | 0.0398 |
| S600 vs S100P | 0.9999995 | 0.9999 | 0.0017 |
| policy's own spread across latents | 0.9960 | Β±0.009 | 0.408 |
The two boards agree with each other twenty times more closely than either agrees with the reference they were both built from, and both sit an order of magnitude inside the policy's own stochasticity.
Quantisation
int8 per-output-channel weights, int16 static per-tensor activations, on all three
graphs. Each stage is calibrated on the previous stage's compiled output β the
vision .hbm's real tokens set the prefix graph's ranges, the prefix .hbm's real
KV set the expert's β rather than on float intermediates.
Replaying that exact scheme inside the reference model on a host costs 0.09 % of action magnitude. 16-bit activations are a decision rather than a default: at int8 the same replay flips the gripper (cosine 0.9926, position 0.291).
What is in the package
| file | why the runtime cannot infer it |
|---|---|
visual.hbm, model.hbm, expert.hbm |
SmolVLM vision tower, SmolLM2 prefix, interleaved action expert |
embed_tokens.bin |
the tied embedding table, 49280 Γ 960 fp16, pre-scaled by sqrt(960) (see below) |
tokenizer.json |
SmolVLM2's tokenizer |
state_proj.f32 |
the state projection, [960, 32] + bias |
cond_table.f32 |
flow-matching conditioning, precomputed per denoising step |
norm_stats (in model.json) |
MEAN_STD for the 8 state and 7 action dims |
smolvla.cameras, .prompt_len, .horizon, .steps |
compile-time shapes |
embed_prescaled is not bookkeeping. embed_prefix multiplies the tied
embedding row by sqrt(hidden) after calling embed_language_tokens, and does
the same to the image tokens. A package shipping the raw rows makes the prompt 31Γ
too small β a policy that can barely read its own instruction β and nothing catches
it, because cosine is scale-invariant and a golden built the same way carries the
same error. The table is therefore stored pre-scaled and the runtime refuses a
manifest that does not say so.
Limits
- One configuration. 2 cameras, 48 language slots, horizon 50, state 8, action 7 β all compile-time shapes. A finetune with three cameras or a longer instruction budget needs its own build. All 40 LIBERO instructions fit in 22 of the 48 slots; a longer one is a hard error, never a silent truncation.
- Mixed BPU core counts on S600. The vision tower is compiled for four
cores (59.0 β 29.8 ms per camera), the prefix and expert for one β the runtime
reads each graph's compiled count independently, so one package can mix them.
Output is bit-identical to an all-single-core build and
libero_spatialis identical episode for episode (71/100, 355 policy calls, both): multi-core changes scheduling, not arithmetic. The prefix gains only 7 % at two cores and the expert does not compile at four, which is why they stay at one. S100P has a single BPU core, so none of this applies there. - Image preprocessing is a plain bilinear resize with aspect-preserving padding on the left and top (SmolVLA's convention, not openpi's centred one). Feed 512Γ512 directly to remove the question.
Provenance and terms
These are derived weights: lerobot/smolvla_libero compiled for the BPU.
| licence | |
|---|---|
lerobot/smolvla_libero (the checkpoint compiled here) |
Apache-2.0 |
lerobot/libero (source of the normalisation statistics) |
Apache-2.0 |
HuggingFaceTB/SmolVLM2-500M-Video-Instruct (tokenizer) |
Apache-2.0 |
lerobot/smolvla_base (upstream base of the finetune) |
not declared |
The checkpoint these are derived from declares Apache-2.0, which is the operative statement; the base model declaring nothing is recorded here rather than glossed. Check upstream terms before redistributing.
Conversion method and acceptance record:
bllm-model-zoo β models/smolvla-libero/.
Model tree for ruisv/bllm-smolvla-libero-512
Base model
HuggingFaceTB/SmolLM2-360M