Qwen3.5-0.8B โ BPU-native build for D-Robotics RDK S100 / S100P
Compiled, quantised (int8 weights / int16 activations), and board-verified build
of Qwen/Qwen3.5-0.8B for D-Robotics'
Nash BPU, running entirely through BLLM โ an
on-device runtime that talks to the BPU directly (hbDNN/hbUCP), with no
GPU/CUDA and no cloud round-trip.
This is a compiled BPU binary, not a transformers-loadable checkpoint โ
closer in kind to a GGUF quantisation than a HF AutoModel. There is no
Inference API widget on this page; see "How to run" below.
New here? Start with BLLM (the runtime) and bllm-model-zoo (how this build was made, full acceptance data, what was tried and rejected). This repo is the downloadable artifact; that repo is the method.
Quick facts
| Base model | Qwen/Qwen3.5-0.8B (Apache-2.0) |
| Architecture | Hybrid: 18 Gated-DeltaNet (linear-attention SSM) layers + 6 full-attention layers |
| Quantisation | int8 weight (per-channel) / int16 activation (static, data-free) |
| Target hardware | D-Robotics RDK S100 (s100/, nash-e) and S100P (s100p/, nash-m) โ two separate builds, see below |
| Modality | Image + text (vision tower included) |
| BPU utilization | 100% โ CPU_inference_time_cost == 0 on the text decoder |
| Runtime | BLLM (native, no libxlm) |
This architecture (linear-attention SSM + head_dim=256) is not covered by
D-Robotics' official OE-LLM toolchain (libxlm) โ this build exists because
BLLM's native runtime compiles the graph directly. It is the smallest member of
the family โ see also ruisv/bllm-qwen3.5-2b
and ruisv/bllm-qwen3.5-4b.
Files in this repo
s100p/ # RDK S100P (Journey-6 J6M, 128 TOPS) โ compiled march nash-m
ctx4k/ # cache_len=4096 โ default, recommended context
model.hbm # text decoder (prefill+decode dual-graph)
model_prefill.hbm # chunked-prefill graph (N=32)
model.json
embed_tokens.bin
tokenizer.json
visual.hbm # vision tower, 320px bucket, 100 tokens/image
ctx512/ # cache_len=512 โ faster decode, short-context use
...same layout
s100/ # RDK S100 (Journey-6 J6E, 80 TOPS) โ compiled march nash-e
ctx4k/ # same six files, same graph and quantisation, different BPU target
ctx512/
Which folder do I need?
A BPU march names a Journey-6 SoC tier, not a board family, and an .hbm
loads only on its own tier:
| march | SoC | board |
|---|---|---|
nash-e |
J6E, 80 TOPS, 1 BPU core | RDK S100 (12 GB) |
nash-m |
J6M, 128 TOPS, 1 BPU core | RDK S100P (24 GB) |
nash-p |
J6P, 4 BPU cores | RDK S600 |
Take s100/ for an RDK S100 and s100p/ for an RDK S100P. Do not read the
march out of the runtime's load-failure message if you are checking a file by
hand โ that message resolves march names against a mismatched enum table and
mislabels them; read it from the .hbm itself (hbdk4.compiler.hbm_extract_desc,
or the nash-x string in the binary).
cache_len is a performance knob as well as a capacity one โ a larger KV
window costs more DDR traffic per decoded token. Pick ctx512 if your
prompts/replies genuinely fit in ~512 tokens; it decodes faster than ctx4k.
How to run
conda install -c https://mirrors.ruis.ai/conda -c conda-forge bllm
import bllm
m = bllm.load("s100p/ctx4k") # RDK S100P; use "s100/ctx4k" on an RDK S100
print(m.chat("่ฟๅผ ๅพ้ๆไปไน?", images=["photo.jpg"]))
# text-only, without loading the vision tower into memory
m = bllm.load("s100p/ctx4k", vision=False)
print(m.chat("ไป็ปไธไธไฝ ่ชๅทฑ"))
Requires a D-Robotics RDK S100 or S100P (matching folder, see above),
sw >= 4.0.5, and an enlarged ION carveout (hb_switch_ion.sh balanced +
reboot) โ see BLLM's docs/LLM_ONBOARD.md.
Measured on hardware
s100p/ โ RDK S100P, performance mode, 3 repeated runs
| variant | decode | TTFT (image) | prefillโdecode parity | ppl | image contrast | task anchors |
|---|---|---|---|---|---|---|
| ctx4096 | 18.52 tok/s | 1.32 s | PASS | 83.09 | PASS | PASS* |
| ctx512 | 23.18 tok/s | 1.19 s | PASS | 83.60 | PASS | PASS |
s100/ โ the nash-e build, and how far it is verified
The S100 build was compiled from the identical graph, weights and quantisation
as the S100P one; only the BPU code generation differs. It was accepted by
running it on S100P silicon (a nash-e binary loads and runs there) against
the S100P build as a same-session baseline, on the same board, same harness,
same prompts, greedy decoding:
| build | decode | TTFT (image) | parity | ppl | image answer | contrast | anchors |
|---|---|---|---|---|---|---|---|
s100/ctx4k (nash-e) |
18.55 tok/s | 1.478 s | PASS | 83.0899 | correct | PASS | PASS* |
s100p/ctx4k (nash-m, baseline) |
18.54 tok/s | 1.479 s | PASS | 83.0899 | correct | PASS | PASS* |
s100/ctx512 (nash-e) |
23.04 tok/s | 1.321 s | PASS | 83.5979 | correct | PASS | PASS |
s100p/ctx512 (nash-m, baseline) |
23.15 tok/s | 1.318 s | PASS | 83.5979 | correct | PASS | PASS |
Every generated answer was token-for-token identical between the two builds,
on every prompt, and the perplexities match to four decimals โ which is the
strongest statement available that the nash-e code generation did not change
the model.
What is not verified: no RDK S100 board was available, so these numbers are S100P timings, not S100 timings. An S100 is 80 TOPS with a 1.5 GHz CPU (vs 128 TOPS / 2.0 GHz), so expect it to be slower โ by how much is unmeasured, and this card will not guess. The S100 also has 12 GB of RAM against the S100P's 24 GB; at this size that is comfortable, but see the 2B/4B cards before assuming it holds there.
* ctx4096 answers one arithmetic anchor (17ร23) incorrectly โ 401 in the
2026-07-30 run, 411 in the 2026-08-31 run, where both the nash-e and
nash-m builds returned the same wrong 411 (which is itself the parity point);
ctx512 โ same weights, only cache_len differs โ answers it correctly (391) in
both builds and both runs.
This is 0.8B sitting right at its arithmetic capability edge, not a
quantisation or infra bug (everything else, including the identical prompt in
Chinese, passes on both). Full methodology, acceptance protocol
(prefillโdecode parity + perplexity + task-coherence anchors +
with/without-image contrast โ not just cosine similarity, see why in the
model-zoo README),
and every rejected build are in
bllm-model-zoo/models/qwen3.5-0.8b
and .../qwen3.5-vlm.
Known limitations
- The all-BPU chunked-prefill path (image ingestion, no CPU fallback) is
numerically solvable but impractically slow to compile โ this build's
prefill graph uses a float
[C,C]CPU fallback for one small intermediate instead (the same design as the 2B/4B siblings). A compile-time/build- engineering tradeoff, not an accuracy or runtime limitation of the shipped model. - int16 activation quantisation is data-free (no calibration set) at this size โ verified sufficient at 0.8B; the 4B sibling needs per-linear calibration (see its model card).
License
The compiled .hbm is a derivative of the upstream Qwen3.5-0.8B
weights and inherits its Apache-2.0 license โ check the upstream model card
for the current terms. BLLM (the runtime that loads this file) is open source;
see its repository for its own license.