Qwen3.8-27B with DFlash2 on two Blackhole chips, 262K context (v6 thin bundle)
Qwen/Qwen3.8-27B served with vLLM on 2 chips (one p300-class board), with
DFlash2 speculative decoding (incoai/Qwen3.8-27B-DFlash2),
at the model's full 262,144-token context and up to 4 concurrent users.
Packaged as a v6 thin tt-model bundle, which is a beta, unsupported format: flags and layout may change.
Not reviewed or endorsed by anyone but the person who built and measured it.
A single-chip sibling with a 16K context is episod/qwen3.8-27b-dflash2-p150.
| Hardware | 2 chips of one p300-class board (mesh P150x2), TP=2, FABRIC_1D; firmware bundle 19.15.0; tt-kmd 2.11.0 |
| Context / users | 262,144 tokens, up to 4 concurrent users (bf8 KV) |
| Decoding | DFlash2 speculative decoding, fused GDN verify, greedy only |
| Weights | Qwen/Qwen3.8-27B@1d4bf0f2 (bfp4 gate/up, bfp8 rest); drafter incoai/Qwen3.8-27B-DFlash2@dedf8df6; vision disabled |
Does it work for coding agents?
The person who built this tried it with qwencode and judged it a success: it could read a project directory, answer questions about it,
and reason about two repositories (tt-model-manager and tenstorrent/skills) and whether a model could be brought up with them.
That was one person's informal trial of read-and-reason work. Tool versions and prompts were not recorded, and it is not a benchmark:
long multi-step editing sessions were not evaluated. The measured numbers below are decode speed, long-context retrieval and a small pass@1 check.
Use
tt-model pull episod/qwen3.8-27b-dflash2-p300
tt-model serve episod/qwen3.8-27b-dflash2-p300 # OpenAI-compatible server, model id Qwen/Qwen3.8-27B
The bundle builds its own venv (install.sh: interpreter, ttnn, tt-metal-models, vLLM 0.26.0 built for the empty target, the plugin).
Host needs: two Blackhole chips of the same board with firmware and driver, 1G hugepages mounted at /dev/hugepages-1G, the SFPI toolchain (/opt/tenstorrent/sfpi),
network access for the install, and roughly 60 GB of disk for weights plus tens of GB for caches.
Tool calling (qwen3_coder) and reasoning (qwen3) parsers are enabled; temperature and sampling controls are not supported by this decoder.
What was verified on this bundle (2026-09-30)
install.shcompleted on a clean directory;run.shbooted toApplication startup completeon board 0 of a 2-board, 4-chip machine. Cold kernel compilation made that first boot take about 25 minutes.- The published repo was then pulled from the Hub and served with the stock commands (
tt-model pull, thentt-model serve, a fresh install into~/.cache/tt-model/models/): it booted toApplication startup complete(about 30 minutes from launch, again mostly cold kernel compilation) with the same per-chip memory as in the next bullet. Passkey on that server: 2,075 tokens correct, 132,784 tokens correct, 255,924 tokens correct; the first 20 coding prompts ran at 87.2 tok/s per user at 1 user.tt-model servelistened on port 20000, not 8000. - Memory per chip after prefill warmup: 24.90 of 30.83 GiB allocated, 5.93 GiB free.
- Passkey retrieval on the packaged server (one trial each):
| prompt tokens | passkey retrieved | wall (prefill + 24 tokens) |
|---|---|---|
| 2,075 | yes | 5 s |
| 13,321 | yes | 8 s |
| 132,784 | yes | 46 s |
| 255,924 | yes | 111 s |
- The first 20 SPEED-Bench coding prompts on the packaged server: 86.9 tok/s per user at 1 user and 57.2 at 4 users (190.7 aggregate). On the same 20 prompts the source build measured 85.9 and 55.9, with identical output token counts (8,827), so the bundle matches it. These 20 prompts are faster than the 80-prompt average below; do not compare the two sets.
- Not verified on the bundle: HumanEval and the full grid (those are source-build numbers below), any host other than the one above, and
cold weight conversion and the weight download: both bundle runs reused a ttnn tensor cache already produced by the source build and the Hugging Face cache already on
the machine (
TT_CACHE_PATH/HF_HOMEoverrides), because the test disk was nearly full. A first run on a clean machine will download about 54 GB and convert the weights before the compile step.
Performance
Measured on the equivalent source-tree build of the same code (the wheels in this bundle were built from that tree), not on this packaged bundle.
vllm bench serve, /v1/completions, random prompts of exactly ISL tokens, --ignore-eos, temperature 0; t/s aggregate output tokens per second, (t/s/u) per user = OSL divided by mean end-to-end latency, so it includes prefill (which is why long-ISL rows look slow; TPOT is decode only). Means of 2-8 requests; - = pool cannot seat that many.
| ISL / OSL | 1 user | 2 users | 4 users |
|---|---|---|---|
| 128 / 128 | 46.8 (46.8) | 65.8 (32.9) | 104 (26.1) |
| 1,024 / 128 | 60.2 (60.2) | 88.5 (44.3) | 123 (30.7) |
| 2,048 / 128 | 51.2 (51.2) | 79.1 (39.6) | 106 (26.5) |
| 4,096 / 128 | 41.0 (41.0) | 58.0 (29.0) | 76.8 (19.2) |
| 8,192 / 128 | 33.5 (33.5) | 42.7 (21.4) | 49.2 (12.3) |
| 16,384 / 128 | 23.8 (23.8) | 27.7 (13.9) | 30.4 (7.6) |
| 32,768 / 128 | 12.5 (12.5) | 13.9 (6.9) | 14.9 (3.7) |
| 65,536 / 128 | 6.5 (6.5) | 6.8 (3.4) | - |
| 131,072 / 128 | 2.8 (2.8) | - | - |
| 200,000 / 128 | 1.6 (1.6) | - | - |
| 128 / 1,024 | 68.3 (68.3) | 103 (51.5) | 160 (40.1) |
| 8,192 / 1,024 | 69.0 (69.0) | 107 (53.6) | 146 (36.5) |
TTFT ms / TPOT ms:
| ISL / OSL | 1 user TTFT ms / TPOT ms | 2 users TTFT ms / TPOT ms | 4 users TTFT ms / TPOT ms |
|---|---|---|---|
| 128 / 128 | 151 / 20.3 | 227 / 28.9 | 479 / 34.8 |
| 1,024 / 128 | 301 / 14.4 | 447 / 19.3 | 960 / 25.2 |
| 2,048 / 128 | 461 / 16.0 | 688 / 20.0 | 1,480 / 26.4 |
| 4,096 / 128 | 904 / 17.5 | 1,349 / 24.1 | 2,907 / 29.6 |
| 8,192 / 128 | 1,816 / 15.8 | 2,716 / 25.8 | 5,886 / 35.6 |
| 16,384 / 128 | 3,720 / 13.1 | 5,586 / 28.7 | 12,041 / 38.0 |
| 32,768 / 128 | 7,885 / 18.3 | 11,779 / 52.5 | 25,457 / 70.6 |
| 65,536 / 128 | 17,462 / 18.2 | 26,155 / 88.6 | - |
| 131,072 / 128 | 42,041 / 26.9 | - | - |
| 200,000 / 128 | 75,937 / 28.8 | - | - |
| 128 / 1,024 | 151 / 14.5 | 228 / 19.2 | 487 / 24.5 |
| 8,192 / 1,024 | 1,820 / 12.7 | 2,725 / 16.0 | 5,879 / 21.6 |
Long-context passkey retrieval, source build, one trial each (prefill time dominates the wall time):
| prompt tokens | passkey retrieved | wall (prefill + 24 tokens) |
|---|---|---|
| 31,821 | yes | 8 s |
| 64,889 | yes | 18 s |
| 101,848 | yes | 31 s |
| 130,548 | yes | 42 s |
| 204,619 | yes | 78 s |
| 204,495 | yes | 78 s |
| 256,154 | yes | 109 s |
SPEED-Bench coding (80 single-turn prompts, so this measures decode speed, not agentic use; greedy, thinking off, 2,048-token cap; evaluation-license dataset, not redistributed):
| users | per-user tok/s | mean TPOT | aggregate t/s | mean TTFT | hit cap |
|---|---|---|---|---|---|
| 1 | 78.9 | 12.67 ms | 64.8 | 184 ms | 7 |
| 4 | 45.9 | 21.78 ms | 152.6 | 711 ms | 7 |
Plain (non-speculative) decode on the same two chips runs at about 17 tok/s per user (128/128, 1 user). HumanEval pass@1 (164 problems, greedy, thinking off, sandbox with no network): 149/164 = 90.9% (5 hit the 1,024-token cap); plain decode on the same harness scored 150/164. A single greedy sample per problem on easy problems: a sanity check, not a quality evaluation.
Limits and caveats
- Prefill is slow at long context: about 42 s for 130K tokens and 109 s for 256K. Decode speed also falls with context (grid above).
- Two chips of the same board only. Selecting the two chips of the second board of the test machine through
TT_VISIBLE_DEVICESfailed inttnn.get_num_devices()(unordered_map::at), for this build and for a source build. The first board and single-chip picks worked. The cause was not found. - Greedy only; at most 4 users. DFlash2 output is not byte-identical to plain greedy decoding (on the 20-prompt agreement check 11 of 20 were identical).
- Vision is disabled (
QWEN36_SKIP_VISION=1). - Needs
TT_METAL_PINNED_MEMORY_CACHE_LIMIT_BYTES=0(set inrun.sh): without it weight load hangs on the test box, in this tt-metal build. run.shwas edited after generation to add the fixed vLLM arguments (--additional-config, prefill chunk, parsers,--no-async-scheduling), becausepackage-thinhas no flag for them.install.shandrun.share not marked executable in the bundle.tt-model pull/tt-model servehandle that; if you run them by hand, usebash install.shandbash run.sh.tt-model servebinds the API to0.0.0.0:20000(all interfaces) on the test machine; put it behind a firewall or check the serve options if that matters to you.- The wheels come from a tree that also carries a local patch allowing a single-chip DFlash configuration (
episod/qwen3.8-27b-dflash2-p150); this two-chip bundle does not use it.
Provenance
| tt-metal | tenstorrent/tt-metal branch qwen36-p150x2 at 9f6b02d2d779 (2026-09-29); wheels versioned 0.79.0.dev20260929+qwen36p150x2.g9f6b02d |
| Plugin | changh95/vllm-tt-plugin ab5f7f3d795e (Apache-2.0) |
| vLLM | 0.26.0, built with VLLM_TARGET_DEVICE=empty at install |
| Python | 3.12; transformers 5.17.0, tokenizers 0.23.2, CPU torch |
Licenses
Base model and drafter: Apache-2.0. tt-metal and vllm-tt-plugin: Apache-2.0. This bundle redistributes built ttnn and tt-metal-models wheels from the tt-metal sources above.
Model tree for episod/qwen3.8-27b-dflash2-p300
Base model
Qwen/Qwen3.8-27B