Qwen3.8-27B with DFlash2 on two Blackhole chips, 262K context (v6 thin bundle)

Qwen/Qwen3.8-27B served with vLLM on 2 chips (one p300-class board), with DFlash2 speculative decoding (incoai/Qwen3.8-27B-DFlash2), at the model's full 262,144-token context and up to 4 concurrent users. Packaged as a v6 thin tt-model bundle, which is a beta, unsupported format: flags and layout may change. Not reviewed or endorsed by anyone but the person who built and measured it. A single-chip sibling with a 16K context is episod/qwen3.8-27b-dflash2-p150.

Hardware 2 chips of one p300-class board (mesh P150x2), TP=2, FABRIC_1D; firmware bundle 19.15.0; tt-kmd 2.11.0
Context / users 262,144 tokens, up to 4 concurrent users (bf8 KV)
Decoding DFlash2 speculative decoding, fused GDN verify, greedy only
Weights Qwen/Qwen3.8-27B@1d4bf0f2 (bfp4 gate/up, bfp8 rest); drafter incoai/Qwen3.8-27B-DFlash2@dedf8df6; vision disabled

Does it work for coding agents?

The person who built this tried it with qwencode and judged it a success: it could read a project directory, answer questions about it, and reason about two repositories (tt-model-manager and tenstorrent/skills) and whether a model could be brought up with them. That was one person's informal trial of read-and-reason work. Tool versions and prompts were not recorded, and it is not a benchmark: long multi-step editing sessions were not evaluated. The measured numbers below are decode speed, long-context retrieval and a small pass@1 check.

Use

tt-model pull  episod/qwen3.8-27b-dflash2-p300
tt-model serve episod/qwen3.8-27b-dflash2-p300          # OpenAI-compatible server, model id Qwen/Qwen3.8-27B

The bundle builds its own venv (install.sh: interpreter, ttnn, tt-metal-models, vLLM 0.26.0 built for the empty target, the plugin). Host needs: two Blackhole chips of the same board with firmware and driver, 1G hugepages mounted at /dev/hugepages-1G, the SFPI toolchain (/opt/tenstorrent/sfpi), network access for the install, and roughly 60 GB of disk for weights plus tens of GB for caches. Tool calling (qwen3_coder) and reasoning (qwen3) parsers are enabled; temperature and sampling controls are not supported by this decoder.

What was verified on this bundle (2026-09-30)

  • install.sh completed on a clean directory; run.sh booted to Application startup complete on board 0 of a 2-board, 4-chip machine. Cold kernel compilation made that first boot take about 25 minutes.
  • The published repo was then pulled from the Hub and served with the stock commands (tt-model pull, then tt-model serve, a fresh install into ~/.cache/tt-model/models/): it booted to Application startup complete (about 30 minutes from launch, again mostly cold kernel compilation) with the same per-chip memory as in the next bullet. Passkey on that server: 2,075 tokens correct, 132,784 tokens correct, 255,924 tokens correct; the first 20 coding prompts ran at 87.2 tok/s per user at 1 user. tt-model serve listened on port 20000, not 8000.
  • Memory per chip after prefill warmup: 24.90 of 30.83 GiB allocated, 5.93 GiB free.
  • Passkey retrieval on the packaged server (one trial each):
prompt tokens passkey retrieved wall (prefill + 24 tokens)
2,075 yes 5 s
13,321 yes 8 s
132,784 yes 46 s
255,924 yes 111 s
  • The first 20 SPEED-Bench coding prompts on the packaged server: 86.9 tok/s per user at 1 user and 57.2 at 4 users (190.7 aggregate). On the same 20 prompts the source build measured 85.9 and 55.9, with identical output token counts (8,827), so the bundle matches it. These 20 prompts are faster than the 80-prompt average below; do not compare the two sets.
  • Not verified on the bundle: HumanEval and the full grid (those are source-build numbers below), any host other than the one above, and cold weight conversion and the weight download: both bundle runs reused a ttnn tensor cache already produced by the source build and the Hugging Face cache already on the machine (TT_CACHE_PATH / HF_HOME overrides), because the test disk was nearly full. A first run on a clean machine will download about 54 GB and convert the weights before the compile step.

Performance

Measured on the equivalent source-tree build of the same code (the wheels in this bundle were built from that tree), not on this packaged bundle. vllm bench serve, /v1/completions, random prompts of exactly ISL tokens, --ignore-eos, temperature 0; t/s aggregate output tokens per second, (t/s/u) per user = OSL divided by mean end-to-end latency, so it includes prefill (which is why long-ISL rows look slow; TPOT is decode only). Means of 2-8 requests; - = pool cannot seat that many.

ISL / OSL 1 user 2 users 4 users
128 / 128 46.8 (46.8) 65.8 (32.9) 104 (26.1)
1,024 / 128 60.2 (60.2) 88.5 (44.3) 123 (30.7)
2,048 / 128 51.2 (51.2) 79.1 (39.6) 106 (26.5)
4,096 / 128 41.0 (41.0) 58.0 (29.0) 76.8 (19.2)
8,192 / 128 33.5 (33.5) 42.7 (21.4) 49.2 (12.3)
16,384 / 128 23.8 (23.8) 27.7 (13.9) 30.4 (7.6)
32,768 / 128 12.5 (12.5) 13.9 (6.9) 14.9 (3.7)
65,536 / 128 6.5 (6.5) 6.8 (3.4) -
131,072 / 128 2.8 (2.8) - -
200,000 / 128 1.6 (1.6) - -
128 / 1,024 68.3 (68.3) 103 (51.5) 160 (40.1)
8,192 / 1,024 69.0 (69.0) 107 (53.6) 146 (36.5)

TTFT ms / TPOT ms:

ISL / OSL 1 user TTFT ms / TPOT ms 2 users TTFT ms / TPOT ms 4 users TTFT ms / TPOT ms
128 / 128 151 / 20.3 227 / 28.9 479 / 34.8
1,024 / 128 301 / 14.4 447 / 19.3 960 / 25.2
2,048 / 128 461 / 16.0 688 / 20.0 1,480 / 26.4
4,096 / 128 904 / 17.5 1,349 / 24.1 2,907 / 29.6
8,192 / 128 1,816 / 15.8 2,716 / 25.8 5,886 / 35.6
16,384 / 128 3,720 / 13.1 5,586 / 28.7 12,041 / 38.0
32,768 / 128 7,885 / 18.3 11,779 / 52.5 25,457 / 70.6
65,536 / 128 17,462 / 18.2 26,155 / 88.6 -
131,072 / 128 42,041 / 26.9 - -
200,000 / 128 75,937 / 28.8 - -
128 / 1,024 151 / 14.5 228 / 19.2 487 / 24.5
8,192 / 1,024 1,820 / 12.7 2,725 / 16.0 5,879 / 21.6

Long-context passkey retrieval, source build, one trial each (prefill time dominates the wall time):

prompt tokens passkey retrieved wall (prefill + 24 tokens)
31,821 yes 8 s
64,889 yes 18 s
101,848 yes 31 s
130,548 yes 42 s
204,619 yes 78 s
204,495 yes 78 s
256,154 yes 109 s

SPEED-Bench coding (80 single-turn prompts, so this measures decode speed, not agentic use; greedy, thinking off, 2,048-token cap; evaluation-license dataset, not redistributed):

users per-user tok/s mean TPOT aggregate t/s mean TTFT hit cap
1 78.9 12.67 ms 64.8 184 ms 7
4 45.9 21.78 ms 152.6 711 ms 7

Plain (non-speculative) decode on the same two chips runs at about 17 tok/s per user (128/128, 1 user). HumanEval pass@1 (164 problems, greedy, thinking off, sandbox with no network): 149/164 = 90.9% (5 hit the 1,024-token cap); plain decode on the same harness scored 150/164. A single greedy sample per problem on easy problems: a sanity check, not a quality evaluation.

Limits and caveats

  • Prefill is slow at long context: about 42 s for 130K tokens and 109 s for 256K. Decode speed also falls with context (grid above).
  • Two chips of the same board only. Selecting the two chips of the second board of the test machine through TT_VISIBLE_DEVICES failed in ttnn.get_num_devices() (unordered_map::at), for this build and for a source build. The first board and single-chip picks worked. The cause was not found.
  • Greedy only; at most 4 users. DFlash2 output is not byte-identical to plain greedy decoding (on the 20-prompt agreement check 11 of 20 were identical).
  • Vision is disabled (QWEN36_SKIP_VISION=1).
  • Needs TT_METAL_PINNED_MEMORY_CACHE_LIMIT_BYTES=0 (set in run.sh): without it weight load hangs on the test box, in this tt-metal build.
  • run.sh was edited after generation to add the fixed vLLM arguments (--additional-config, prefill chunk, parsers, --no-async-scheduling), because package-thin has no flag for them.
  • install.sh and run.sh are not marked executable in the bundle. tt-model pull / tt-model serve handle that; if you run them by hand, use bash install.sh and bash run.sh.
  • tt-model serve binds the API to 0.0.0.0:20000 (all interfaces) on the test machine; put it behind a firewall or check the serve options if that matters to you.
  • The wheels come from a tree that also carries a local patch allowing a single-chip DFlash configuration (episod/qwen3.8-27b-dflash2-p150); this two-chip bundle does not use it.

Provenance

tt-metal tenstorrent/tt-metal branch qwen36-p150x2 at 9f6b02d2d779 (2026-09-29); wheels versioned 0.79.0.dev20260929+qwen36p150x2.g9f6b02d
Plugin changh95/vllm-tt-plugin ab5f7f3d795e (Apache-2.0)
vLLM 0.26.0, built with VLLM_TARGET_DEVICE=empty at install
Python 3.12; transformers 5.17.0, tokenizers 0.23.2, CPU torch

Licenses

Base model and drafter: Apache-2.0. tt-metal and vllm-tt-plugin: Apache-2.0. This bundle redistributes built ttnn and tt-metal-models wheels from the tt-metal sources above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for episod/qwen3.8-27b-dflash2-p300

Base model

Qwen/Qwen3.8-27B
Finetuned
(456)
this model