qwen3.8-27b-p150x2
Qwen/Qwen3.8-27B served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: tp2 (default, tensor-parallel over both chips), p1d1 (one chip prefills, the other decodes), tp2-dflash2 (tp2 + DFlash2 speculative decoding). Sibling package for the 4-chip P300x2: qwen3.8-27b-p300x2 (profiles plain, batch8-dflash2, single-user-dflash2).
Runs on p150x2 or p150x2 or p150x2 — see the serve profiles below.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart
tt-model pull changh95/qwen3.8-27b-p150x2 --with-weights
tt-model serve changh95/qwen3.8-27b-p150x2
pull --with-weights downloads the Docker image and the Qwen/Qwen3.8-27B weights at 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Run
tt-model pull changh95/qwen3.8-27b-p150x2 --with-weights
tt-model serve changh95/qwen3.8-27b-p150x2 # tp2 (default)
tt-model serve changh95/qwen3.8-27b-p150x2 --profile p1d1
tt-model serve changh95/qwen3.8-27b-p150x2 --profile tp2-dflash2 # needs: hf download incoai/Qwen3.8-27B-DFlash2
- OpenAI-compatible API on port 20000, model id
Qwen/Qwen3.8-27B; streaming, tool calls (qwen3_coder) and reasoning content (qwen3) work as in vLLM;"chat_template_kwargs": {"enable_thinking": false}turns thinking off. - First boot converts the weights (6-8 min); later boots take about 2 min.
curl -s http://localhost:20000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model": "Qwen/Qwen3.8-27B", "messages": [{"role": "user", "content": "Explain KV caching in two sentences."}], "max_tokens": 256}'
Profiles at a glance
| profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
|---|---|---|---|---|---|---|
tp2 (default) |
32 | 64K | 622K tokens | 17.5 | 0.2 s / 1.8 s / 6.6 s | speed, long prompts, many users, all sampling options |
p1d1 |
8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
tp2-dflash2 |
4 | 64K | 262K tokens | 41.9 | 0.15 s / 1.9 s / 6.8 s | 1-4 greedy users, code, prompts under ~16k tokens |
tp2is 1.3x faster per user thanp1d1at 1 user and has 3.8x the KV pool;p1d1keeps TPOT at 71-87 ms at 1-8 users across every prompt length, wheretp2at 8 users degrades to 134-255 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokensp1d1delivers 84-90% oftp2's aggregate throughput.tp2-dflash2is lossless (greedy trajectory oftp2up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.9x at 8k, 1.4x at 32k) and grows with answer length (3.1-3.2x at 1k-token answers).
Performance
Method: tt-inference-server --workflow benchmarks grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). t/s = aggregate output tokens/s, t/s/u = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
tp2: throughput, t/s (t/s/u)
| ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 128 / 128 | 17 (17.5) | 33 (16.3) | 58 (14.4) | 100 (12.5) | 162 (10.1) | 245 (7.7) |
| 1,024 / 128 | 17 (17.1) | 32 (15.8) | 54 (13.5) | 90 (11.2) | 138 (8.6) | 192 (6.0) |
| 2,048 / 128 | 17 (16.8) | 30 (15.2) | 52 (13.0) | 80 (10.0) | 118 (7.4) | 154 (4.8) |
| 4,096 / 128 | 16 (15.9) | 28 (13.8) | 45 (11.2) | 64 (8.0) | 84 (5.3) | 103 (3.2) |
| 8,192 / 128 | 14 (14.3) | 23 (11.6) | 35 (8.7) | 45 (5.6) | 54 (3.4) | 61 (1.9) |
| 16,384 / 128 | 12 (12.3) | 18 (9.0) | 24 (6.0) | 28 (3.6) | 32 (2.0) | 34 (1.1) |
| 32,768 / 128 | 9 (9.2) | 12 (6.0) | 14 (3.5) | 15 (1.9) | 16 (1.0) | 16 (0.9) @18 |
| 128 / 1,024 | 18 (17.7) | 34 (17.1) | 62 (15.6) | 116 (14.5) | 205 (12.8) | 368 (11.5) |
| 8,192 / 1,024 | 17 (17.1) | 32 (16.0) | 59 (14.8) | 97 (12.1) | 151 (9.4) | 219 (6.8) |
@18: the KV pool seats 18 prompts of that length.
tp2: TTFT ms / TPOT ms
| ISL / OSL | 1 user | 8 users | 32 users |
|---|---|---|---|
| 128 / 128 | 197 / 56 | 1,436 / 70 | 6,272 / 82 |
| 2,048 / 128 | 501 / 56 | 3,600 / 73 | 15,447 / 88 |
| 8,192 / 128 | 1,778 / 57 | 12,352 / 81 | 34,181 / 260 |
| 32,768 / 128 | 6,631 / 58 | 34,573 / 255 | - |
| 128 / 1,024 | 197 / 56 | 1,444 / 68 | 6,273 / 81 |
p1d1: throughput, t/s (t/s/u)
| ISL / OSL | 1 user | 2 | 4 | 8 |
|---|---|---|---|---|
| 128 / 128 | 14 (13.7) | 26 (12.8) | 48 (12.2) | 87 (11.0) |
| 1,024 / 128 | 13 (12.9) | 24 (12.1) | 44 (11.3) | 76 (10.0) |
| 2,048 / 128 | 12 (12.5) | 23 (11.7) | 42 (10.9) | 72 (9.6) |
| 4,096 / 128 | 12 (11.5) | 21 (10.7) | 37 (9.8) | 56 (7.6) |
| 8,192 / 128 | 10 (9.9) | 17 (8.9) | 26 (7.4) | 29 (4.5) |
| 16,384 / 128 | 8 (7.6) | 12 (6.7) | 15 (4.3) | 16 (2.5) |
| 32,768 / 128 | 5 (5.0) | 6 (3.8) | 7 (2.6) | - |
| 128 / 1,024 | 14 (14.0) | 27 (13.5) | 52 (13.1) | 98 (12.3) |
| 8,192 / 1,024 | 13 (13.2) | 25 (12.5) | 46 (11.8) | 79 (10.4) |
p1d1: TTFT ms / TPOT ms
| ISL / OSL | 1 user | 8 users |
|---|---|---|
| 128 / 128 | 339 / 71 | 628 / 86 |
| 2,048 / 128 | 1,161 / 71 | 2,303 / 87 |
| 8,192 / 128 | 3,819 / 72 | 18,507 / 79 |
| 32,768 / 128 | 16,308 / 74 | - |
| 128 / 1,024 | 347 / 71 | 747 / 80 |
tp2-dflash2 vs tp2, 1 user
| ISL / OSL | tp2 t/s/u |
tp2-dflash2 t/s/u |
speed-up | tp2-dflash2 TTFT / TPOT ms |
|---|---|---|---|---|
| 128 / 128 | 17.5 | 41.9 | 2.39x | 149 / 23 |
| 2,048 / 128 | 16.8 | 33.7 | 2.01x | 517 / 26 |
| 8,192 / 128 | 14.3 | 26.6 | 1.86x | 1,874 / 23 |
| 32,768 / 128 | 9.2 | 12.7 | 1.38x | 6,778 / 26 |
| 128 / 1,024 | 17.7 | 56.8 | 3.21x | 149 / 18 |
| 8,192 / 1,024 | 17.1 | 52.6 | 3.08x | 1,767 / 17 |
- 4 users: 113 vs 58 t/s aggregate at 128/128 (1.9x), 151 vs 62 at 128/1,024 (2.4x), 47 vs 35 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
Notes
- Not yet run on real p150a cards: all numbers above come from the P300x2 emulation (two chips on different cards, one ethernet cable = 2 links).
tp2andtp2-dflash2ship a 2-channel mesh descriptor (models/demos/blackhole/qwen36/mesh/p150_x2_2link_mesh_graph_descriptor.textproto): the channel count is the minimum number of trained links the fabric requires and it uses every link it finds, so a pair joined by one cable (2 links) or two cables (4 links) both boot. tt-metal's stockp150_x2descriptor asks for 4 links and refuses a one-cable pair at start-up (Expected 4 eth links).p1d1uses a separate two-mesh descriptor whose link count isTT_PD_CONNS(default 2 = one cable). tp2-dflash2produces greedy output only and serves at most 4 users;tp2andp1d1support every vLLM sampling option.- Speculative decoding is not available in
p1d1.
Serve profiles
One image serves every profile below; pick one with --profile.
| profile | hardware | mesh | max_num_seqs | max_model_len |
|---|---|---|---|---|
tp2 (default) |
p150x2 | P150x2 | 32 | 65536 |
p1d1 |
p150x2 | P150x2 | 8 | 65536 |
tp2-dflash2 |
p150x2 | P150x2 | 4 | 65536 |
Provenance
The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | fee9e0d35948111be29083c4eb144023d86890d8 |
| vLLM | v0.26.0 |
| vllm-tt-plugin | ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 |
code/ digest |
b0c865d431db3b43 (sha256, first 16 hex digits) |
| built | 2026-10-02T20:06:49+00:00 by tt-model 0.1.0 |