qwen3.8-27b-p150x2

Qwen/Qwen3.8-27B served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: tp2 (default, tensor-parallel over both chips), p1d1 (one chip prefills, the other decodes), tp2-dflash2 (tp2 + DFlash2 speculative decoding). Sibling package for the 4-chip P300x2: qwen3.8-27b-p300x2 (profiles plain, batch8-dflash2, single-user-dflash2).

Runs on p150x2 or p150x2 or p150x2 — see the serve profiles below.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  changh95/qwen3.8-27b-p150x2 --with-weights
tt-model serve changh95/qwen3.8-27b-p150x2

pull --with-weights downloads the Docker image and the Qwen/Qwen3.8-27B weights at 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Run

tt-model pull  changh95/qwen3.8-27b-p150x2 --with-weights
tt-model serve changh95/qwen3.8-27b-p150x2                        # tp2 (default)
tt-model serve changh95/qwen3.8-27b-p150x2 --profile p1d1
tt-model serve changh95/qwen3.8-27b-p150x2 --profile tp2-dflash2  # needs: hf download incoai/Qwen3.8-27B-DFlash2
  • OpenAI-compatible API on port 20000, model id Qwen/Qwen3.8-27B; streaming, tool calls (qwen3_coder) and reasoning content (qwen3) work as in vLLM; "chat_template_kwargs": {"enable_thinking": false} turns thinking off.
  • First boot converts the weights (6-8 min); later boots take about 2 min.
curl -s http://localhost:20000/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model": "Qwen/Qwen3.8-27B", "messages": [{"role": "user", "content": "Explain KV caching in two sentences."}], "max_tokens": 256}'

Profiles at a glance

profile users context KV pool 1 user tok/s TTFT 128 / 8k / 32k tokens pick it for
tp2 (default) 32 64K 622K tokens 17.5 0.2 s / 1.8 s / 6.6 s speed, long prompts, many users, all sampling options
p1d1 8 64K 164K tokens 13.7 0.3 s / 3.8 s / 16 s steady per-token latency while long prompts keep arriving (decode never pauses for a prefill)
tp2-dflash2 4 64K 262K tokens 41.9 0.15 s / 1.9 s / 6.8 s 1-4 greedy users, code, prompts under ~16k tokens
  • tp2 is 1.3x faster per user than p1d1 at 1 user and has 3.8x the KV pool; p1d1 keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where tp2 at 8 users degrades to 134-255 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens p1d1 delivers 84-90% of tp2's aggregate throughput.
  • tp2-dflash2 is lossless (greedy trajectory of tp2 up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.9x at 8k, 1.4x at 32k) and grows with answer length (3.1-3.2x at 1k-token answers).

Performance

Method: tt-inference-server --workflow benchmarks grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). t/s = aggregate output tokens/s, t/s/u = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.

tp2: throughput, t/s (t/s/u)

ISL / OSL 1 user 2 4 8 16 32
128 / 128 17 (17.5) 33 (16.3) 58 (14.4) 100 (12.5) 162 (10.1) 245 (7.7)
1,024 / 128 17 (17.1) 32 (15.8) 54 (13.5) 90 (11.2) 138 (8.6) 192 (6.0)
2,048 / 128 17 (16.8) 30 (15.2) 52 (13.0) 80 (10.0) 118 (7.4) 154 (4.8)
4,096 / 128 16 (15.9) 28 (13.8) 45 (11.2) 64 (8.0) 84 (5.3) 103 (3.2)
8,192 / 128 14 (14.3) 23 (11.6) 35 (8.7) 45 (5.6) 54 (3.4) 61 (1.9)
16,384 / 128 12 (12.3) 18 (9.0) 24 (6.0) 28 (3.6) 32 (2.0) 34 (1.1)
32,768 / 128 9 (9.2) 12 (6.0) 14 (3.5) 15 (1.9) 16 (1.0) 16 (0.9) @18
128 / 1,024 18 (17.7) 34 (17.1) 62 (15.6) 116 (14.5) 205 (12.8) 368 (11.5)
8,192 / 1,024 17 (17.1) 32 (16.0) 59 (14.8) 97 (12.1) 151 (9.4) 219 (6.8)

@18: the KV pool seats 18 prompts of that length.

tp2: TTFT ms / TPOT ms

ISL / OSL 1 user 8 users 32 users
128 / 128 197 / 56 1,436 / 70 6,272 / 82
2,048 / 128 501 / 56 3,600 / 73 15,447 / 88
8,192 / 128 1,778 / 57 12,352 / 81 34,181 / 260
32,768 / 128 6,631 / 58 34,573 / 255 -
128 / 1,024 197 / 56 1,444 / 68 6,273 / 81

p1d1: throughput, t/s (t/s/u)

ISL / OSL 1 user 2 4 8
128 / 128 14 (13.7) 26 (12.8) 48 (12.2) 87 (11.0)
1,024 / 128 13 (12.9) 24 (12.1) 44 (11.3) 76 (10.0)
2,048 / 128 12 (12.5) 23 (11.7) 42 (10.9) 72 (9.6)
4,096 / 128 12 (11.5) 21 (10.7) 37 (9.8) 56 (7.6)
8,192 / 128 10 (9.9) 17 (8.9) 26 (7.4) 29 (4.5)
16,384 / 128 8 (7.6) 12 (6.7) 15 (4.3) 16 (2.5)
32,768 / 128 5 (5.0) 6 (3.8) 7 (2.6) -
128 / 1,024 14 (14.0) 27 (13.5) 52 (13.1) 98 (12.3)
8,192 / 1,024 13 (13.2) 25 (12.5) 46 (11.8) 79 (10.4)

p1d1: TTFT ms / TPOT ms

ISL / OSL 1 user 8 users
128 / 128 339 / 71 628 / 86
2,048 / 128 1,161 / 71 2,303 / 87
8,192 / 128 3,819 / 72 18,507 / 79
32,768 / 128 16,308 / 74 -
128 / 1,024 347 / 71 747 / 80

tp2-dflash2 vs tp2, 1 user

ISL / OSL tp2 t/s/u tp2-dflash2 t/s/u speed-up tp2-dflash2 TTFT / TPOT ms
128 / 128 17.5 41.9 2.39x 149 / 23
2,048 / 128 16.8 33.7 2.01x 517 / 26
8,192 / 128 14.3 26.6 1.86x 1,874 / 23
32,768 / 128 9.2 12.7 1.38x 6,778 / 26
128 / 1,024 17.7 56.8 3.21x 149 / 18
8,192 / 1,024 17.1 52.6 3.08x 1,767 / 17
  • 4 users: 113 vs 58 t/s aggregate at 128/128 (1.9x), 151 vs 62 at 128/1,024 (2.4x), 47 vs 35 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.

Notes

  • Not yet run on real p150a cards: all numbers above come from the P300x2 emulation (two chips on different cards, one ethernet cable = 2 links). tp2 and tp2-dflash2 ship a 2-channel mesh descriptor (models/demos/blackhole/qwen36/mesh/p150_x2_2link_mesh_graph_descriptor.textproto): the channel count is the minimum number of trained links the fabric requires and it uses every link it finds, so a pair joined by one cable (2 links) or two cables (4 links) both boot. tt-metal's stock p150_x2 descriptor asks for 4 links and refuses a one-cable pair at start-up (Expected 4 eth links). p1d1 uses a separate two-mesh descriptor whose link count is TT_PD_CONNS (default 2 = one cable).
  • tp2-dflash2 produces greedy output only and serves at most 4 users; tp2 and p1d1 support every vLLM sampling option.
  • Speculative decoding is not available in p1d1.

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh max_num_seqs max_model_len
tp2 (default) p150x2 P150x2 32 65536
p1d1 p150x2 P150x2 8 65536
tp2-dflash2 p150x2 P150x2 4 65536

Provenance

The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal fee9e0d35948111be29083c4eb144023d86890d8
vLLM v0.26.0
vllm-tt-plugin ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7
code/ digest b0c865d431db3b43 (sha256, first 16 hex digits)
built 2026-10-02T20:06:49+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including changh95/qwen3.8-27b-p150x2