changh95 commited on
Commit
1466f35
·
verified ·
1 Parent(s): ca51ee5

Add files using upload-large-folder tool

Browse files
Files changed (23) hide show
  1. README.md +31 -33
  2. code/models/demos/blackhole/qwen36/tt-model-p300x2.yaml +327 -0
  3. code/models/demos/blackhole/qwen36/tt-model.yaml +27 -28
  4. code/models/demos/blackhole/qwen36/tt/chunked_prefill.py +114 -0
  5. code/models/demos/blackhole/qwen36/tt/qwen36_vllm.py +127 -7
  6. code/models/demos/blackhole/qwen36/tt/qwen36_vllm_dflash.py +340 -22
  7. image/blobs/sha256/1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2 +0 -0
  8. image/blobs/sha256/45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d +0 -0
  9. image/blobs/sha256/4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9 +0 -0
  10. image/blobs/sha256/58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2 +28 -0
  11. image/blobs/sha256/70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9 +0 -0
  12. image/blobs/sha256/8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146 +121 -0
  13. image/blobs/sha256/aba4bce132b5625823266e9eb71834b271fbc91ad4b8693c3bd34ecf068eb41e +19 -0
  14. image/blobs/sha256/bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7 +0 -0
  15. image/blobs/sha256/d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d +1 -0
  16. image/blobs/sha256/e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691 +0 -0
  17. image/blobs/sha256/ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9 +0 -0
  18. image/blobs/sha256/ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7 +1 -0
  19. image/blobs/sha256/f08a1fdc23ba07b7b094692072f62ca467a03cbe754bc49130dabddf907da186 +1 -0
  20. image/index.json +1 -1
  21. image/manifest.json +1 -1
  22. requirements.lock +24 -28
  23. tt_kernel_manifest.json +11 -11
README.md CHANGED
@@ -3,16 +3,15 @@ tags:
3
  - blackhole
4
  - p150x2
5
  - tt-model-cache
6
- - tt-model-catalog
7
  - tt-model-container
8
  - vllm-plugin
9
  ---
10
 
11
  # qwen3.8-27b-p150x2
12
 
13
- [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling packages for the 4-chip P300x2: [qwen3.8-27b-p300x2](https://huggingface.co/changh95/qwen3.8-27b-p300x2), [qwen3.8-27b-dflash2-p300x2](https://huggingface.co/changh95/qwen3.8-27b-dflash2-p300x2).
14
 
15
- Runs on **p150x2** — see the serve profiles below.
16
 
17
  Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).
18
 
@@ -46,30 +45,30 @@ curl -s http://localhost:20000/v1/chat/completions -H 'Content-Type: application
46
 
47
  | profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
48
  |---|---|---|---|---|---|---|
49
- | `tp2` (default) | 32 | 64K | 622K tokens | 17.1 | 0.3 s / 1.9 s / 8.2 s | speed, long prompts, many users, all sampling options |
50
  | `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
51
- | `tp2-dflash2` | 4 | 64K | 262K tokens | 40.8 (58 on code) | 0.15 s / 2.1 s / 9.2 s | 1-4 greedy users, code, prompts under ~16k tokens |
52
 
53
- * `tp2` is 1.2x faster per user than `p1d1` at 1 user and has 3.8x the KV pool; `p1d1` keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where `tp2` at 8 users degrades to 146-285 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens `p1d1` is within 4-9% of `tp2`'s aggregate throughput.
54
- * `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.6x at 8k, 1.2x at 32k) and grows with answer length (3.3-3.4x at 1k-token answers).
55
 
56
  ### Performance
57
 
58
- Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Code: SPEED-Bench coding, 80 prompts, thinking off, temperature 0, 1 user. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
59
 
60
  #### `tp2`: throughput, t/s (t/s/u)
61
 
62
  | ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
63
  |---|---|---|---|---|---|---|
64
- | 128 / 128 | 17 (17.1) | 29 (14.7) | 54 (13.4) | 92 (11.4) | 143 (9.0) | 193 (6.0) |
65
- | 1,024 / 128 | 17 (16.7) | 28 (14.1) | 50 (12.6) | 83 (10.4) | 123 (7.7) | 156 (4.9) |
66
- | 2,048 / 128 | 16 (16.3) | 27 (13.7) | 48 (11.9) | 75 (9.3) | 107 (6.7) | 131 (4.1) |
67
- | 4,096 / 128 | 15 (15.4) | 25 (12.4) | 41 (10.2) | 59 (7.4) | 78 (4.9) | 90 (2.8) |
68
- | 8,192 / 128 | 14 (13.9) | 21 (10.5) | 31 (7.8) | 41 (5.1) | 50 (3.1) | 54 (1.7) |
69
- | 16,384 / 128 | 11 (11.4) | 16 (7.9) | 21 (5.2) | 25 (3.1) | 28 (1.8) | 29 (0.9) |
70
- | 32,768 / 128 | 8 (8.2) | 10 (5.0) | 12 (3.0) | 13 (1.7) | 14 (0.9) | 14 (0.8) @18 |
71
- | 128 / 1,024 | 18 (17.6) | 33 (16.6) | 65 (16.1) | 123 (15.4) | 218 (13.6) | 361 (11.3) |
72
- | 8,192 / 1,024 | 17 (16.9) | 31 (15.6) | 57 (14.3) | 98 (12.2) | 153 (9.6) | 210 (6.6) |
73
 
74
  `@18`: the KV pool seats 18 prompts of that length.
75
 
@@ -77,11 +76,11 @@ Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exac
77
 
78
  | ISL / OSL | 1 user | 8 users | 32 users |
79
  |---|---|---|---|
80
- | 128 / 128 | 316 / 57 | 2,394 / 69 | 10,226 / 87 |
81
- | 2,048 / 128 | 610 / 57 | 4,518 / 72 | 19,498 / 93 |
82
- | 8,192 / 128 | 1,941 / 57 | 14,217 / 84 | 39,474 / 285 |
83
- | 32,768 / 128 | 8,243 / 58 | 40,999 / 285 | - |
84
- | 128 / 1,024 | 317 / 57 | 2,379 / 63 | 10,322 / 79 |
85
 
86
  #### `p1d1`: throughput, t/s (t/s/u)
87
 
@@ -111,15 +110,14 @@ Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exac
111
 
112
  | ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
113
  |---|---|---|---|---|
114
- | 128 / 128 | 17.1 | 40.8 | 2.39x | 152 / 24 |
115
- | 2,048 / 128 | 16.3 | 33.6 | 2.06x | 511 / 26 |
116
- | 8,192 / 128 | 13.9 | 22.6 | 1.63x | 2,060 / 28 |
117
- | 32,768 / 128 | 8.2 | 9.6 | 1.17x | 9,224 / 32 |
118
- | 128 / 1,024 | 17.6 | 58.8 | 3.34x | 154 / 17 |
119
- | 8,192 / 1,024 | 16.9 | 56.7 | 3.36x | 2,060 / 16 |
120
- | code (SPEED-Bench, EOS-terminated) | 16.2 | 58.4 | 3.6x | 194 / 13 |
121
 
122
- * 4 users: 117 vs 54 t/s aggregate at 128/128 (2.2x), 163 vs 65 at 128/1,024 (2.5x), 42 vs 31 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
123
 
124
  ### Notes
125
 
@@ -143,9 +141,9 @@ The exact sources the image was built from — `code/` in this repo is byte-iden
143
 
144
  | component | built from |
145
  | --- | --- |
146
- | tt-metal | [`aa20afaf172ac35370b8d112edc0f771ee6d2571`](https://github.com/tenstorrent/tt-metal/commit/aa20afaf172ac35370b8d112edc0f771ee6d2571) |
147
  | vLLM | [`v0.26.0`](https://github.com/vllm-project/vllm/releases/tag/v0.26.0) |
148
  | vllm-tt-plugin | [`ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7`](https://github.com/changh95/vllm-tt-plugin/commit/ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7) |
149
- | `code/` digest | `5510f5995b022260` (sha256, first 16 hex digits) |
150
- | built | 2026-09-25T16:58:52+00:00 by tt-model 0.1.0 |
151
 
 
3
  - blackhole
4
  - p150x2
5
  - tt-model-cache
 
6
  - tt-model-container
7
  - vllm-plugin
8
  ---
9
 
10
  # qwen3.8-27b-p150x2
11
 
12
+ [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling package for the 4-chip P300x2: [qwen3.8-27b-p300x2](https://huggingface.co/changh95/qwen3.8-27b-p300x2) (profiles `plain`, `batch8-dflash2`, `single-user-dflash2`).
13
 
14
+ Runs on **p150x2** or **p150x2** or **p150x2** — see the serve profiles below.
15
 
16
  Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).
17
 
 
45
 
46
  | profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
47
  |---|---|---|---|---|---|---|
48
+ | `tp2` (default) | 32 | 64K | 622K tokens | 17.5 | 0.2 s / 1.8 s / 6.6 s | speed, long prompts, many users, all sampling options |
49
  | `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
50
+ | `tp2-dflash2` | 4 | 64K | 262K tokens | 41.9 | 0.15 s / 1.9 s / 6.8 s | 1-4 greedy users, code, prompts under ~16k tokens |
51
 
52
+ * `tp2` is 1.3x faster per user than `p1d1` at 1 user and has 3.8x the KV pool; `p1d1` keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where `tp2` at 8 users degrades to 134-255 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens `p1d1` delivers 84-90% of `tp2`'s aggregate throughput.
53
+ * `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.9x at 8k, 1.4x at 32k) and grows with answer length (3.1-3.2x at 1k-token answers).
54
 
55
  ### Performance
56
 
57
+ Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
58
 
59
  #### `tp2`: throughput, t/s (t/s/u)
60
 
61
  | ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
62
  |---|---|---|---|---|---|---|
63
+ | 128 / 128 | 17 (17.5) | 33 (16.3) | 58 (14.4) | 100 (12.5) | 162 (10.1) | 245 (7.7) |
64
+ | 1,024 / 128 | 17 (17.1) | 32 (15.8) | 54 (13.5) | 90 (11.2) | 138 (8.6) | 192 (6.0) |
65
+ | 2,048 / 128 | 17 (16.8) | 30 (15.2) | 52 (13.0) | 80 (10.0) | 118 (7.4) | 154 (4.8) |
66
+ | 4,096 / 128 | 16 (15.9) | 28 (13.8) | 45 (11.2) | 64 (8.0) | 84 (5.3) | 103 (3.2) |
67
+ | 8,192 / 128 | 14 (14.3) | 23 (11.6) | 35 (8.7) | 45 (5.6) | 54 (3.4) | 61 (1.9) |
68
+ | 16,384 / 128 | 12 (12.3) | 18 (9.0) | 24 (6.0) | 28 (3.6) | 32 (2.0) | 34 (1.1) |
69
+ | 32,768 / 128 | 9 (9.2) | 12 (6.0) | 14 (3.5) | 15 (1.9) | 16 (1.0) | 16 (0.9) @18 |
70
+ | 128 / 1,024 | 18 (17.7) | 34 (17.1) | 62 (15.6) | 116 (14.5) | 205 (12.8) | 368 (11.5) |
71
+ | 8,192 / 1,024 | 17 (17.1) | 32 (16.0) | 59 (14.8) | 97 (12.1) | 151 (9.4) | 219 (6.8) |
72
 
73
  `@18`: the KV pool seats 18 prompts of that length.
74
 
 
76
 
77
  | ISL / OSL | 1 user | 8 users | 32 users |
78
  |---|---|---|---|
79
+ | 128 / 128 | 197 / 56 | 1,436 / 70 | 6,272 / 82 |
80
+ | 2,048 / 128 | 501 / 56 | 3,600 / 73 | 15,447 / 88 |
81
+ | 8,192 / 128 | 1,778 / 57 | 12,352 / 81 | 34,181 / 260 |
82
+ | 32,768 / 128 | 6,631 / 58 | 34,573 / 255 | - |
83
+ | 128 / 1,024 | 197 / 56 | 1,444 / 68 | 6,273 / 81 |
84
 
85
  #### `p1d1`: throughput, t/s (t/s/u)
86
 
 
110
 
111
  | ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
112
  |---|---|---|---|---|
113
+ | 128 / 128 | 17.5 | 41.9 | 2.39x | 149 / 23 |
114
+ | 2,048 / 128 | 16.8 | 33.7 | 2.01x | 517 / 26 |
115
+ | 8,192 / 128 | 14.3 | 26.6 | 1.86x | 1,874 / 23 |
116
+ | 32,768 / 128 | 9.2 | 12.7 | 1.38x | 6,778 / 26 |
117
+ | 128 / 1,024 | 17.7 | 56.8 | 3.21x | 149 / 18 |
118
+ | 8,192 / 1,024 | 17.1 | 52.6 | 3.08x | 1,767 / 17 |
 
119
 
120
+ * 4 users: 113 vs 58 t/s aggregate at 128/128 (1.9x), 151 vs 62 at 128/1,024 (2.4x), 47 vs 35 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
121
 
122
  ### Notes
123
 
 
141
 
142
  | component | built from |
143
  | --- | --- |
144
+ | tt-metal | [`fee9e0d35948111be29083c4eb144023d86890d8`](https://github.com/tenstorrent/tt-metal/commit/fee9e0d35948111be29083c4eb144023d86890d8) |
145
  | vLLM | [`v0.26.0`](https://github.com/vllm-project/vllm/releases/tag/v0.26.0) |
146
  | vllm-tt-plugin | [`ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7`](https://github.com/changh95/vllm-tt-plugin/commit/ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7) |
147
+ | `code/` digest | `b0c865d431db3b43` (sha256, first 16 hex digits) |
148
+ | built | 2026-10-02T20:06:49+00:00 by tt-model 0.1.0 |
149
 
code/models/demos/blackhole/qwen36/tt-model-p300x2.yaml ADDED
@@ -0,0 +1,327 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # tt-model-p300x2.yaml — Qwen/Qwen3.8-27B on a P300x2 (2x P300 boards = 4 Blackhole chips, (1,4) mesh) via vLLM.
2
+ # ONE image, three serve profiles:
3
+ # plain plain batched decode, up to 32 users, full sampling (DEFAULT; was `batch32` of the old plain package)
4
+ # batch8-dflash2 DFlash2 speculative decoding, up to 8 users, two verify buckets, chunked prefill
5
+ # single-user-dflash2 DFlash2 speculative decoding, one user
6
+ # This file replaces both earlier P300x2 manifests: the plain package's models/demos/blackhole/qwen36/tt-model.yaml on
7
+ # branch qwen36-prefill-opt-package (image 0becf4834925, tt-metal 0e988de173f, plugin 751ec33) and
8
+ # tt-model-dflash2-p300x2.yaml (changh95/qwen3.8-27b-dflash2-p300x2, retired: tt-metal 8645cc34677, plugin 96d4416).
9
+ # tt-model.yaml on this branch is the p150x2 package (changh95/qwen3.8-27b-p150x2).
10
+ #
11
+ # tt-model package --container models/demos/blackhole/qwen36/tt-model-p300x2.yaml --out <builds>
12
+ # tt-model serve <builds>/qwen3.8-27b-p300x2/tt_kernel_manifest.json [--profile batch8-dflash2]
13
+ # tt-model push <builds>/qwen3.8-27b-p300x2 --private
14
+ schema: "5.1"
15
+ repo: changh95/qwen3.8-27b-p300x2
16
+ name: qwen3.8-27b-p300x2
17
+ weights:
18
+ repo: Qwen/Qwen3.8-27B
19
+ revision: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 # refs/main as validated, 2026-09-12
20
+ kind: vllm-plugin
21
+ arch: blackhole
22
+
23
+ source:
24
+ # Branch qwen36-p300x2-pkg = qwen36-p150x2 (round-5 integration a7a8aeb8d41: TP=2 prefill, tp2/TP=4 chunked prefill,
25
+ # fast GDN remap, long-prompt TTFT) + the SDPA width pin default-off commit + this manifest.
26
+ tt_metal: /home/ttuser/experiments/qwen36_27b/tt-metal-pkg5
27
+ code: # import closure of qwen36_vllm.py + the demo/spec-decode path
28
+ - models/common
29
+ - models/tt_transformers
30
+ - models/demos/blackhole/qwen36 # model, DFlash drafters, demo, tests, vllm_bundle
31
+ - models/demos/qwen3_vl/tt/common.py
32
+ - models/demos/utils
33
+ - models/experimental/gated_attention_gated_deltanet
34
+ - models/perf
35
+ - models/tt_dit/utils/tensor.py
36
+ - models/model_trace_region_sizes.yaml
37
+ - models/model_targets.yaml
38
+ ubuntu: "22.04"
39
+ python: "3.12"
40
+
41
+ runtime:
42
+ vllm: {version: "0.26.0"} # the validated serving venv: vllm 0.26.0 (empty target)
43
+ # vllm-tt-plugin branch pd-disagg: superset of both earlier pins (751ec33 of the plain package, 7ddf2c8 of the DFlash2
44
+ # v13 package): the block-output / batched / ragged contracts the DFlash2 profiles use, the TT chunk policy and the
45
+ # block-output chunked-prefill capability batch8-dflash2 uses, the generic TT_MODEL_CLASS_OVERRIDES the plain profile
46
+ # uses (a67ef80, already in 751ec33), plus PD connector, host sampler and scheduler fixes. Pinned to the pushed commit.
47
+ plugin: {repo: https://github.com/changh95/vllm-tt-plugin, ref: 96d4416d31f79973c9ff10b3b41e15d0b06e3430}
48
+ extra_models_dir: models/demos/blackhole/qwen36/vllm_bundle # main_class Qwen36DFlashForCausalLM (the DFlash2 profiles);
49
+ # `plain` selects Qwen36ForCausalLM via TT_MODEL_CLASS_OVERRIDES
50
+
51
+ serve: # common to every profile; a profile's env MERGES over this (it cannot unset a key), so
52
+ # nothing profile-specific lives here: no QWEN36_DRAFTER / DFLASH_* / QWEN36_MTP keys
53
+ hardware: p300x2
54
+ mesh_device: P150x4 # tt-metal's name for the (1,4) Blackhole mesh a P300x2 opens
55
+ port: 8000
56
+ max_model_len: 262144
57
+ max_num_seqs: 32 # overridden per profile below
58
+ block_size: 64
59
+ capabilities:
60
+ tool_parser: qwen3_coder
61
+ reasoning_parser: qwen3
62
+ additional_config:
63
+ tt:
64
+ fabric_config: FABRIC_1D
65
+ trace_region_size: 1073741824
66
+ l1_small_size: 24576
67
+ sample_on_device_mode: decode_only # plain: the ttnn device sampler (1x4); DFlash2: REQUIRED by the block-output contract
68
+ env:
69
+ ARCH_NAME: blackhole
70
+ TT_QWEN35_TEXT_VER: qwen36_blackhole
71
+ TT_MESH_GRAPH_DESC_PATH: /opt/tt-metal/tt_metal/fabric/mesh_graph_descriptors/p300_x2_mesh_graph_descriptor.textproto
72
+ QWEN36_MAX_TOKENS_ALL_USERS: "525312" # bf8 paged KV pool, every profile (= both earlier packages)
73
+ VLLM_RPC_TIMEOUT: "900000"
74
+ VLLM_CONFIGURE_LOGGING: "1"
75
+ TORCHDYNAMO_DISABLE: "1"
76
+ HF_HUB_OFFLINE: "1"
77
+ args:
78
+ - [--max-num-batched-tokens, "262144"]
79
+
80
+ serve_profiles:
81
+ - name: plain
82
+ description: Plain batched decode, up to 32 concurrent users, full sampling support (formerly `batch32`).
83
+ max_num_seqs: 32
84
+ env:
85
+ # The bundle's main_class is the DFlash class; the plugin registers the checkpoint arch under its TT-prefixed name
86
+ # (platform.py: "TT" + arch, the bundle too), so the override MUST be keyed on TTQwen3_5ForConditionalGeneration to
87
+ # beat it; the raw arch is kept for the plain lookup path (same string as the p150x2 tp2 profile, boot-verified there).
88
+ TT_MODEL_CLASS_OVERRIDES: "TTQwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM,Qwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM"
89
+ QWEN36_MTP: "0" # explicit: the plain class must not load the MTP drafter head / its paged KV cache
90
+ QWEN36_PLAIN_GDN_SLOT_FAST: "1" # fast bit-exact GDN slot hand-off at admission (code default: TP=2 only). Bit-exact at
91
+ # TP=4 (slot test 1x4, B=32, decode between admissions); ~130 ms less per admitted
92
+ # request. Set here, not as a TP=4 code default, so the DFlash2 profiles are unchanged
93
+ # (since 1bec7b836a8 QWEN36_DRAFTER=mtp alone would build it)
94
+ QWEN36_DRAFTER: mtp # carried over from batch32; the plain class never speculates. Kept as a backstop: if the
95
+ # override ever failed to apply, the DFlash class would serve plain (speculation OFF)
96
+ # instead of starting speculative serving at 32 slots
97
+ args:
98
+ - [--max-num-batched-tokens, "262144"]
99
+ - name: single-user-dflash2
100
+ description: One request at a time with lossless DFlash2 speculative decoding (block-output serving).
101
+ max_num_seqs: 1
102
+ env:
103
+ QWEN36_DRAFTER: dflash2
104
+ QWEN36_GDN_SPEC_FUSED: "1" # fused gdn_spec_step verify (M2): verify 49.8 -> 37.7 ms at 8 rows, 70 -> 39-41 at 32 rows
105
+ # DFlash2 speculative decoding (incoai/Qwen3.8-27B-DFlash2, block 8 -> 7 drafts per verify, lossless greedy). The
106
+ # drafter is a second checkpoint the single `weights` pointer cannot carry: it is resolved from the HF cache mounted
107
+ # at /hf, so run `hf download incoai/Qwen3.8-27B-DFlash2 --revision dedf8df68adfb1afeaf7b7480c0a0243108177b4` once
108
+ # before `tt-model serve` (quickstart below).
109
+ DFLASH_WEIGHTS: incoai/Qwen3.8-27B-DFlash2@dedf8df68adfb1afeaf7b7480c0a0243108177b4
110
+ QWEN36_DFLASH_BLOCK: "8"
111
+ QWEN36_DFLASH_MAX_PROMPT: "2048" # prompts up to this take the single masked-bucket prefill; longer ones the
112
+ # eager chunked spec prefill. EVERY text prompt speculates.
113
+ QWEN36_DFLASH_SERVE_BLOCK: "32" # tokens committed per vLLM step on a speculative request
114
+ args:
115
+ - [--max-num-batched-tokens, "262144"]
116
+ - --no-async-scheduling # the plugin's block-output contract is synchronous
117
+ - name: batch8-dflash2
118
+ description: Up to 8 concurrent requests with lossless DFlash2 speculative decoding in two verify buckets — K=7 drafts per step (4x8) while up to 4 requests are live, K=3 (8x4) at 5-8 — switched at runtime with no recapture (batched ragged block-output serving).
119
+ max_num_seqs: 8
120
+ env:
121
+ QWEN36_DRAFTER: dflash2
122
+ QWEN36_GDN_SPEC_FUSED: "1" # fused gdn_spec_step verify (M2): verify 49.8 -> 37.7 ms at 8 rows, 70 -> 39-41 at 32 rows
123
+ DFLASH_WEIGHTS: incoai/Qwen3.8-27B-DFlash2@dedf8df68adfb1afeaf7b7480c0a0243108177b4 # see single-user-dflash2
124
+ QWEN36_DFLASH_BLOCK: "8" # K_max = 7: the 4x8 bucket drafts the checkpoint's whole block; every bucket's K <= BLOCK-1
125
+ QWEN36_DFLASH_MAX_PROMPT: "2048"
126
+ QWEN36_DFLASH_SERVE_BLOCK: "32"
127
+ QWEN36_DFLASH_BUCKETS: "8x4,4x8" # verify geometries B x T (profiles/dual_bucket_spec.json): the decoder verifies in the
128
+ # smallest bucket that seats the live requests and reseeds them across a switch bit-exactly
129
+ QWEN36_DFLASH_RAGGED: "1" # one speculative iteration per decode step: a bucket switch stays inside one step
130
+ QWEN36_DFLASH_FOLD_SEED: "1" # (both verify traces + both buckets' draft/extend traces fit: TRACE region 137 MiB of 1 GiB)
131
+ QWEN36_DFLASH_CHUNKED_PREFILL: "1" # chunked prefill (qwen36_vllm_dflash.py declares supports_chunked_prefill +
132
+ # tt_prefill_chunk_tokens 2048 + tt_block_output_chunked_prefill; refused at
133
+ # TP != 4 and max_num_seqs 1). Evaluated OFF vs ON: profiles/tp4_chunked/EVAL.md
134
+ additional_config:
135
+ tt: # deep-merged over serve.additional_config.tt
136
+ prefill_chunk_tokens: 2048 # chunk size (a multiple of the model's tt_prefill_chunk_tokens and of block_size)
137
+ chunked_prefill_decode_steps: 4 # speculative decode steps between two chunk steps while requests decode
138
+ # Other chunk-policy keys keep their block-output defaults (vllm_tt_plugin/config.py
139
+ # resolve_tt_prefill_chunk_policy): only a prompt with >= 8192 tokens left (chunked_prefill_min_tokens) that
140
+ # arrives while an EARLIER request decodes (chunked_prefill_protect_prior_decoders) is chunked; bursts
141
+ # (burst_longs 2) and idle steps prefill whole, exactly as without chunking.
142
+ args:
143
+ - [--max-num-batched-tokens, "262144"]
144
+ - --no-async-scheduling # the block-output contract and the chunk cadence are synchronous
145
+ default_profile: plain
146
+
147
+ verify:
148
+ # plain: the class the override names imports, and the plugin parses the profile's override string to it (vllm first:
149
+ # vllm_tt_plugin.platform imported cold hits a vllm.platforms circular import)
150
+ - "from models.demos.blackhole.qwen36.tt.qwen36_vllm import Qwen36ForCausalLM; assert Qwen36ForCausalLM"
151
+ - "import os, vllm; os.environ['TT_MODEL_CLASS_OVERRIDES']='TTQwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM,Qwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM'; from vllm_tt_plugin.platform import _tt_model_class_overrides as f; o=f(); assert o['TTQwen3_5ForConditionalGeneration'].endswith('qwen36_vllm:Qwen36ForCausalLM') and len(o) == 2, o"
152
+ - "import os; os.environ['QWEN36_MTP']='0'; os.environ['QWEN36_DRAFTER']='mtp'; from models.demos.blackhole.qwen36.tt.model_config import mtp_requested_by_env; assert not mtp_requested_by_env()"
153
+ - "import os; os.environ['QWEN36_DRAFTER']='mtp'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; assert not C.model_capabilities['tt_adaptive_block_output'] and C.model_capabilities['output_tokens_per_step'] == 1, dict(C.model_capabilities)"
154
+ # DFlash2 profiles
155
+ - "import os; os.environ['QWEN36_DRAFTER']='dflash2'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; assert C.model_capabilities['tt_adaptive_block_output'] and C.model_capabilities['tt_adaptive_block_batched'] and C.model_capabilities['output_tokens_per_step'] == 32, C.model_capabilities"
156
+ - "from models.demos.blackhole.qwen36.tt.dflash2_serving import DFlash2ServingDecoder; assert DFlash2ServingDecoder"
157
+ - "from models.demos.blackhole.qwen36.tt.dflash2_serving import DFlash2DualBucketDecoder, BucketPlanner; assert DFlash2DualBucketDecoder and BucketPlanner"
158
+ - "import os; os.environ['QWEN36_DRAFTER']='dflash2'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import parse_buckets, check_buckets; assert check_buckets(parse_buckets('8x4,4x8'), 8, 7, True) == ((8, 4), (4, 8))"
159
+ - "import json; m=json.load(open('/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle/qwen36/vllm_metadata.json')); assert m['main_class'].endswith(':Qwen36DFlashForCausalLM'), m"
160
+ - "from vllm_tt_plugin.config import get_tt_block_output_kv_lookahead_tokens, is_tt_adaptive_block_output_model, is_tt_adaptive_block_batched, is_tt_adaptive_block_ragged"
161
+ - "import os; os.environ['QWEN36_DRAFTER']='dflash2'; os.environ['QWEN36_DFLASH_RAGGED']='1'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; assert C.model_capabilities['tt_adaptive_block_batched'] and C.model_capabilities['tt_adaptive_block_ragged']"
162
+ - "from models.demos.blackhole.qwen36.tt.dflash2_decode import DFlash2Decoder; from models.demos.blackhole.qwen36.tt.dflash2_tp import DFlash2DrafterTP; assert DFlash2Decoder and DFlash2DrafterTP"
163
+ # batch8-dflash2 chunked prefill: the model declares the block-output chunk contract only with the knob on (never with it
164
+ # off: single-user-dflash2), and the plugin carries the chunk policy that reads prefill_chunk_tokens / chunked_prefill_decode_steps.
165
+ - "import os; os.environ['QWEN36_DRAFTER']='dflash2'; os.environ['QWEN36_DFLASH_CHUNKED_PREFILL']='1'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; c=C.model_capabilities; assert c['supports_chunked_prefill'] and c['tt_block_output_chunked_prefill'] and c['tt_prefill_chunk_tokens'] == 2048, dict(c)"
166
+ - "import os; os.environ['QWEN36_DRAFTER']='dflash2'; os.environ.pop('QWEN36_DFLASH_CHUNKED_PREFILL', None); from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; c=C.model_capabilities; assert not c.get('supports_chunked_prefill') and 'tt_block_output_chunked_prefill' not in c, dict(c)"
167
+ - "from vllm_tt_plugin.config import resolve_tt_prefill_chunk_policy, is_tt_block_output_chunked_prefill, get_tt_prefill_chunk_extras"
168
+ # common
169
+ - "import vllm.model_executor.models.qwen3_5"
170
+ - "import transformers.models.qwen3_5"
171
+ - "from pathlib import Path; assert Path('/opt/tt-metal/tt_metal/fabric/mesh_graph_descriptors/p300_x2_mesh_graph_descriptor.textproto').is_file()"
172
+ - "from pathlib import Path; assert Path('/opt/tt-metal/models/model_trace_region_sizes.yaml').is_file()"
173
+ - "from pathlib import Path; assert Path('/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle/qwen36/vllm_metadata.json').is_file()"
174
+
175
+ # Card: rendered by profiles/render_card.py p300x2 <this file> from profiles/card_p300x2_bench.md. Every {{...}} is a
176
+ # placeholder for a number measured on this build (lane B2 benchmarks); fill the source, then re-render.
177
+ card:
178
+ description: |
179
+ [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on a Tenstorrent P300x2 (2× P300 = 4 Blackhole chips, 4-way tensor parallel), 256K context. One image, three profiles: `plain` (default; plain batched decode for up to 32 users with every sampling option), `batch8-dflash2` (lossless [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding for 1 to 8 users, 1.2-1.8× `plain` per user) and `single-user-dflash2` (the fastest single stream: 60 tokens/s at short prompts and 83 with 1k-token answers, vs 32 for `plain`). It replaces [changh95/qwen3.8-27b-dflash2-p300x2](https://huggingface.co/changh95/qwen3.8-27b-dflash2-p300x2), whose profiles now live here.
180
+ quickstart: |
181
+ ### Run
182
+
183
+ ```bash
184
+ tt-model pull changh95/qwen3.8-27b-p300x2 --with-weights
185
+ tt-model serve changh95/qwen3.8-27b-p300x2 # plain (default)
186
+ tt-model serve changh95/qwen3.8-27b-p300x2 --profile batch8-dflash2 # needs the drafter weights below
187
+ tt-model serve changh95/qwen3.8-27b-p300x2 --profile single-user-dflash2 # needs the drafter weights below
188
+ tt-model profiles changh95/qwen3.8-27b-p300x2 # list the profiles
189
+ ```
190
+
191
+ * `batch32` is now called `plain`: same configuration, new name. `--profile batch32` no longer exists; use `--profile plain` or no flag.
192
+ * First boot converts the weights and compiles the kernels (about 7 min); later boots take about 2-3 min.
193
+
194
+ ### Drafter weights (DFlash2 profiles only, once)
195
+
196
+ ```bash
197
+ hf download incoai/Qwen3.8-27B-DFlash2 --revision dedf8df68adfb1afeaf7b7480c0a0243108177b4
198
+ ```
199
+
200
+ The drafter is a second checkpoint; the server reads it from your Hugging Face cache. `plain` does not need it.
201
+
202
+ ### Query
203
+
204
+ OpenAI-compatible API on port 20000, model id `Qwen/Qwen3.8-27B`:
205
+
206
+ ```bash
207
+ curl -s http://localhost:20000/v1/chat/completions -H 'Content-Type: application/json' \
208
+ -d '{"model": "Qwen/Qwen3.8-27B", "messages": [{"role": "user", "content": "Write a Python function that merges two sorted lists."}], "max_tokens": 512}'
209
+ ```
210
+
211
+ * Streaming, tool calls (`qwen3_coder` parser) and reasoning content (`qwen3` parser) work as in vLLM on every profile; `"chat_template_kwargs": {"enable_thinking": false}` turns thinking off.
212
+ * `plain` supports every vLLM sampling option. The DFlash2 profiles return the target model's greedy trajectory: `temperature`/`top_p` affect only the first token; `logprobs`, structured outputs and images are not supported.
213
+
214
+ ### Profiles at a glance
215
+
216
+ | profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
217
+ |---|---|---|---|---|---|---|
218
+ | `plain` (default) | 32 | 256K | 525K tokens | 31.6 | 0.2 s / 1.1 s / 4.0 s | many users, long prompts, all sampling options (temperature, top-p/top-k, penalties, logprobs, structured outputs) |
219
+ | `batch8-dflash2` | 8 | 256K | 525K tokens | 57.3 | 0.1 s / 1.1 s / 4.1 s | 1-8 greedy users |
220
+ | `single-user-dflash2` | 1 | 256K | 525K tokens | 60.1 | 0.1 s / 1.1 s / 4.1 s | exactly one greedy user, fastest per token |
221
+
222
+ * `plain` is the old `batch32` profile under a new name. At 1 user `batch8-dflash2` is 1.8x faster per user; at 8 users it still leads in aggregate throughput (271 vs 163 t/s at 128 / 128). Beyond 8 users, or for sampled output, use `plain` (up to 688 t/s at 32 users).
223
+ * The DFlash2 profiles are lossless: each request follows the target model's greedy trajectory (up to bf16 near-ties). Their gain is largest with long answers (1.6x at 1,024-token answers, 1 user) and fades with prompt length (1.19x at 32k tokens).
224
+
225
+ ## Benchmarks
226
+
227
+ tt-inference-server `--workflow benchmarks` (standard P300X2 grid: random prompts of exactly ISL tokens, `ignore_eos`, greedy, streaming) on one P300x2, measured 2026-10-01/02 on this image, one profile at a time. `t/s` = aggregate output tokens per second, `(t/s/u)` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Points with several runs (the 1-user rows repeat in every run; `single-user-dflash2` ran three times) show the median.
228
+
229
+ ### `plain` (default), t/s (t/s/u)
230
+
231
+ | ISL / OSL | 1 user | 2 users | 4 users | 8 users | 16 users | 32 users |
232
+ |---|---|---|---|---|---|---|
233
+ | 128 / 128 | 31.6 (31.6) | 50.9 (25.4) | 94.0 (23.5) | 163.4 (20.4) | 256.8 (16.1) | 359.0 (11.2) |
234
+ | 1,024 / 128 | 30.8 (30.8) | 48.9 (24.5) | 88.4 (22.1) | 146.1 (18.3) | 218.1 (13.6) | 283.9 (8.9) |
235
+ | 2,048 / 128 | 30.0 (30.0) | 47.2 (23.6) | 82.2 (20.6) | 131.4 (16.4) | 186.9 (11.7) | 235.1 (7.3) |
236
+ | 4,096 / 128 | 28.3 (28.3) | 43.3 (21.6) | 71.0 (17.8) | 104.4 (13.0) | 137.4 (8.6) | 160.8 (5.0) |
237
+ | 8,192 / 128 | 25.4 (25.4) | 36.8 (18.4) | 55.9 (14.0) | 73.5 (9.2) | 88.9 (5.6) | 98.3 (3.1) |
238
+ | 16,384 / 128 | 21.5 (21.5) | 28.9 (14.4) | 38.5 (9.6) | 47.0 (5.9) | 52.6 (3.3) | – |
239
+ | 32,768 / 128 | 16.0 (16.0) | 19.7 (9.8) | 23.4 (5.9) | 25.7 (3.2) | – | – |
240
+ | 65,536 / 128 | 9.8 (9.8) | 11.2 (5.6) | 12.3 (3.1) | 12.5 (1.6) | – | – |
241
+ | 131,072 / 128 | 4.8 (4.8) | 5.3 (2.6) | 5.5 (1.4) | – | – | – |
242
+ | 128 / 1,024 | 32.4 (32.4) | 59.7 (29.9) | 117.0 (29.3) | 224.1 (28.0) | 415.8 (26.0) | 687.7 (21.5) |
243
+ | 8,192 / 1,024 | 31.1 (31.1) | 56.0 (28.0) | 103.5 (25.9) | 179.7 (22.5) | 286.8 (17.9) | 393.8 (12.3) |
244
+ | 10,000 / 1,024 | 30.9 (30.9) | 54.9 (27.4) | 100.1 (25.0) | 170.6 (21.3) | 266.1 (16.6) | 356.1 (11.1) |
245
+
246
+ * A dash = the 525K-token KV pool does not seat that many prompts of that length (the workflow skips the point).
247
+
248
+ ### `plain`, TTFT ms / TPOT ms
249
+
250
+ | ISL / OSL | 1 user | 8 users | 32 users |
251
+ |---|---|---|---|
252
+ | 128 / 128 | 157 / 30.7 | 1,183 / 40.0 | 5,019 / 50.3 |
253
+ | 1,024 / 128 | 245 / 30.8 | 1,810 / 40.9 | 7,920 / 51.2 |
254
+ | 2,048 / 128 | 345 / 30.9 | 2,495 / 41.7 | 10,772 / 52.3 |
255
+ | 4,096 / 128 | 587 / 30.9 | 4,216 / 44.1 | 18,361 / 55.9 |
256
+ | 8,192 / 128 | 1,103 / 31.0 | 7,776 / 48.4 | 33,750 / 62.3 |
257
+ | 16,384 / 128 | 1,982 / 31.3 | 14,673 / 56.2 | – |
258
+ | 32,768 / 128 | 4,003 / 31.6 | 30,395 / 74.2 | – |
259
+ | 65,536 / 128 | 9,024 / 32.3 | 53,087 / 224.8 | – |
260
+ | 131,072 / 128 | 22,314 / 33.9 | – | – |
261
+ | 128 / 1,024 | 158 / 30.7 | 1,202 / 34.6 | 5,169 / 41.5 |
262
+ | 8,192 / 1,024 | 1,113 / 31.1 | 7,753 / 37.0 | 33,766 / 48.3 |
263
+ | 10,000 / 1,024 | 1,255 / 31.1 | 9,347 / 37.8 | 36,777 / 54.0 |
264
+
265
+ ### `batch8-dflash2`, t/s (t/s/u)
266
+
267
+ | ISL / OSL | 1 user | 2 users | 4 users | 8 users | `plain` 1 user t/s/u |
268
+ |---|---|---|---|---|---|
269
+ | 128 / 128 | 57.3 (57.3) | 100.6 (51.4) | 173.9 (45.8) | 270.5 (34.9) | 31.6 |
270
+ | 1,024 / 128 | 55.7 (55.7) | 101.2 (51.0) | 162.8 (43.5) | 245.6 (31.5) | 30.8 |
271
+ | 2,048 / 128 | 48.7 (48.7) | 87.1 (45.0) | 140.0 (36.0) | 195.1 (24.9) | 30.0 |
272
+ | 4,096 / 128 | 45.0 (45.0) | 73.2 (38.1) | 108.9 (28.1) | 141.9 (18.0) | 28.3 |
273
+ | 8,192 / 128 | 38.4 (38.4) | 54.2 (27.4) | 69.9 (18.7) | 85.1 (10.8) | 25.4 |
274
+ | 16,384 / 128 | 28.2 (28.2) | 37.8 (19.2) | 43.4 (11.4) | 51.9 (6.6) | 21.5 |
275
+ | 32,768 / 128 | 19.0 (19.0) | 23.4 (11.8) | 22.7 (6.2) | 25.5 (3.4) | 16.0 |
276
+ | 65,536 / 128 | 10.7 (10.7) | 12.1 (6.1) | 11.7 (3.1) | 12.3 (1.6) | 9.8 |
277
+ | 131,072 / 128 | 4.9 (4.9) | 5.4 (2.7) | 5.5 (1.4) | – | 4.8 |
278
+ | 128 / 1,024 | 53.2 (53.2) | 88.1 (44.6) | 162.2 (45.2) | 279.0 (39.5) | 32.4 |
279
+ | 8,192 / 1,024 | 43.2 (43.2) | 62.2 (32.7) | 121.1 (35.3) | 219.4 (30.9) | 31.1 |
280
+ | 10,000 / 1,024 | 52.7 (52.7) | 64.8 (36.8) | 153.6 (40.4) | 191.5 (29.4) | 30.9 |
281
+
282
+ * A dash = the KV pool does not seat that many users at that length. At 1 to 4 users the server drafts 7 tokens per step; at 5 to 8, 3 per step. Long prompts (8,192 tokens or more) that arrive while other users decode are prefilled in 2,048-token chunks, which keeps the other users' tokens flowing and raises that prompt's own TTFT.
283
+
284
+ ### `batch8-dflash2`, TTFT ms / TPOT ms
285
+
286
+ | ISL / OSL | 1 user | 8 users |
287
+ |---|---|---|
288
+ | 128 / 128 | 138 / 16.5 | 331 / 26.3 |
289
+ | 1,024 / 128 | 190 / 16.6 | 565 / 27.5 |
290
+ | 2,048 / 128 | 288 / 18.4 | 890 / 33.5 |
291
+ | 4,096 / 128 | 550 / 18.0 | 1,840 / 41.6 |
292
+ | 8,192 / 128 | 1,079 / 17.7 | 5,603 / 49.3 |
293
+ | 16,384 / 128 | 2,003 / 20.0 | 10,244 / 73.1 |
294
+ | 32,768 / 128 | 4,085 / 20.9 | 30,933 / 55.6 |
295
+ | 65,536 / 128 | 9,208 / 22.2 | 54,090 / 206.9 |
296
+ | 131,072 / 128 | 22,733 / 27.3 | – |
297
+ | 128 / 1,024 | 139 / 18.7 | 376 / 25.0 |
298
+ | 8,192 / 1,024 | 1,080 / 22.1 | 4,411 / 28.0 |
299
+ | 10,000 / 1,024 | 1,231 / 17.8 | 5,891 / 28.3 |
300
+
301
+ ### `single-user-dflash2`, one user
302
+
303
+ | ISL / OSL | TTFT ms | TPOT ms | t/s/u | `plain` t/s/u |
304
+ |---|---|---|---|---|
305
+ | 128 / 128 | 134 | 15.7 | 60.1 | 31.6 |
306
+ | 1,024 / 128 | 184 | 16.1 | 57.6 | 30.8 |
307
+ | 2,048 / 128 | 289 | 17.4 | 51.3 | 30.0 |
308
+ | 4,096 / 128 | 537 | 16.9 | 47.5 | 28.3 |
309
+ | 8,192 / 128 | 1,067 | 16.1 | 41.1 | 25.4 |
310
+ | 16,384 / 128 | 1,991 | 17.4 | 30.5 | 21.5 |
311
+ | 32,768 / 128 | 4,086 | 17.2 | 20.4 | 16.0 |
312
+ | 65,536 / 128 | 9,187 | 18.8 | 11.1 | 9.8 |
313
+ | 131,072 / 128 | 22,620 | 18.5 | 5.1 | 4.8 |
314
+ | 128 / 1,024 | 134 | 12.5 | 79.2 | 32.4 |
315
+ | 8,192 / 1,024 | 1,082 | 11.0 | 83.2 | 31.1 |
316
+ | 10,000 / 1,024 | 1,221 | 13.1 | 70.0 | 30.9 |
317
+
318
+ * Median of three runs per point.
319
+
320
+ ### Changes in this release
321
+
322
+ * One package for the P300x2. The DFlash2 profiles moved here from [changh95/qwen3.8-27b-dflash2-p300x2](https://huggingface.co/changh95/qwen3.8-27b-dflash2-p300x2), which is retired (its page points here). The old plain profile `batch32` is now `plain`.
323
+ * `plain` now includes two correctness fixes. Earlier plain images could produce degenerate greedy text on short prompts, and their output could change from the second request on. The traced short-prompt prefill no longer trusts stale page-table rows, and each request's Gated-DeltaNet state is written into its decode slot exactly.
324
+ * `batch8-dflash2` prefills long prompts in chunks. A prompt of 8,192 tokens or more that arrives while other requests decode is prefilled in 2,048-token chunks between their steps, so they keep streaming (first shipped in the last DFlash2 release, 2026-09-30).
325
+ * Faster prefill of long prompts: TTFT at 32k tokens 5,149 -> 4,085 ms on `batch8-dflash2`, 5,077 -> 4,086 ms on `single-user-dflash2` and 4,457 -> 4,003 ms on `plain` (1 user).
326
+ * `plain` throughput: the exact slot write costs time per admitted request, so short prompts at many users are slower than with the previous plain image: 128 / 128 tokens at 32 users 491 -> 359 tokens/s (-27%), at 8 users 188 -> 163 (-13%), TTFT at 128 tokens 94 -> 157 ms. One user decodes at the same speed (32.3 -> 31.6 t/s/u); long answers at 32 users lose 7% (739 -> 688 t/s at 128 / 1,024). The previous image's higher numbers came from a slot write that was not exact.
327
+ * Intended output change: on the DFlash2 profiles, prompts of 2,048 tokens or more now prefill with exactly the numerics of the `plain` path. Greedy output for those prompts can differ from the previous DFlash2 package at bf16 near-ties. Shorter prompts give the same output as before.
code/models/demos/blackhole/qwen36/tt-model.yaml CHANGED
@@ -194,7 +194,7 @@ verify:
194
  # benchmarks land; the text below is the pre-benchmark draft (placeholders in {{...}}).
195
  card:
196
  description: |
197
- [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling packages for the 4-chip P300x2: [qwen3.8-27b-p300x2](https://huggingface.co/changh95/qwen3.8-27b-p300x2), [qwen3.8-27b-dflash2-p300x2](https://huggingface.co/changh95/qwen3.8-27b-dflash2-p300x2).
198
  quickstart: |
199
  ### Run
200
 
@@ -217,30 +217,30 @@ card:
217
 
218
  | profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
219
  |---|---|---|---|---|---|---|
220
- | `tp2` (default) | 32 | 64K | 622K tokens | 17.1 | 0.3 s / 1.9 s / 8.2 s | speed, long prompts, many users, all sampling options |
221
  | `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
222
- | `tp2-dflash2` | 4 | 64K | 262K tokens | 40.8 (58 on code) | 0.15 s / 2.1 s / 9.2 s | 1-4 greedy users, code, prompts under ~16k tokens |
223
 
224
- * `tp2` is 1.2x faster per user than `p1d1` at 1 user and has 3.8x the KV pool; `p1d1` keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where `tp2` at 8 users degrades to 146-285 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens `p1d1` is within 4-9% of `tp2`'s aggregate throughput.
225
- * `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.6x at 8k, 1.2x at 32k) and grows with answer length (3.3-3.4x at 1k-token answers).
226
 
227
  ### Performance
228
 
229
- Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Code: SPEED-Bench coding, 80 prompts, thinking off, temperature 0, 1 user. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
230
 
231
  #### `tp2`: throughput, t/s (t/s/u)
232
 
233
  | ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
234
  |---|---|---|---|---|---|---|
235
- | 128 / 128 | 17 (17.1) | 29 (14.7) | 54 (13.4) | 92 (11.4) | 143 (9.0) | 193 (6.0) |
236
- | 1,024 / 128 | 17 (16.7) | 28 (14.1) | 50 (12.6) | 83 (10.4) | 123 (7.7) | 156 (4.9) |
237
- | 2,048 / 128 | 16 (16.3) | 27 (13.7) | 48 (11.9) | 75 (9.3) | 107 (6.7) | 131 (4.1) |
238
- | 4,096 / 128 | 15 (15.4) | 25 (12.4) | 41 (10.2) | 59 (7.4) | 78 (4.9) | 90 (2.8) |
239
- | 8,192 / 128 | 14 (13.9) | 21 (10.5) | 31 (7.8) | 41 (5.1) | 50 (3.1) | 54 (1.7) |
240
- | 16,384 / 128 | 11 (11.4) | 16 (7.9) | 21 (5.2) | 25 (3.1) | 28 (1.8) | 29 (0.9) |
241
- | 32,768 / 128 | 8 (8.2) | 10 (5.0) | 12 (3.0) | 13 (1.7) | 14 (0.9) | 14 (0.8) @18 |
242
- | 128 / 1,024 | 18 (17.6) | 33 (16.6) | 65 (16.1) | 123 (15.4) | 218 (13.6) | 361 (11.3) |
243
- | 8,192 / 1,024 | 17 (16.9) | 31 (15.6) | 57 (14.3) | 98 (12.2) | 153 (9.6) | 210 (6.6) |
244
 
245
  `@18`: the KV pool seats 18 prompts of that length.
246
 
@@ -248,11 +248,11 @@ card:
248
 
249
  | ISL / OSL | 1 user | 8 users | 32 users |
250
  |---|---|---|---|
251
- | 128 / 128 | 316 / 57 | 2,394 / 69 | 10,226 / 87 |
252
- | 2,048 / 128 | 610 / 57 | 4,518 / 72 | 19,498 / 93 |
253
- | 8,192 / 128 | 1,941 / 57 | 14,217 / 84 | 39,474 / 285 |
254
- | 32,768 / 128 | 8,243 / 58 | 40,999 / 285 | - |
255
- | 128 / 1,024 | 317 / 57 | 2,379 / 63 | 10,322 / 79 |
256
 
257
  #### `p1d1`: throughput, t/s (t/s/u)
258
 
@@ -282,15 +282,14 @@ card:
282
 
283
  | ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
284
  |---|---|---|---|---|
285
- | 128 / 128 | 17.1 | 40.8 | 2.39x | 152 / 24 |
286
- | 2,048 / 128 | 16.3 | 33.6 | 2.06x | 511 / 26 |
287
- | 8,192 / 128 | 13.9 | 22.6 | 1.63x | 2,060 / 28 |
288
- | 32,768 / 128 | 8.2 | 9.6 | 1.17x | 9,224 / 32 |
289
- | 128 / 1,024 | 17.6 | 58.8 | 3.34x | 154 / 17 |
290
- | 8,192 / 1,024 | 16.9 | 56.7 | 3.36x | 2,060 / 16 |
291
- | code (SPEED-Bench, EOS-terminated) | 16.2 | 58.4 | 3.6x | 194 / 13 |
292
 
293
- * 4 users: 117 vs 54 t/s aggregate at 128/128 (2.2x), 163 vs 65 at 128/1,024 (2.5x), 42 vs 31 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
294
 
295
  ### Notes
296
 
 
194
  # benchmarks land; the text below is the pre-benchmark draft (placeholders in {{...}}).
195
  card:
196
  description: |
197
+ [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling package for the 4-chip P300x2: [qwen3.8-27b-p300x2](https://huggingface.co/changh95/qwen3.8-27b-p300x2) (profiles `plain`, `batch8-dflash2`, `single-user-dflash2`).
198
  quickstart: |
199
  ### Run
200
 
 
217
 
218
  | profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
219
  |---|---|---|---|---|---|---|
220
+ | `tp2` (default) | 32 | 64K | 622K tokens | 17.5 | 0.2 s / 1.8 s / 6.6 s | speed, long prompts, many users, all sampling options |
221
  | `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
222
+ | `tp2-dflash2` | 4 | 64K | 262K tokens | 41.9 | 0.15 s / 1.9 s / 6.8 s | 1-4 greedy users, code, prompts under ~16k tokens |
223
 
224
+ * `tp2` is 1.3x faster per user than `p1d1` at 1 user and has 3.8x the KV pool; `p1d1` keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where `tp2` at 8 users degrades to 134-255 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens `p1d1` delivers 84-90% of `tp2`'s aggregate throughput.
225
+ * `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.9x at 8k, 1.4x at 32k) and grows with answer length (3.1-3.2x at 1k-token answers).
226
 
227
  ### Performance
228
 
229
+ Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
230
 
231
  #### `tp2`: throughput, t/s (t/s/u)
232
 
233
  | ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
234
  |---|---|---|---|---|---|---|
235
+ | 128 / 128 | 17 (17.5) | 33 (16.3) | 58 (14.4) | 100 (12.5) | 162 (10.1) | 245 (7.7) |
236
+ | 1,024 / 128 | 17 (17.1) | 32 (15.8) | 54 (13.5) | 90 (11.2) | 138 (8.6) | 192 (6.0) |
237
+ | 2,048 / 128 | 17 (16.8) | 30 (15.2) | 52 (13.0) | 80 (10.0) | 118 (7.4) | 154 (4.8) |
238
+ | 4,096 / 128 | 16 (15.9) | 28 (13.8) | 45 (11.2) | 64 (8.0) | 84 (5.3) | 103 (3.2) |
239
+ | 8,192 / 128 | 14 (14.3) | 23 (11.6) | 35 (8.7) | 45 (5.6) | 54 (3.4) | 61 (1.9) |
240
+ | 16,384 / 128 | 12 (12.3) | 18 (9.0) | 24 (6.0) | 28 (3.6) | 32 (2.0) | 34 (1.1) |
241
+ | 32,768 / 128 | 9 (9.2) | 12 (6.0) | 14 (3.5) | 15 (1.9) | 16 (1.0) | 16 (0.9) @18 |
242
+ | 128 / 1,024 | 18 (17.7) | 34 (17.1) | 62 (15.6) | 116 (14.5) | 205 (12.8) | 368 (11.5) |
243
+ | 8,192 / 1,024 | 17 (17.1) | 32 (16.0) | 59 (14.8) | 97 (12.1) | 151 (9.4) | 219 (6.8) |
244
 
245
  `@18`: the KV pool seats 18 prompts of that length.
246
 
 
248
 
249
  | ISL / OSL | 1 user | 8 users | 32 users |
250
  |---|---|---|---|
251
+ | 128 / 128 | 197 / 56 | 1,436 / 70 | 6,272 / 82 |
252
+ | 2,048 / 128 | 501 / 56 | 3,600 / 73 | 15,447 / 88 |
253
+ | 8,192 / 128 | 1,778 / 57 | 12,352 / 81 | 34,181 / 260 |
254
+ | 32,768 / 128 | 6,631 / 58 | 34,573 / 255 | - |
255
+ | 128 / 1,024 | 197 / 56 | 1,444 / 68 | 6,273 / 81 |
256
 
257
  #### `p1d1`: throughput, t/s (t/s/u)
258
 
 
282
 
283
  | ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
284
  |---|---|---|---|---|
285
+ | 128 / 128 | 17.5 | 41.9 | 2.39x | 149 / 23 |
286
+ | 2,048 / 128 | 16.8 | 33.7 | 2.01x | 517 / 26 |
287
+ | 8,192 / 128 | 14.3 | 26.6 | 1.86x | 1,874 / 23 |
288
+ | 32,768 / 128 | 9.2 | 12.7 | 1.38x | 6,778 / 26 |
289
+ | 128 / 1,024 | 17.7 | 56.8 | 3.21x | 149 / 18 |
290
+ | 8,192 / 1,024 | 17.1 | 52.6 | 3.08x | 1,767 / 17 |
 
291
 
292
+ * 4 users: 113 vs 58 t/s aggregate at 128/128 (1.9x), 151 vs 62 at 128/1,024 (2.4x), 47 vs 35 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
293
 
294
  ### Notes
295
 
code/models/demos/blackhole/qwen36/tt/chunked_prefill.py ADDED
@@ -0,0 +1,114 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SPDX-FileCopyrightText: © 2026 Tenstorrent USA, Inc.
2
+ # SPDX-License-Identifier: Apache-2.0
3
+ """Scheduler-driven chunked prefill for batched TP serving: the host-side row plan (no ttnn, host-testable).
4
+
5
+ vLLM (with the plugin's chunk policy, QWEN36_CHUNKED_PREFILL=1) may split one long prompt into 2048-aligned chunks that
6
+ run in separate prefill calls, with decode steps and other prompts' prefills in between. The per-sequence GDN state of
7
+ the partial prompt (recurrent state, conv taps, cross-chunk conv carry) lives in the persistent B=1 prefill scratch
8
+ between its calls (its decode slot row is NOT safe storage: the batched decode rewrites idle rows inside the pow2 bucket
9
+ and slot remaps gather every row). Any other prompt prefilled while a partial is in flight resets that scratch, so the
10
+ partial's state is first copied to a same-shape park buffer ("park") and copied back before its next chunk ("unpark").
11
+
12
+ ``ChunkedPrefillPlanner.plan`` turns one prefill call's rows into an execution plan and the scratch owner that results:
13
+
14
+ * a row is a RESUME row only when the caller flags it (``resume_mask``; the plugin sets it for a scheduler chunk
15
+ continuation). A resume row must continue the current owner exactly: same first KV block, ``start == next_pos``,
16
+ ``start`` a positive multiple of the chunk size. Anything else raises (never silently re-prefill with a wrong state).
17
+ * every other row re-prefills from position 0 (``start`` ignored), which is the pre-chunking behaviour for any row.
18
+ * an intermediate row (``final_mask`` False) must end on a chunk boundary; it becomes the scratch owner, writes no
19
+ decode slot and reads back no logits.
20
+ * the resume row runs first. Before any row that resets the scratch (every non-resume row) an unparked owner is parked;
21
+ before a resume row a parked owner is unparked. So a call that carries only short prompts between two chunks of the
22
+ partial parks it too (review B1), and a stale owner left by an aborted/preempted partial costs one park at most.
23
+ """
24
+
25
+ from dataclasses import dataclass, replace
26
+
27
+
28
+ @dataclass(frozen=True)
29
+ class ScratchOwner:
30
+ """The partial prompt whose GDN state the B=1 prefill scratch (or its park buffer) holds."""
31
+
32
+ first_block: int # page_table[row, 0] of the request: its identity for the continuation check
33
+ next_pos: int # tokens [0, next_pos) are in the paged KV and in the state; the next chunk must start here
34
+ parked: bool # True: the state is in the park buffer (the scratch was reused since)
35
+
36
+
37
+ @dataclass(frozen=True)
38
+ class RowPlan:
39
+ row: int # index into the call's rows (logits are returned in call order)
40
+ start: int # absolute first position this call computes (0 unless resume)
41
+ end: int # exclusive end (the chunk end / prompt length)
42
+ resume: bool
43
+ final: bool # False = intermediate chunk: no slot write, no logits readback
44
+ park_before: bool
45
+ unpark_before: bool
46
+
47
+
48
+ class ChunkedPrefillPlanner:
49
+ def __init__(self, chunk_tokens):
50
+ if int(chunk_tokens) <= 0:
51
+ raise ValueError(f"chunk_tokens must be positive, got {chunk_tokens}")
52
+ self.chunk = int(chunk_tokens)
53
+ self.owner = None # ScratchOwner | None
54
+
55
+ def plan(self, starts, ends, resume_mask, final_mask, first_blocks):
56
+ """Return (list[RowPlan] in execution order, owner after the call). Does not modify ``self.owner``."""
57
+ n = len(ends)
58
+ for name, seq in (("starts", starts), ("resume_mask", resume_mask), ("final_mask", final_mask)):
59
+ if len(seq) != n:
60
+ raise ValueError(f"chunked prefill: {name} has {len(seq)} entries for {n} rows")
61
+ if len(first_blocks) != n:
62
+ raise ValueError(f"chunked prefill: first_blocks has {len(first_blocks)} entries for {n} rows")
63
+ C = self.chunk
64
+ resume_rows = [u for u in range(n) if resume_mask[u]]
65
+ if len(resume_rows) > 1:
66
+ raise ValueError(f"chunked prefill: {len(resume_rows)} resume rows {resume_rows} in one call (max 1)")
67
+ new_partials = [u for u in range(n) if not resume_mask[u] and not final_mask[u]]
68
+ if len(new_partials) + (1 if resume_rows and not final_mask[resume_rows[0]] else 0) > 1:
69
+ raise ValueError(
70
+ f"chunked prefill: more than one partial prompt in one call (resume rows {resume_rows}, "
71
+ f"new intermediate rows {new_partials}); the scheduler admits one long prompt at a time"
72
+ )
73
+ for u in range(n):
74
+ if int(ends[u]) < 1:
75
+ raise ValueError(f"chunked prefill: row {u} has end {ends[u]}")
76
+ if not final_mask[u] and int(ends[u]) % C != 0:
77
+ raise ValueError(
78
+ f"chunked prefill: intermediate row {u} ends at {ends[u]}, not a multiple of the chunk {C}"
79
+ )
80
+ owner = self.owner
81
+ if resume_rows:
82
+ u = resume_rows[0]
83
+ s = int(starts[u])
84
+ if s <= 0 or s % C != 0 or s >= int(ends[u]):
85
+ raise ValueError(
86
+ f"chunked prefill: resume row {u} has start {s}, end {ends[u]}: the start must be a positive "
87
+ f"multiple of {C} below the end"
88
+ )
89
+ if owner is None or owner.first_block != int(first_blocks[u]) or owner.next_pos != s:
90
+ raise ValueError(
91
+ f"chunked prefill: resume row {u} (first_block={int(first_blocks[u])}, start={s}) does not continue "
92
+ f"the scratch owner {owner}"
93
+ )
94
+ order = resume_rows + [u for u in range(n) if u not in resume_rows]
95
+ plans = []
96
+ for u in order:
97
+ resume = bool(resume_mask[u])
98
+ final = bool(final_mask[u])
99
+ park = unpark = False
100
+ if resume:
101
+ unpark = owner.parked
102
+ start = int(starts[u])
103
+ else:
104
+ park = owner is not None and not owner.parked
105
+ if park:
106
+ owner = replace(owner, parked=True)
107
+ start = 0
108
+ plans.append(RowPlan(u, start, int(ends[u]), resume, final, park, unpark))
109
+ if final:
110
+ if resume:
111
+ owner = None
112
+ else:
113
+ owner = ScratchOwner(int(first_blocks[u]), int(ends[u]), parked=False)
114
+ return plans, owner
code/models/demos/blackhole/qwen36/tt/qwen36_vllm.py CHANGED
@@ -84,19 +84,51 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
84
  # so no device-resident token/position chain is ever assumed; the plugin then overlaps only the
85
  # engine's scheduling/output work with the device step (its steady-decode fast path needs device sampling).
86
  # The gate is read when the platform queries the capability (_ModelCapabilities), not at import time.
 
 
 
 
 
 
 
 
 
 
 
 
87
  class _ModelCapabilities(dict):
88
- """dict whose ``supports_async_decode`` entry reads QWEN36_ASYNC_DECODE_OK at access time."""
 
89
 
90
  _ASYNC = "supports_async_decode"
 
 
 
 
 
 
91
 
92
  def __getitem__(self, key):
93
  if key == self._ASYNC:
94
  return os.environ.get("QWEN36_ASYNC_DECODE_OK", "0") == "1"
 
 
 
 
 
 
95
  return super().__getitem__(key)
96
 
 
 
 
 
 
97
  def get(self, key, default=None):
98
  if key == self._ASYNC:
99
  return self[key]
 
 
100
  return super().get(key, default)
101
 
102
  model_capabilities = _ModelCapabilities(
@@ -351,7 +383,26 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
351
  "batched (max_num_seqs>1) serving is text-only; multimodal is single-sequence "
352
  "(max_concurrency=1). Run the model at max_num_seqs=1 for image/video requests."
353
  )
354
- return self._prefill_forward_tp_batched(model, tokens, page_table, prompt_lens, kwargs.get("empty_slots"))
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
355
  vision_tokens = self._compute_vision_tokens(model, kwargs)
356
  if model.use_tp:
357
  return self._prefill_forward_tp(model, tokens, page_table, prompt_lens, vision_tokens=vision_tokens)
@@ -417,7 +468,9 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
417
  logger.info(f"Finished prefill up to {T} tokens, starting decode...")
418
  return logits, torch.zeros(1, dtype=torch.long)
419
 
420
- def _prefill_forward_tp_batched(self, model, tokens, page_table, prompt_lens, empty_slots):
 
 
421
  """TP batched (max_num_seqs>1) prefill: prefill each request in this step into its decode slot.
422
 
423
  vLLM prefills new requests while other slots decode, so each user's B=1 state is written into
@@ -428,6 +481,10 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
428
  page_table: torch [N, max_blocks] — row u = request u's blocks.
429
  prompt_lens: per-request real lengths (row u trimmed to prompt_lens[u]).
430
  empty_slots: per-request decode slot; defaults to range(N) (mirrors Generator.prefill_forward_text).
 
 
 
 
431
  Returns ([N, 1, vocab] host logits, [N] zero rope_deltas — text M-RoPE delta is 0, applied model-side).
432
  """
433
  N = tokens.shape[0]
@@ -437,8 +494,27 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
437
  empty_slots = [int(s) for s in empty_slots]
438
  token_ids_list = [tokens[u : u + 1, : plens[u]].to(torch.int32) for u in range(N)]
439
  pt = page_table if isinstance(page_table, torch.Tensor) else ttnn.to_torch(page_table)
440
- logger.info(f"Prefilling {N} user(s) into slots {empty_slots} (TP batched masked-bucket)")
441
- host_logits = model.prefill_paged_slots(token_ids_list, pt, empty_slots, valid_lens=plens)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
442
  logits = torch.cat([hl.reshape(1, 1, -1) for hl in host_logits], dim=0) # [N, 1, vocab]
443
  logger.info(f"Finished batched prefill of {N} user(s), starting decode...")
444
  return logits, torch.zeros(N, dtype=torch.long)
@@ -454,10 +530,26 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
454
  # BEFORE the decode trace reads it. The plugin remaps its own buffers (and the seed RNG via
455
  # super().decode_forward), but GDN state is model-internal, so mirror the same reindex here.
456
  # slot_remap is passed through unchanged so the seed-RNG remap inside super() still runs.
 
 
 
 
 
 
 
 
 
 
 
 
 
457
  if model.use_tp and model.args.max_batch_size > 1:
458
  slot_remap = kwargs.get("slot_remap")
459
  if slot_remap is not None:
460
- model._remap_gdn_slots(slot_remap)
 
 
 
461
  # Decode bucketing (default on; TT_DECODE_BUCKETING=0 off): slice host inputs to the
462
  # smallest power-of-2 width >= active prefix [0:num_active) before the base forward.
463
  # No runner edit / output re-pad — plugin reads unpadded_batch_size in slot order.
@@ -509,7 +601,23 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
509
  _sm = getattr(_m, "sampling", None)
510
  if _sm is not None and hasattr(_sm, "set_trace_bucket"):
511
  _sm.set_trace_bucket(B)
512
- return super().decode_forward(*args, **kwargs)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
513
 
514
  def warmup_model_prefill(self, kv_cache, enable_trace, *args, **kwargs):
515
  # Capture the chunk-prefill trace + warm the masked-bucket set so requests only replay
@@ -540,6 +648,9 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
540
  )
541
  prev = model._bind_gdn_prefill_scratch() if batched else None
542
  try:
 
 
 
543
  model.capture_prefill_trace_chunked(
544
  self.mesh_device, page_table, chunk_size=_PREFILL_WARMUP_CHUNK, capture_chunk_trace=True
545
  )
@@ -550,11 +661,20 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
550
  # Compile the device-side slot-write programs (QWEN36_GDN_SLOT_DEVICE_COPY=2: fill_cache + masked where) and
551
  # upload the per-slot row masks now, so the first real request does not pay ~450 ms for it.
552
  model.warmup_gdn_slot_write()
 
 
 
553
  # Steady-state view: weights + KV pool + GDN slot state + the persistent prefill buffers are all allocated
554
  # (the decode traces are captured earlier by warmup_model_decode). The free DRAM here, minus a margin for
555
  # the transient prefill activations, is the headroom QWEN36_MAX_TOKENS_ALL_USERS can grow into.
556
  _log_device_memory(self.mesh_device, "after prefill warmup")
557
 
 
 
 
 
 
 
558
  def warmup_model_decode(self, *args, **kwargs):
559
  # Defer to WarmupForwardMixin, which warms the paged-SDPA + GDN decode path at pos 0.
560
  # Drop stale `non_greedy_decoding_on_device` from the old vLLM plugin; no-op for Qwen.
 
84
  # so no device-resident token/position chain is ever assumed; the plugin then overlaps only the
85
  # engine's scheduling/output work with the device step (its steady-decode fast path needs device sampling).
86
  # The gate is read when the platform queries the capability (_ModelCapabilities), not at import time.
87
+ # QWEN36_CHUNKED_PREFILL=1 (default off; batched TP text serving) declares scheduler-driven chunked prefill:
88
+ # ``supports_chunked_prefill`` True plus ``tt_prefill_chunk_tokens`` (the 2048 chunk-trace unit; the plugin's chunk
89
+ # size must be a multiple of it), consumed by the plugin's chunk policy. A resumed chunk continues from the GDN
90
+ # state parked in the B=1 prefill scratch (model.prefill_paged_slots, tt/chunked_prefill.py). With the knob off
91
+ # both keys read as absent (``in`` False, ``.get`` returns the caller's default, ``[]`` of the chunk size raises),
92
+ # so every profile resolves exactly as before. The two keys are DYNAMIC: they live only in the accessors, not in
93
+ # the dict storage, so a copy (``dict(caps)``, ``{**caps}``) or ``.items()`` / iteration sees only the static
94
+ # entries. Consumers must query the class attribute through ``[]`` / ``in`` / ``.get`` (the plugin does), and a
95
+ # subclass that copies the dict (the DFlash class) must state its own chunked-prefill entries explicitly.
96
+ # The knob is model-side only: with QWEN36_CHUNKED_PREFILL=1 the park buffer (~147 MiB/device at TP=2) is
97
+ # allocated at warmup even when the plugin later leaves the policy off (async scheduling, kv_transfer_config /
98
+ # p1d1, --no-enable-chunked-prefill), which only costs KV-pool headroom; set the knob only on a chunking profile.
99
  class _ModelCapabilities(dict):
100
+ """dict whose ``supports_async_decode`` (QWEN36_ASYNC_DECODE_OK) and chunked-prefill (QWEN36_CHUNKED_PREFILL)
101
+ entries are read from the environment at access time."""
102
 
103
  _ASYNC = "supports_async_decode"
104
+ _CHUNKED = "supports_chunked_prefill"
105
+ _CHUNK_TOKENS = "tt_prefill_chunk_tokens"
106
+
107
+ @staticmethod
108
+ def _chunked_on():
109
+ return os.environ.get("QWEN36_CHUNKED_PREFILL", "0") == "1"
110
 
111
  def __getitem__(self, key):
112
  if key == self._ASYNC:
113
  return os.environ.get("QWEN36_ASYNC_DECODE_OK", "0") == "1"
114
+ if key == self._CHUNKED:
115
+ return self._chunked_on()
116
+ if key == self._CHUNK_TOKENS:
117
+ if not self._chunked_on():
118
+ raise KeyError(key)
119
+ return _PREFILL_WARMUP_CHUNK
120
  return super().__getitem__(key)
121
 
122
+ def __contains__(self, key):
123
+ if key in (self._CHUNK_TOKENS, self._CHUNKED):
124
+ return self._chunked_on()
125
+ return super().__contains__(key)
126
+
127
  def get(self, key, default=None):
128
  if key == self._ASYNC:
129
  return self[key]
130
+ if key in (self._CHUNKED, self._CHUNK_TOKENS):
131
+ return self[key] if self._chunked_on() else default
132
  return super().get(key, default)
133
 
134
  model_capabilities = _ModelCapabilities(
 
383
  "batched (max_num_seqs>1) serving is text-only; multimodal is single-sequence "
384
  "(max_concurrency=1). Run the model at max_num_seqs=1 for image/video requests."
385
  )
386
+ return self._prefill_forward_tp_batched(
387
+ model,
388
+ tokens,
389
+ page_table,
390
+ prompt_lens,
391
+ kwargs.get("empty_slots"),
392
+ start_pos=kwargs.get("start_pos"),
393
+ resume_mask=kwargs.get("prefill_resume_mask"),
394
+ final_mask=kwargs.get("prefill_final_mask"),
395
+ )
396
+ resume_mask = kwargs.get("prefill_resume_mask")
397
+ final_mask = kwargs.get("prefill_final_mask")
398
+ if (resume_mask is not None and any(bool(m) for m in resume_mask)) or (
399
+ final_mask is not None and not all(bool(m) for m in final_mask)
400
+ ):
401
+ # Chunks are only split while other requests decode, which needs max_num_seqs > 1 (the batched path).
402
+ raise ValueError(
403
+ f"chunked prefill needs the batched TP path (use_tp={model.use_tp}, "
404
+ f"max_batch_size={model.args.max_batch_size}); resume_mask={resume_mask} final_mask={final_mask}"
405
+ )
406
  vision_tokens = self._compute_vision_tokens(model, kwargs)
407
  if model.use_tp:
408
  return self._prefill_forward_tp(model, tokens, page_table, prompt_lens, vision_tokens=vision_tokens)
 
468
  logger.info(f"Finished prefill up to {T} tokens, starting decode...")
469
  return logits, torch.zeros(1, dtype=torch.long)
470
 
471
+ def _prefill_forward_tp_batched(
472
+ self, model, tokens, page_table, prompt_lens, empty_slots, start_pos=None, resume_mask=None, final_mask=None
473
+ ):
474
  """TP batched (max_num_seqs>1) prefill: prefill each request in this step into its decode slot.
475
 
476
  vLLM prefills new requests while other slots decode, so each user's B=1 state is written into
 
481
  page_table: torch [N, max_blocks] — row u = request u's blocks.
482
  prompt_lens: per-request real lengths (row u trimmed to prompt_lens[u]).
483
  empty_slots: per-request decode slot; defaults to range(N) (mirrors Generator.prefill_forward_text).
484
+ start_pos / resume_mask / final_mask: scheduler-driven chunked prefill (QWEN36_CHUNKED_PREFILL=1): a row with
485
+ resume_mask True continues its prompt at start_pos (the chunk start); final_mask False marks an
486
+ intermediate chunk (prompt_lens = the chunk end). Rows are always passed from token 0. Without the
487
+ masks start_pos is ignored and every row prefills from 0 (the pre-chunking behaviour).
488
  Returns ([N, 1, vocab] host logits, [N] zero rope_deltas — text M-RoPE delta is 0, applied model-side).
489
  """
490
  N = tokens.shape[0]
 
494
  empty_slots = [int(s) for s in empty_slots]
495
  token_ids_list = [tokens[u : u + 1, : plens[u]].to(torch.int32) for u in range(N)]
496
  pt = page_table if isinstance(page_table, torch.Tensor) else ttnn.to_torch(page_table)
497
+ chunked = resume_mask is not None or final_mask is not None
498
+ chunk_desc = ""
499
+ if chunked and (
500
+ (resume_mask is not None and any(resume_mask)) or (final_mask is not None and not all(final_mask))
501
+ ):
502
+ # (start, end, resume, final) per row: the one line that shows a split prefill in the server log
503
+ res = [bool(resume_mask[u]) if resume_mask is not None else False for u in range(N)]
504
+ fin = [bool(final_mask[u]) if final_mask is not None else True for u in range(N)]
505
+ chunk_desc = " chunked rows " + str(
506
+ [(int(start_pos[u]) if res[u] else 0, plens[u], res[u], fin[u]) for u in range(N)]
507
+ )
508
+ logger.info(f"Prefilling {N} user(s) into slots {empty_slots} (TP batched masked-bucket){chunk_desc}")
509
+ host_logits = model.prefill_paged_slots(
510
+ token_ids_list,
511
+ pt,
512
+ empty_slots,
513
+ valid_lens=plens,
514
+ start_positions=[int(start_pos[u]) for u in range(N)] if chunked and start_pos is not None else None,
515
+ resume_mask=[bool(resume_mask[u]) for u in range(N)] if resume_mask is not None else None,
516
+ final_mask=[bool(final_mask[u]) for u in range(N)] if final_mask is not None else None,
517
+ )
518
  logits = torch.cat([hl.reshape(1, 1, -1) for hl in host_logits], dim=0) # [N, 1, vocab]
519
  logger.info(f"Finished batched prefill of {N} user(s), starting decode...")
520
  return logits, torch.zeros(N, dtype=torch.long)
 
530
  # BEFORE the decode trace reads it. The plugin remaps its own buffers (and the seed RNG via
531
  # super().decode_forward), but GDN state is model-internal, so mirror the same reindex here.
532
  # slot_remap is passed through unchanged so the seed-RNG remap inside super() still runs.
533
+ # QWEN36_DECODE_SLOW_LOG=1 (perf triage, off by default): drain the device FIRST, so neither the step time
534
+ # nor the remap time includes work still queued from the previous step (a prefill's slot-write copies, the
535
+ # conv-history sync), then time the GDN slot remap phase by phase (model._remap_gdn_slots timing=: device
536
+ # synced between phases) and log every remap step and every other step slower than 150 ms. The syncs slow
537
+ # the logged steps a little; do not quote TPOT from a run with it on.
538
+ _slow_log = os.environ.get("QWEN36_DECODE_SLOW_LOG", "0") == "1"
539
+ if _slow_log:
540
+ _tq0 = time.perf_counter()
541
+ ttnn.synchronize_device(self.mesh_device)
542
+ _t_queued = time.perf_counter() - _tq0
543
+ _td0 = time.perf_counter() if _slow_log else 0.0
544
+ _t_remap = 0.0
545
+ _remap_timing = {} if _slow_log else None
546
  if model.use_tp and model.args.max_batch_size > 1:
547
  slot_remap = kwargs.get("slot_remap")
548
  if slot_remap is not None:
549
+ model._remap_gdn_slots(slot_remap, timing=_remap_timing)
550
+ if _slow_log:
551
+ ttnn.synchronize_device(self.mesh_device)
552
+ _t_remap = time.perf_counter() - _td0
553
  # Decode bucketing (default on; TT_DECODE_BUCKETING=0 off): slice host inputs to the
554
  # smallest power-of-2 width >= active prefix [0:num_active) before the base forward.
555
  # No runner edit / output re-pad — plugin reads unpadded_batch_size in slot order.
 
601
  _sm = getattr(_m, "sampling", None)
602
  if _sm is not None and hasattr(_sm, "set_trace_bucket"):
603
  _sm.set_trace_bucket(B)
604
+ if not _slow_log:
605
+ return super().decode_forward(*args, **kwargs)
606
+ out = super().decode_forward(*args, **kwargs)
607
+ _dt = time.perf_counter() - _td0
608
+ if _dt > 0.15 or kwargs.get("slot_remap") is not None:
609
+ _rt = _remap_timing or {}
610
+ _phases = " ".join(
611
+ f"{k}={1e3 * _rt[k]:.0f}" for k in ("taps_s", "rec_s", "conv_s", "packed_s", "hist_sync_s") if k in _rt
612
+ )
613
+ logger.info(
614
+ f"[DECODE_SLOW] {1e3 * _dt:.0f} ms B={int(tokens.shape[0]) if tokens is not None else None} "
615
+ f"remap={'yes' if kwargs.get('slot_remap') is not None else 'no'} remap_ms={1e3 * _t_remap:.0f} "
616
+ f"queued_ms={1e3 * _t_queued:.0f} moved={_rt.get('moved', 0)} cross_parity={_rt.get('cross_parity', 0)} "
617
+ f"packed={'yes' if _rt.get('packed') else 'no'} phases_ms[{_phases}] "
618
+ f"kw={sorted(k for k in kwargs if kwargs[k] is not None and k not in ('tokens', 'page_table', 'kv_cache'))}"
619
+ )
620
+ return out
621
 
622
  def warmup_model_prefill(self, kv_cache, enable_trace, *args, **kwargs):
623
  # Capture the chunk-prefill trace + warm the masked-bucket set so requests only replay
 
648
  )
649
  prev = model._bind_gdn_prefill_scratch() if batched else None
650
  try:
651
+ if batched and self.model_capabilities.get("supports_chunked_prefill", False):
652
+ # Chunked prefill parks a partial prompt's scratch state: allocate + compile it before any trace.
653
+ model.ensure_gdn_park_buffer()
654
  model.capture_prefill_trace_chunked(
655
  self.mesh_device, page_table, chunk_size=_PREFILL_WARMUP_CHUNK, capture_chunk_trace=True
656
  )
 
661
  # Compile the device-side slot-write programs (QWEN36_GDN_SLOT_DEVICE_COPY=2: fill_cache + masked where) and
662
  # upload the per-slot row masks now, so the first real request does not pay ~450 ms for it.
663
  model.warmup_gdn_slot_write()
664
+ # Same for the fast GDN slot remap (QWEN36_GDN_REMAP_FAST): a first-seen remap program costs ~300 ms of JIT.
665
+ if self._warm_gdn_remap():
666
+ model.warmup_gdn_remap()
667
  # Steady-state view: weights + KV pool + GDN slot state + the persistent prefill buffers are all allocated
668
  # (the decode traces are captured earlier by warmup_model_decode). The free DRAM here, minus a margin for
669
  # the transient prefill activations, is the headroom QWEN36_MAX_TOKENS_ALL_USERS can grow into.
670
  _log_device_memory(self.mesh_device, "after prefill warmup")
671
 
672
+ def _warm_gdn_remap(self):
673
+ """Whether warmup_model_prefill compiles the fast GDN slot-remap programs: yes for every class whose decode
674
+ applies the plugin's slot_remap to the device GDN state (this one). Subclasses that never remap on device
675
+ (the speculative DFlash decode composes slot_remap into a row indirection) return False."""
676
+ return True
677
+
678
  def warmup_model_decode(self, *args, **kwargs):
679
  # Defer to WarmupForwardMixin, which warms the paged-SDPA + GDN decode path at pos 0.
680
  # Drop stale `non_greedy_decoding_on_device` from the old vLLM plugin; no-op for Qwen.
code/models/demos/blackhole/qwen36/tt/qwen36_vllm_dflash.py CHANGED
@@ -123,6 +123,63 @@ _PREFILL_CHUNK = 2048
123
  # verify's reach past the block (W committed positions + K+1 candidate rows).
124
  _MAX_DRAFT = 15
125
  _DEBUG = os.environ.get("QWEN36_DFLASH_DEBUG", "0") == "1"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
126
 
127
 
128
  def parse_buckets(spec):
@@ -206,24 +263,36 @@ def bucket_id(bucket):
206
  class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
207
  """Qwen36ForCausalLM + model-internal multi-user DFlash2 speculation on decode steps (see module doc)."""
208
 
209
- model_capabilities = {
210
- **Qwen36ForCausalLM.model_capabilities,
211
- # A decode step commits exactly _W tokens per request (EOS-filled at a stop).
212
- "output_tokens_per_step": _W,
213
- # Block on decode steps; prefill anchors are plain width-1 rows.
214
- "tt_adaptive_block_output": _W > 1,
215
- # ...for EVERY request of the step (one multi-user speculative step), not only when solo.
216
- "tt_adaptive_block_batched": _W > 1,
217
- # Ragged rows: one speculative iteration per step, each row 1..W real ids then -1 padding
218
- # (QWEN36_DFLASH_RAGGED=1; the plugin must implement the same contract).
219
- "tt_adaptive_block_ragged": _W > 1 and _RAGGED,
220
- # EVERY text prompt speculates (0 = no prompt-length frontier): the plain decode / chunk-prefill
221
- # traces never run on the request path (a spec replay after them hangs the device).
222
- "tt_adaptive_block_max_prompt_tokens": 0,
223
- # The block step writes the W committed positions AND the last verify's K+1 candidate rows into
224
- # the paged KV inside one step: have the scheduler allocate that reach up front.
225
- "tt_block_output_kv_lookahead_tokens": (_W + _MAX_DRAFT + 1) if _W > 1 else 0,
226
- }
 
 
 
 
 
 
 
 
 
 
 
 
227
 
228
  def __init__(self, *args, **kwargs):
229
  super().__init__(*args, **kwargs)
@@ -247,6 +316,22 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
247
  self._buckets = () # (B_cfg, T_cfg) verify geometries, largest B == max_num_seqs (module doc)
248
  self._multi_bucket = False
249
  self._last_bucket = None # bucket id plan() last put in force (logged on change)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
250
  if _W > 1:
251
  if not model.use_tp:
252
  raise RuntimeError(
@@ -262,6 +347,11 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
262
  f"{_BUCKETS_ENV}={os.environ.get(_BUCKETS_ENV)!r} needs the multi-bucket decoder "
263
  "(dflash2_serving.DFlash2DualBucketDecoder), which this tree does not have"
264
  )
 
 
 
 
 
265
  logger.info(
266
  f"Qwen36DFlash serving: slots={B} block W={_W} tokens/step{' (ragged: one iteration/step)' if _RAGGED else ''}, "
267
  f"buckets={','.join(bucket_id(bt) for bt in self._buckets)} "
@@ -416,10 +506,15 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
416
  return pt[:, :nb].contiguous()
417
  return torch.cat([pt, torch.zeros(1, nb - pt.shape[1], dtype=torch.int32)], dim=1)
418
 
419
- def _spec_prefill(self, model, dec, phys, prompt, T, pt_row):
420
  """Eager tap-capturing prefill of ONE request into physical slot ``phys`` (masked bucket for a
421
  short prompt, 2048-token chunks + masked tail for a long one), each chunk's taps ingested into the
422
- drafter's ring for that slot. Returns host logits [1, vocab] (float)."""
 
 
 
 
 
423
 
424
  def on_chunk(hidden, chunk_start, valid_len):
425
  taps = model.take_dflash_eager_taps()
@@ -427,12 +522,34 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
427
  raise RuntimeError("Qwen36DFlash: eager prefill captured no drafter taps (bucket trace gate on?)")
428
  dec.ingest_prompt(phys, taps, chunk_start + valid_len, chunk_start=chunk_start)
429
 
430
- dec.ctx_len[phys] = 0
 
 
 
 
 
 
 
431
  model._dflash_tap = True
432
  try:
433
- logits_dev = model.prefill_for_spec(prompt, self._spec_pref_pt(model, pt_row), T, on_chunk, slot=phys)
 
 
 
 
 
434
  finally:
435
  model._dflash_tap = False
 
 
 
 
 
 
 
 
 
 
436
  if _FAST_LOGITS:
437
  # Replicated, tile-padded to 32 rows: untilize on device (31 padding rows dropped, 16 MB -> 0.5 MB per
438
  # device) and read device 0 only -- the QWEN36_PREFILL_LOGITS_FAST path of prefill_paged_slots. The
@@ -469,6 +586,14 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
469
  self._in_warmup = False
470
  return out
471
 
 
 
 
 
 
 
 
 
472
  def warmup_model_prefill(self, *args, **kwargs):
473
  self._in_warmup = True
474
  try:
@@ -522,6 +647,11 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
522
  S0 = min(buckets[0], _PREFILL_CHUNK - 1)
523
  for u in range(1, B):
524
  self._spec_prefill(model, dec, u, dummy_prompt(S0, seed=100 + u), S0, rows[u])
 
 
 
 
 
525
  ttnn.synchronize_device(model.mesh_device)
526
  logger.info(f"Qwen36DFlash phase-1 warmup (alloc + compile) done in {time.perf_counter() - t0:.1f}s")
527
  self._warm_rows = rows
@@ -706,6 +836,9 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
706
  dec.set_bucket(ids[bt])
707
  dec.step()
708
  dec.end(0)
 
 
 
709
  ttnn.synchronize_device(dev)
710
  n1 = int(count())
711
  if n1 != n0:
@@ -721,6 +854,7 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
721
  logger.info(
722
  f"Qwen36DFlash post-capture guard: program cache unchanged at {n0} entries over the warm sweep + one "
723
  f"traced step per bucket ({','.join(ids[bt] for bt in self._buckets)})"
 
724
  )
725
 
726
  def _spec_capture(self):
@@ -762,6 +896,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
762
  self._reset_bucket_policy(dec)
763
  self._spec = dec
764
  self._spec_pre = None
 
 
 
 
765
  logger.info(f"Qwen36DFlash phase-2 warmup (captures) done in {time.perf_counter() - t0:.1f}s: {stats} (W={_W})")
766
 
767
  def _reset_bucket_policy(self, dec):
@@ -818,6 +956,14 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
818
 
819
  def prefill_forward(self, tokens, page_table, kv_cache, prompt_lens, **kwargs):
820
  if _W <= 1 or not self._spec_ready():
 
 
 
 
 
 
 
 
821
  return super().prefill_forward(tokens, page_table, kv_cache, prompt_lens, **kwargs)
822
  model = self.model[0]
823
  dec = self._spec
@@ -830,6 +976,41 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
830
  empty_slots = kwargs.get("empty_slots")
831
  logical = [int(s) for s in empty_slots] if empty_slots is not None else list(range(N))
832
  pt = torch.as_tensor(page_table)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
833
  out = []
834
  for u in range(N):
835
  phys = self._phys[logical[u]]
@@ -847,6 +1028,114 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
847
  logger.info(f"Finished prefill of {N} request(s), starting decode...")
848
  return torch.cat(out, dim=0), torch.zeros(N, dtype=torch.long)
849
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
850
  # ------------------------------------------------------------------ decode: one block per step, all live slots
851
  def _no_device_sampler(self):
852
  """True when no model on the mesh has the ttnn device sampler (TP=2: 124,160 logits/device > its 64K cap)."""
@@ -854,6 +1143,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
854
 
855
  def decode_forward(self, *args, **kwargs):
856
  if _W <= 1 or not self._spec_ready():
 
 
 
 
857
  if _W > 1 and kwargs.get("sampling_params") is not None and self._no_device_sampler():
858
  # Pre-arm plain step on a mesh without a device sampler: host logits (the runner samples).
859
  kwargs = dict(kwargs, sampling_params=None)
@@ -891,6 +1184,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
891
  # user), so the begins below land in the geometry in force; a mis-ordering is a decoder assert, never a
892
  # wrong-geometry seed. Single bucket: a host no-op.
893
  plan = getattr(dec, "plan", None)
 
 
 
 
894
  if plan is not None:
895
  live_after = {phys for phys in range(B) if dec.active[phys]}
896
  for i in range(Bp):
@@ -900,6 +1197,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
900
  if bucket != self._last_bucket:
901
  logger.info(f"Qwen36DFlash: verify bucket {bucket} in force ({len(live_after)} live slot(s))")
902
  self._last_bucket = bucket
 
 
 
 
903
  live_rows = []
904
  nosession_rows = [] # live rows without a session: EOS so the request ends (ragged: [EOS, -1, ...])
905
  for i in range(Bp):
@@ -934,9 +1235,21 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
934
  live_rows.append((i, phys))
935
  t0 = time.perf_counter() if _DEBUG else 0.0
936
  if _RAGGED:
 
 
 
937
  out = self._decode_ragged(dec, live_rows, Bp, t0)
938
  for i in nosession_rows:
939
  out[i, 0] = self._eos_fill
 
 
 
 
 
 
 
 
 
940
  return out
941
  # Step until every live row can fill its block. A row whose carry already holds a stop token
942
  # needs nothing: it emits through the stop this step (latency for a normal request); the
@@ -1068,6 +1381,11 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
1068
  self._spec.end(phys)
1069
  self._pending[phys] = None
1070
  self._carry[phys], self._stopped[phys], self._prev_tail[phys] = [], False, None
 
 
 
 
 
1071
 
1072
  def release_persistent_capture(self) -> None:
1073
  if self._spec is None:
 
123
  # verify's reach past the block (W committed positions + K+1 candidate rows).
124
  _MAX_DRAFT = 15
125
  _DEBUG = os.environ.get("QWEN36_DFLASH_DEBUG", "0") == "1"
126
+ # Scheduler-driven chunked prefill (opt-in; profiles/tp4_chunked/DESIGN.md): with the knob on, the class declares
127
+ # supports_chunked_prefill + tt_prefill_chunk_tokens (the 2048 eager unit) + tt_block_output_chunked_prefill, the plugin's
128
+ # TT chunk policy may split a long prompt into 2048-aligned chunks interleaved with the other users' speculative decode
129
+ # steps, and prefill_forward resumes a partial prompt through the eager prefill_for_spec path (GDN state parked in the
130
+ # B=1 scratch between calls, tt/chunked_prefill.py; drafter context in the slot's own ring). Off (default): the
131
+ # capability dict, prefill_forward, the warm-up and the guards are exactly today's.
132
+ _CP_ENV = "QWEN36_DFLASH_CHUNKED_PREFILL"
133
+ # Per-row prefill timing of the chunked path (device-synchronised; measurement only, off by default).
134
+ _CP_TIMING = os.environ.get("QWEN36_DFLASH_CP_TIMING", "0") == "1"
135
+ # Drained per-step timers in decode_forward (plan/switch, begins, step; measurement only, off by default).
136
+ _STEP_LOG = os.environ.get("QWEN36_DFLASH_STEP_LOG", "0") == "1"
137
+
138
+
139
+ def dflash_chunked_prefill_on():
140
+ """The chunked-prefill knob, read at access time (speculation must be on: a plain-serving DFlash class is the base
141
+ class's behaviour and never declares the block-output chunk capability)."""
142
+ return _W > 1 and os.environ.get(_CP_ENV, "0") == "1"
143
+
144
+
145
+ class _DFlashCapabilities(dict):
146
+ """The DFlash class's capability dict: the static entries are exactly today's (a ``dict()`` / ``{**}`` copy, ``items()``
147
+ and iteration see only them, ``supports_chunked_prefill`` present and False); with QWEN36_DFLASH_CHUNKED_PREFILL=1
148
+ (read at access time) ``[]`` / ``in`` / ``.get`` report chunked prefill on, the 2048-token chunk unit and the
149
+ block-output chunk contract. Only these three keys are dynamic (``supports_async_decode`` stays the static False).
150
+ """
151
+
152
+ _CHUNKED = "supports_chunked_prefill"
153
+ _DYNAMIC_ONLY = ("tt_prefill_chunk_tokens", "tt_block_output_chunked_prefill")
154
+
155
+ @staticmethod
156
+ def _on_values():
157
+ return {
158
+ "supports_chunked_prefill": True,
159
+ "tt_prefill_chunk_tokens": _PREFILL_CHUNK,
160
+ "tt_block_output_chunked_prefill": True,
161
+ }
162
+
163
+ def __getitem__(self, key):
164
+ if key == self._CHUNKED or key in self._DYNAMIC_ONLY:
165
+ if dflash_chunked_prefill_on():
166
+ return self._on_values()[key]
167
+ if key in self._DYNAMIC_ONLY:
168
+ raise KeyError(key)
169
+ return super().__getitem__(key)
170
+
171
+ def __contains__(self, key):
172
+ if key in self._DYNAMIC_ONLY:
173
+ return dflash_chunked_prefill_on()
174
+ return super().__contains__(key)
175
+
176
+ def get(self, key, default=None):
177
+ if key == self._CHUNKED or key in self._DYNAMIC_ONLY:
178
+ if dflash_chunked_prefill_on():
179
+ return self._on_values()[key]
180
+ if key in self._DYNAMIC_ONLY:
181
+ return default
182
+ return super().get(key, default)
183
 
184
 
185
  def parse_buckets(spec):
 
263
  class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
264
  """Qwen36ForCausalLM + model-internal multi-user DFlash2 speculation on decode steps (see module doc)."""
265
 
266
+ model_capabilities = _DFlashCapabilities(
267
+ {
268
+ **Qwen36ForCausalLM.model_capabilities,
269
+ # No scheduler-driven chunked prefill by default (QWEN36_CHUNKED_PREFILL is the plain class's knob and must not
270
+ # reach this class): the plain class reads its knob at access time; this dict copies only its static entries, so
271
+ # state the default explicitly (with _W == 1 the block-output refusals in the plugin would not catch it).
272
+ # QWEN36_DFLASH_CHUNKED_PREFILL=1 turns it on through _DFlashCapabilities' accessors (module constants).
273
+ "supports_chunked_prefill": False,
274
+ # A decode step commits exactly _W tokens per request (EOS-filled at a stop).
275
+ "output_tokens_per_step": _W,
276
+ # Block on decode steps; prefill anchors are plain width-1 rows.
277
+ "tt_adaptive_block_output": _W > 1,
278
+ # ...for EVERY request of the step (one multi-user speculative step), not only when solo.
279
+ "tt_adaptive_block_batched": _W > 1,
280
+ # Ragged rows: one speculative iteration per step, each row 1..W real ids then -1 padding
281
+ # (QWEN36_DFLASH_RAGGED=1; the plugin must implement the same contract).
282
+ "tt_adaptive_block_ragged": _W > 1 and _RAGGED,
283
+ # EVERY text prompt speculates (0 = no prompt-length frontier): the plain decode / chunk-prefill
284
+ # traces never run on the request path (a spec replay after them hangs the device).
285
+ "tt_adaptive_block_max_prompt_tokens": 0,
286
+ # The block step writes the W committed positions AND the last verify's K+1 candidate rows into
287
+ # the paged KV inside one step: have the scheduler allocate that reach up front.
288
+ "tt_block_output_kv_lookahead_tokens": (_W + _MAX_DRAFT + 1) if _W > 1 else 0,
289
+ }
290
+ )
291
+
292
+ # Chunked-prefill state (set per instance in __init__; class defaults for instances built without it, e.g. host tests)
293
+ _cp_on = False
294
+ _cp_owner_phys = None
295
+ _forbid_plain = False
296
 
297
  def __init__(self, *args, **kwargs):
298
  super().__init__(*args, **kwargs)
 
316
  self._buckets = () # (B_cfg, T_cfg) verify geometries, largest B == max_num_seqs (module doc)
317
  self._multi_bucket = False
318
  self._last_bucket = None # bucket id plan() last put in force (logged on change)
319
+ # Chunked prefill (QWEN36_DFLASH_CHUNKED_PREFILL=1): the physical slot whose partial prompt the B=1 scratch (or its
320
+ # park buffer) holds, the third ownership check next to the planner's first block and next position.
321
+ self._cp_on = dflash_chunked_prefill_on()
322
+ self._cp_owner_phys = None
323
+ self._forbid_plain = False # armed at the end of the spec capture under the knob (plain-trace tripwire)
324
+ if self._cp_on:
325
+ if B <= 1:
326
+ raise RuntimeError(
327
+ f"{_CP_ENV}=1 needs the batched B=1 prefill scratch (--max-num-seqs > 1); got max_num_seqs={B} "
328
+ "(single-user-dflash2 has nothing to interleave a chunk with)"
329
+ )
330
+ if int(model.num_devices) != 4:
331
+ raise RuntimeError(
332
+ f"{_CP_ENV}=1 is validated at TP=4 (P150x4 / P300x2) only; got {int(model.num_devices)} device(s) "
333
+ "(the TP=2 spec prefill path is changing under lane Q; unset the knob)"
334
+ )
335
  if _W > 1:
336
  if not model.use_tp:
337
  raise RuntimeError(
 
347
  f"{_BUCKETS_ENV}={os.environ.get(_BUCKETS_ENV)!r} needs the multi-bucket decoder "
348
  "(dflash2_serving.DFlash2DualBucketDecoder), which this tree does not have"
349
  )
350
+ if self._cp_on:
351
+ logger.info(
352
+ f"Qwen36DFlash serving: scheduler-driven chunked prefill ON ({_CP_ENV}=1, chunk unit "
353
+ f"{_PREFILL_CHUNK}): partial prompts resume through the eager spec prefill"
354
+ )
355
  logger.info(
356
  f"Qwen36DFlash serving: slots={B} block W={_W} tokens/step{' (ragged: one iteration/step)' if _RAGGED else ''}, "
357
  f"buckets={','.join(bucket_id(bt) for bt in self._buckets)} "
 
506
  return pt[:, :nb].contiguous()
507
  return torch.cat([pt, torch.zeros(1, nb - pt.shape[1], dtype=torch.int32)], dim=1)
508
 
509
+ def _spec_prefill(self, model, dec, phys, prompt, T, pt_row, start=0, final=True):
510
  """Eager tap-capturing prefill of ONE request into physical slot ``phys`` (masked bucket for a
511
  short prompt, 2048-token chunks + masked tail for a long one), each chunk's taps ingested into the
512
+ drafter's ring for that slot. Returns host logits [1, vocab] (float).
513
+
514
+ Chunked prefill: ``start`` > 0 resumes the slot's partial prompt at ``start`` (its drafter frontier
515
+ ``dec.ctx_len[phys]`` must equal ``start``; the GDN state must be in the bound scratch, see prefill_for_spec);
516
+ ``final=False`` is an intermediate chunk ending at ``T``: no logits (returns None), the drafter frontier is
517
+ left at ``T``."""
518
 
519
  def on_chunk(hidden, chunk_start, valid_len):
520
  taps = model.take_dflash_eager_taps()
 
522
  raise RuntimeError("Qwen36DFlash: eager prefill captured no drafter taps (bucket trace gate on?)")
523
  dec.ingest_prompt(phys, taps, chunk_start + valid_len, chunk_start=chunk_start)
524
 
525
+ chunked = bool(start) or not final
526
+ if not start:
527
+ dec.ctx_len[phys] = 0
528
+ elif dec.ctx_len[phys] != start:
529
+ raise RuntimeError(
530
+ f"Qwen36DFlash chunked prefill: slot {phys} resumes at {start} but its drafter context ends at "
531
+ f"{dec.ctx_len[phys]} (the continuation moved slots or a chunk was lost)"
532
+ )
533
  model._dflash_tap = True
534
  try:
535
+ if chunked:
536
+ logits_dev = model.prefill_for_spec(
537
+ prompt, self._spec_pref_pt(model, pt_row), T, on_chunk, slot=phys, start=int(start), final=final
538
+ )
539
+ else:
540
+ logits_dev = model.prefill_for_spec(prompt, self._spec_pref_pt(model, pt_row), T, on_chunk, slot=phys)
541
  finally:
542
  model._dflash_tap = False
543
+ if not final:
544
+ if logits_dev is not None:
545
+ ttnn.deallocate(logits_dev)
546
+ raise RuntimeError("Qwen36DFlash chunked prefill: an intermediate chunk produced logits")
547
+ if dec.ctx_len[phys] != T:
548
+ raise RuntimeError(
549
+ f"Qwen36DFlash chunked prefill: intermediate chunk [{start}, {T}) left slot {phys}'s drafter "
550
+ f"context at {dec.ctx_len[phys]}"
551
+ )
552
+ return None
553
  if _FAST_LOGITS:
554
  # Replicated, tile-padded to 32 rows: untilize on device (31 padding rows dropped, 16 MB -> 0.5 MB per
555
  # device) and read device 0 only -- the QWEN36_PREFILL_LOGITS_FAST path of prefill_paged_slots. The
 
586
  self._in_warmup = False
587
  return out
588
 
589
+ def _warm_gdn_remap(self):
590
+ """Speculative serving (_W > 1) never runs remap_slots: decode_forward composes the plugin's slot_remap into
591
+ _phys, and plain decode is not used once the spec traces exist (the _forbid_plain tripwire). So
592
+ warmup_model_prefill must not compile the fast-remap programs either: at TP=2 (tp2-dflash2, B=4) that would
593
+ run 4 remaps on the B=4 GDN state after the spec traces were captured, shapes no test or served run exercised
594
+ (lane R review R1). _W == 1 is the plain decode path, which does remap."""
595
+ return _W <= 1
596
+
597
  def warmup_model_prefill(self, *args, **kwargs):
598
  self._in_warmup = True
599
  try:
 
647
  S0 = min(buckets[0], _PREFILL_CHUNK - 1)
648
  for u in range(1, B):
649
  self._spec_prefill(model, dec, u, dummy_prompt(S0, seed=100 + u), S0, rows[u])
650
+ if self._cp_on:
651
+ # 3) Chunked prefill: the park buffer (before ANY trace is captured; the base prefill warm-up's own call is
652
+ # then a no-op) and every resume / park program, through the chunk policy's own orchestration.
653
+ model.ensure_gdn_park_buffer()
654
+ self._cp_warm_sequence(model, dec, rows, S0, "phase-1 warm-up")
655
  ttnn.synchronize_device(model.mesh_device)
656
  logger.info(f"Qwen36DFlash phase-1 warmup (alloc + compile) done in {time.perf_counter() - t0:.1f}s")
657
  self._warm_rows = rows
 
836
  dec.set_bucket(ids[bt])
837
  dec.step()
838
  dec.end(0)
839
+ if self._cp_on:
840
+ # ...and one chunked prompt with a rider between its chunks (park, unpark, resumed chunk and tail).
841
+ self._cp_warm_sequence(model, dec, rows, S0, "post-capture guard")
842
  ttnn.synchronize_device(dev)
843
  n1 = int(count())
844
  if n1 != n0:
 
854
  logger.info(
855
  f"Qwen36DFlash post-capture guard: program cache unchanged at {n0} entries over the warm sweep + one "
856
  f"traced step per bucket ({','.join(ids[bt] for bt in self._buckets)})"
857
+ + (" + a chunked prefill with a rider" if self._cp_on else "")
858
  )
859
 
860
  def _spec_capture(self):
 
896
  self._reset_bucket_policy(dec)
897
  self._spec = dec
898
  self._spec_pre = None
899
+ if self._cp_on:
900
+ # Tripwire (hang class): from here on a plain trace replay raises instead of hanging the device.
901
+ model._forbid_plain_traces = True
902
+ self._forbid_plain = True
903
  logger.info(f"Qwen36DFlash phase-2 warmup (captures) done in {time.perf_counter() - t0:.1f}s: {stats} (W={_W})")
904
 
905
  def _reset_bucket_policy(self, dec):
 
956
 
957
  def prefill_forward(self, tokens, page_table, kv_cache, prompt_lens, **kwargs):
958
  if _W <= 1 or not self._spec_ready():
959
+ # Reachable with the tripwire armed only after release_persistent_capture (_spec None) while a plugin
960
+ # warm-up runs again (_in_warmup; otherwise _spec_ready() raises first): a plain forward would replay plain
961
+ # traces in a process that captured spec traces. On the request path _spec_ready() is True whenever the
962
+ # tripwire is armed; there the guard is model._check_plain_trace_allowed at the plain prefill replay sites.
963
+ if self._forbid_plain:
964
+ raise RuntimeError(
965
+ "Qwen36DFlash: plain prefill after the speculative traces were captured (hang class)"
966
+ )
967
  return super().prefill_forward(tokens, page_table, kv_cache, prompt_lens, **kwargs)
968
  model = self.model[0]
969
  dec = self._spec
 
976
  empty_slots = kwargs.get("empty_slots")
977
  logical = [int(s) for s in empty_slots] if empty_slots is not None else list(range(N))
978
  pt = torch.as_tensor(page_table)
979
+ resume_mask = kwargs.get("prefill_resume_mask")
980
+ final_mask = kwargs.get("prefill_final_mask")
981
+ if resume_mask is not None or final_mask is not None:
982
+ # The plugin's TT chunk policy (tt_block_output_chunked_prefill) passes both masks on EVERY prefill step.
983
+ resume_mask = [bool(x) for x in (resume_mask if resume_mask is not None else [False] * N)]
984
+ final_mask = [bool(x) for x in (final_mask if final_mask is not None else [True] * N)]
985
+ if len(resume_mask) != N or len(final_mask) != N:
986
+ raise RuntimeError(
987
+ f"Qwen36DFlash: prefill masks of {len(resume_mask)}/{len(final_mask)} rows for {N} prompt rows"
988
+ )
989
+ trivial = not any(resume_mask) and all(final_mask)
990
+ if not trivial and not self._cp_on:
991
+ raise RuntimeError(
992
+ f"Qwen36DFlash: the runner asked for a chunked prefill (resume={resume_mask}, final={final_mask}) "
993
+ f"but {_CP_ENV} is off"
994
+ )
995
+ # Whole prompts with no partial held in the scratch take today's loop below, byte for byte. Anything else
996
+ # (a resume, an intermediate chunk, or whole prompts while a partial is held: it must be parked first)
997
+ # goes through the planner.
998
+ if not trivial or model._chunked_prefill_planner().owner is not None:
999
+ starts = kwargs.get("start_pos")
1000
+ starts = [int(starts[u]) for u in range(N)] if starts is not None else [0] * N
1001
+ out = self._prefill_planned(
1002
+ model,
1003
+ dec,
1004
+ [torch.as_tensor(tokens)[u : u + 1, : plens[u]].to(torch.int32) for u in range(N)],
1005
+ [pt[u].reshape(-1).clone() for u in range(N)],
1006
+ [self._phys[logical[u]] for u in range(N)],
1007
+ starts,
1008
+ plens,
1009
+ resume_mask,
1010
+ final_mask,
1011
+ )
1012
+ logger.info(f"Finished prefill of {N} request(s), starting decode...")
1013
+ return torch.cat(out, dim=0), torch.zeros(N, dtype=torch.long)
1014
  out = []
1015
  for u in range(N):
1016
  phys = self._phys[logical[u]]
 
1028
  logger.info(f"Finished prefill of {N} request(s), starting decode...")
1029
  return torch.cat(out, dim=0), torch.zeros(N, dtype=torch.long)
1030
 
1031
+ def _prefill_planned(self, model, dec, prompts, rows, phys_of, starts, ends, resume_mask, final_mask, seat=True):
1032
+ """One prefill call under the chunk policy: the rows in ChunkedPrefillPlanner order (a resume row first; park
1033
+ before any row that resets the scratch while a partial is held unparked, unpark before the resume), each through
1034
+ the eager spec prefill. A final row returns its host logits [1, 1, vocab] and (``seat``) marks the slot for
1035
+ begin() at its first decode step; an intermediate row returns zero logits, writes no slot row and becomes the
1036
+ scratch owner. The planner's owner is committed only after every row succeeded (tt/chunked_prefill.py).
1037
+ Returns the logits in call order."""
1038
+ N = len(prompts)
1039
+ planner = model._chunked_prefill_planner()
1040
+ first_blocks = [int(r.reshape(-1)[0]) for r in rows]
1041
+ plans, owner_after = planner.plan(starts, ends, resume_mask, final_mask, first_blocks)
1042
+ out = [None] * N
1043
+ owner_phys = self._cp_owner_phys
1044
+ done = []
1045
+ for p in plans:
1046
+ u = p.row
1047
+ phys = int(phys_of[u])
1048
+ t0 = None
1049
+ if _CP_TIMING:
1050
+ ttnn.synchronize_device(model.mesh_device)
1051
+ t0 = time.perf_counter()
1052
+ if p.park_before:
1053
+ model._park_gdn_scratch()
1054
+ if p.unpark_before:
1055
+ model._unpark_gdn_scratch()
1056
+ t1 = None
1057
+ if _CP_TIMING:
1058
+ ttnn.synchronize_device(model.mesh_device)
1059
+ t1 = time.perf_counter()
1060
+ if p.resume:
1061
+ if phys != self._cp_owner_phys:
1062
+ raise RuntimeError(
1063
+ f"Qwen36DFlash chunked prefill: resume row {u} is on slot {phys}, the partial prompt is in slot "
1064
+ f"{self._cp_owner_phys} (the plugin must keep a continuation on its state slot)"
1065
+ )
1066
+ if dec.active[phys] or self._pending[phys] is not None:
1067
+ raise RuntimeError(
1068
+ f"Qwen36DFlash chunked prefill: resume slot {phys} has a live or pending session"
1069
+ )
1070
+ else:
1071
+ if dec.active[phys]:
1072
+ # vLLM released the previous occupant before reusing its slot; close its session if not.
1073
+ dec.end(phys)
1074
+ self._carry[phys], self._stopped[phys], self._prev_tail[phys] = [], False, None
1075
+ self._pending[phys] = None
1076
+ if owner_phys == phys:
1077
+ owner_phys = None # a new prompt reuses the stale partial's slot (the planner parked its state)
1078
+ T = int(p.end)
1079
+ if p.final:
1080
+ logger.info(
1081
+ f"Prefilling slot {phys} up to {T} tokens (TP eager spec prefill"
1082
+ + (f", resumed at {p.start})" if p.resume else ")")
1083
+ )
1084
+ else:
1085
+ logger.info(
1086
+ f"Prefilling slot {phys} tokens [{p.start}, {T}) (TP eager spec prefill, intermediate chunk)"
1087
+ )
1088
+ lt = self._spec_prefill(model, dec, phys, prompts[u][:, :T], T, rows[u], start=p.start, final=p.final)
1089
+ if p.final:
1090
+ out[u] = lt.view(1, 1, -1)
1091
+ if seat:
1092
+ self._pending[phys] = (T, rows[u])
1093
+ if p.resume:
1094
+ owner_phys = None
1095
+ else:
1096
+ out[u] = torch.zeros(1, 1, model.vocab_size, dtype=torch.float32)
1097
+ owner_phys = phys
1098
+ if _CP_TIMING:
1099
+ ttnn.synchronize_device(model.mesh_device)
1100
+ t2 = time.perf_counter()
1101
+ logger.info(
1102
+ f"[DFLASH_CP] phys={phys} start={p.start} end={T} resume={int(p.resume)} final={int(p.final)} "
1103
+ f"park={int(p.park_before)} unpark={int(p.unpark_before)} park_ms={(t1 - t0) * 1e3:.1f} "
1104
+ f"prefill_ms={(t2 - t1) * 1e3:.1f}"
1105
+ )
1106
+ done.append((phys, p.start, T, int(p.resume), int(p.final), int(p.park_before), int(p.unpark_before)))
1107
+ planner.owner = owner_after
1108
+ self._cp_owner_phys = owner_phys if owner_after is not None else None
1109
+ logger.info(f"Qwen36DFlash chunked rows [(phys, start, end, resume, final, park, unpark)]: {done}")
1110
+ return out
1111
+
1112
+ def _cp_warm_sequence(self, model, dec, rows, S0, tag):
1113
+ """Exercise every program of the chunked path (park / unpark, a resumed chunk with and without logits, a
1114
+ tail-only resume, an intermediate chunk's no-logits exit) with the chunk policy's own orchestration: slot 0
1115
+ takes a long prompt in chunks, slot 1 a short rider between them. Leaves no seated slot and no scratch owner."""
1116
+ L0 = _PREFILL_CHUNK + S0
1117
+ L1 = 2 * _PREFILL_CHUNK
1118
+ p0 = dummy_prompt(max(L0, L1), seed=4242)
1119
+ rider = dummy_prompt(S0, seed=4343)
1120
+ seqs = [
1121
+ # [0, C) intermediate; a rider (parks the partial); [C, C+S0) final: tail-only resume after an unpark
1122
+ ([p0], [rows[0]], [0], [0], [_PREFILL_CHUNK], [False], [False]),
1123
+ ([rider], [rows[1]], [1], [0], [S0], [False], [True]),
1124
+ ([p0], [rows[0]], [0], [_PREFILL_CHUNK], [L0], [True], [True]),
1125
+ # [0, C) intermediate; [C, 2C) final together with a rider: exact-multiple resume (logits from the chunk)
1126
+ ([p0], [rows[0]], [0], [0], [_PREFILL_CHUNK], [False], [False]),
1127
+ ([p0, rider], [rows[0], rows[1]], [0, 1], [_PREFILL_CHUNK, 0], [L1, S0], [True, False], [True, True]),
1128
+ ]
1129
+ for prompts, rws, phys, st, en, rs, fn in seqs:
1130
+ self._prefill_planned(model, dec, prompts, rws, phys, st, en, rs, fn, seat=False)
1131
+ ttnn.synchronize_device(model.mesh_device)
1132
+ owner = model._chunked_prefill_planner().owner
1133
+ if owner is not None or self._cp_owner_phys is not None or any(p is not None for p in self._pending):
1134
+ raise RuntimeError(
1135
+ f"Qwen36DFlash chunked-prefill {tag}: left scratch owner {owner} / slot {self._cp_owner_phys} / "
1136
+ f"pending {self._pending}"
1137
+ )
1138
+
1139
  # ------------------------------------------------------------------ decode: one block per step, all live slots
1140
  def _no_device_sampler(self):
1141
  """True when no model on the mesh has the ttnn device sampler (TP=2: 124,160 logits/device > its 64K cap)."""
 
1143
 
1144
  def decode_forward(self, *args, **kwargs):
1145
  if _W <= 1 or not self._spec_ready():
1146
+ # Same reachability as in prefill_forward (a re-warm-up after release_persistent_capture); the plain decode
1147
+ # trace replay lives in the generator base class and has no model-side guard of its own.
1148
+ if self._forbid_plain:
1149
+ raise RuntimeError("Qwen36DFlash: plain decode after the speculative traces were captured (hang class)")
1150
  if _W > 1 and kwargs.get("sampling_params") is not None and self._no_device_sampler():
1151
  # Pre-arm plain step on a mesh without a device sampler: host logits (the runner samples).
1152
  kwargs = dict(kwargs, sampling_params=None)
 
1184
  # user), so the begins below land in the geometry in force; a mis-ordering is a decoder assert, never a
1185
  # wrong-geometry seed. Single bucket: a host no-op.
1186
  plan = getattr(dec, "plan", None)
1187
+ if _STEP_LOG:
1188
+ ttnn.synchronize_device(self.model[0].mesh_device)
1189
+ _ts = [time.perf_counter()]
1190
+ _bucket_before = getattr(dec, "cur_id", None)
1191
  if plan is not None:
1192
  live_after = {phys for phys in range(B) if dec.active[phys]}
1193
  for i in range(Bp):
 
1197
  if bucket != self._last_bucket:
1198
  logger.info(f"Qwen36DFlash: verify bucket {bucket} in force ({len(live_after)} live slot(s))")
1199
  self._last_bucket = bucket
1200
+ if _STEP_LOG:
1201
+ ttnn.synchronize_device(self.model[0].mesh_device)
1202
+ _ts.append(time.perf_counter())
1203
+ _nbegin = sum(1 for i in range(Bp) if int(poss[i]) >= 0 and self._pending[self._phys[i]] is not None)
1204
  live_rows = []
1205
  nosession_rows = [] # live rows without a session: EOS so the request ends (ragged: [EOS, -1, ...])
1206
  for i in range(Bp):
 
1235
  live_rows.append((i, phys))
1236
  t0 = time.perf_counter() if _DEBUG else 0.0
1237
  if _RAGGED:
1238
+ if _STEP_LOG:
1239
+ ttnn.synchronize_device(self.model[0].mesh_device)
1240
+ _ts.append(time.perf_counter())
1241
  out = self._decode_ragged(dec, live_rows, Bp, t0)
1242
  for i in nosession_rows:
1243
  out[i, 0] = self._eos_fill
1244
+ if _STEP_LOG:
1245
+ ttnn.synchronize_device(self.model[0].mesh_device)
1246
+ _ts.append(time.perf_counter())
1247
+ logger.info(
1248
+ f"[DFLASH_STEP] rows={len(live_rows)} begins={_nbegin} bucket={_bucket_before}->"
1249
+ f"{getattr(dec, 'cur_id', None)} plan_ms={(_ts[1] - _ts[0]) * 1e3:.1f} "
1250
+ f"begin_ms={(_ts[2] - _ts[1]) * 1e3:.1f} step_ms={(_ts[3] - _ts[2]) * 1e3:.1f} "
1251
+ f"tok={int((out >= 0).sum())}"
1252
+ )
1253
  return out
1254
  # Step until every live row can fill its block. A row whose carry already holds a stop token
1255
  # needs nothing: it emits through the stop this step (latency for a normal request); the
 
1381
  self._spec.end(phys)
1382
  self._pending[phys] = None
1383
  self._carry[phys], self._stopped[phys], self._prev_tail[phys] = [], False, None
1384
+ if self._cp_owner_phys == phys:
1385
+ # An aborted / preempted partial prompt: its scratch state is dead; drop the ownership so the next prompt
1386
+ # does not park it (a preempted request re-prefills from 0).
1387
+ self._cp_owner_phys = None
1388
+ self.model[0]._chunked_prefill_planner().owner = None
1389
 
1390
  def release_persistent_capture(self) -> None:
1391
  if self._spec is None:
image/blobs/sha256/1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2 ADDED
Binary file (6.62 kB). View file
 
image/blobs/sha256/45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d ADDED
Binary file (115 Bytes). View file
 
image/blobs/sha256/4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9 ADDED
Binary file (4.39 kB). View file
 
image/blobs/sha256/58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2 ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schemaVersion": 2,
3
+ "mediaType": "application/vnd.oci.image.index.v1+json",
4
+ "manifests": [
5
+ {
6
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
7
+ "digest": "sha256:8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146",
8
+ "size": 4501,
9
+ "platform": {
10
+ "architecture": "amd64",
11
+ "os": "linux"
12
+ }
13
+ },
14
+ {
15
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
16
+ "digest": "sha256:aba4bce132b5625823266e9eb71834b271fbc91ad4b8693c3bd34ecf068eb41e",
17
+ "size": 566,
18
+ "annotations": {
19
+ "vnd.docker.reference.digest": "sha256:8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146",
20
+ "vnd.docker.reference.type": "attestation-manifest"
21
+ },
22
+ "platform": {
23
+ "architecture": "unknown",
24
+ "os": "unknown"
25
+ }
26
+ }
27
+ ]
28
+ }
image/blobs/sha256/70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9 ADDED
Binary file (1.52 kB). View file
 
image/blobs/sha256/8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146 ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schemaVersion": 2,
3
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
4
+ "config": {
5
+ "mediaType": "application/vnd.oci.image.config.v1+json",
6
+ "digest": "sha256:ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7",
7
+ "size": 15070
8
+ },
9
+ "layers": [
10
+ {
11
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
12
+ "digest": "sha256:98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67",
13
+ "size": 29751627
14
+ },
15
+ {
16
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
17
+ "digest": "sha256:6973a7aa7572623f5d0581fa621c6a91869112c028a42eb75fd5dad653cc7550",
18
+ "size": 69482149
19
+ },
20
+ {
21
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
22
+ "digest": "sha256:4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9",
23
+ "size": 4393
24
+ },
25
+ {
26
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
27
+ "digest": "sha256:96ec90bda118ca3b0adc835b3984bab76d0463fd2fb8d18304fa14596868c1a9",
28
+ "size": 9565649
29
+ },
30
+ {
31
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
32
+ "digest": "sha256:a092adbd4319ea0e08680de7f4cadc2d2016ed6b7bb27721f985334ec3ddf72a",
33
+ "size": 138948153
34
+ },
35
+ {
36
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
37
+ "digest": "sha256:f37009d803089ea4de72f86b7b2504b33b0e660cd05baa5b5f0b701ab907154f",
38
+ "size": 66293828
39
+ },
40
+ {
41
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
42
+ "digest": "sha256:1de484bf746902ac142693f9c9c810133dec4e4dead0d07da4049c845dfd4b33",
43
+ "size": 1354630576
44
+ },
45
+ {
46
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
47
+ "digest": "sha256:45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d",
48
+ "size": 115
49
+ },
50
+ {
51
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
52
+ "digest": "sha256:dbc3dfaa9d9011e424bbc5046b70645e8c6c1db180dd7e04178c962df6eb25d7",
53
+ "size": 139072443
54
+ },
55
+ {
56
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
57
+ "digest": "sha256:9f6d55da9fcb57e2f47984e8ee74c0e35973bee3b4f7428af5c4242ce4d5dcdb",
58
+ "size": 38513788
59
+ },
60
+ {
61
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
62
+ "digest": "sha256:bc3fc935ead0134d3ab22f25c6fabfa03f66fe05bc0af2c7f422d5aeb8e59374",
63
+ "size": 38513668
64
+ },
65
+ {
66
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
67
+ "digest": "sha256:6f8be5fbe44d20285f20ab7ddf0f118a155ae79f913784a83545e1eff48c1ef7",
68
+ "size": 103118556
69
+ },
70
+ {
71
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
72
+ "digest": "sha256:295aa94991e93759c4dfa8c455c7779e88483d53f19ae0eba70044ddae690081",
73
+ "size": 11906048
74
+ },
75
+ {
76
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
77
+ "digest": "sha256:b7f5500fbf7b7923cfcb92f65307d234f9cccd846060141ed2aa4b3e56328738",
78
+ "size": 1783263
79
+ },
80
+ {
81
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
82
+ "digest": "sha256:1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2",
83
+ "size": 6618
84
+ },
85
+ {
86
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
87
+ "digest": "sha256:1f118d55ecce2afdad0bcb54c6d62e873bcb62f8be7afd631d4e8d59287c3c51",
88
+ "size": 4030916
89
+ },
90
+ {
91
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
92
+ "digest": "sha256:bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7",
93
+ "size": 1076
94
+ },
95
+ {
96
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
97
+ "digest": "sha256:ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9",
98
+ "size": 1379
99
+ },
100
+ {
101
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
102
+ "digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1",
103
+ "size": 32
104
+ },
105
+ {
106
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
107
+ "digest": "sha256:70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9",
108
+ "size": 1520
109
+ },
110
+ {
111
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
112
+ "digest": "sha256:e808130579a129d44f7bb71d6c7b94b11ec446778b2f2004e540f7bf755f12de",
113
+ "size": 773549
114
+ },
115
+ {
116
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
117
+ "digest": "sha256:e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691",
118
+ "size": 4245
119
+ }
120
+ ]
121
+ }
image/blobs/sha256/aba4bce132b5625823266e9eb71834b271fbc91ad4b8693c3bd34ecf068eb41e ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schemaVersion": 2,
3
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
4
+ "config": {
5
+ "mediaType": "application/vnd.oci.image.config.v1+json",
6
+ "digest": "sha256:f08a1fdc23ba07b7b094692072f62ca467a03cbe754bc49130dabddf907da186",
7
+ "size": 167
8
+ },
9
+ "layers": [
10
+ {
11
+ "mediaType": "application/vnd.in-toto+json",
12
+ "digest": "sha256:d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d",
13
+ "size": 1569,
14
+ "annotations": {
15
+ "in-toto.io/predicate-type": "https://slsa.dev/provenance/v0.2"
16
+ }
17
+ }
18
+ ]
19
+ }
image/blobs/sha256/bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7 ADDED
Binary file (1.08 kB). View file
 
image/blobs/sha256/d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d ADDED
@@ -0,0 +1 @@
 
 
1
+ {"_type":"https://in-toto.io/Statement/v0.1","predicateType":"https://slsa.dev/provenance/v0.2","subject":[{"name":"pkg:docker/tt-model/qwen3.8-27b-p150x2@build-fee9e0d35?platform=linux%2Famd64","digest":{"sha256":"8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146"}}],"predicate":{"builder":{"id":""},"buildType":"https://mobyproject.org/buildkit@v1","materials":[{"uri":"pkg:docker/docker/dockerfile@1.7-labs","digest":{"sha256":"b99fecfe00268a8b556fad7d9c37ee25d716ae08a5d7320e6d51c4dd83246894"}},{"uri":"pkg:docker/ubuntu@22.04?platform=linux%2Famd64","digest":{"sha256":"b1066385161d28ddf6bc7e7b28a9170eec11484c821d1a5150d176cbde41d7f7"}},{"uri":"pkg:docker/ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64@latest?platform=linux%2Famd64","digest":{"sha256":"82d91fd88b21f297844406897cee5bd6c9ded6e39447ff6dda698c83b33404d2"}}],"invocation":{"configSource":{"entryPoint":"Dockerfile"},"parameters":{"frontend":"gateway.v0","args":{"cmdline":"docker/dockerfile:1.7-labs","context:metalsrc":"local:metalsrc","frontend.caps":"moby.buildkit.frontend.contexts+forward","source":"docker/dockerfile:1.7-labs"},"locals":[{"name":"context"},{"name":"dockerfile"},{"name":"metalsrc"}]},"environment":{"platform":"linux/amd64"}},"metadata":{"buildInvocationID":"jrcwl84ug3518kcuik2a5bv47","buildStartedOn":"2026-10-03T05:06:49.883974165+09:00","buildFinishedOn":"2026-10-03T05:12:19.066557014+09:00","completeness":{"parameters":false,"environment":true,"materials":false},"reproducible":false,"https://mobyproject.org/buildkit@v1#metadata":{}}}}
image/blobs/sha256/e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691 ADDED
Binary file (4.25 kB). View file
 
image/blobs/sha256/ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9 ADDED
Binary file (1.38 kB). View file
 
image/blobs/sha256/ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7 ADDED
@@ -0,0 +1 @@
 
 
1
+ {"architecture":"amd64","config":{"User":"tt","Env":["PATH=/opt/tt-venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","VENV=/opt/tt-venv","VIRTUAL_ENV=/opt/tt-venv","TT_METAL_RUNTIME_ROOT=/opt/tt-metal","TT_METAL_HOME=/opt/tt-metal","PYTHONPATH=/opt/tt-metal","LD_LIBRARY_PATH=/opt/tt-metal/build/lib:/opt/openmpi-v5.0.7-ulfm/lib","EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle","TT_VLLM_BUILTIN_MODELS=0","TT_MODEL_KIND=vllm-plugin","HF_HOME=/hf","TT_METAL_CACHE=/cache","HOME=/home/tt","USER=tt","LOGNAME=tt"],"Entrypoint":["/usr/local/bin/entrypoint.sh"],"Cmd":["/usr/local/bin/serve-default.sh"],"WorkingDir":"/home/tt/work","Labels":{"org.opencontainers.image.revision":"fee9e0d35948111be29083c4eb144023d86890d8","org.opencontainers.image.version":"22.04","org.tenstorrent.tt-model":"qwen3.8-27b-p150x2","org.tenstorrent.tt-model.arch":"blackhole","org.tenstorrent.tt-model.kind":"vllm-plugin","org.tenstorrent.tt-model.plugin":"ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7","org.tenstorrent.tt-model.profiles":"tp2,p1d1,tp2-dflash2","org.tenstorrent.tt-model.repo":"changh95/qwen3.8-27b-p150x2","org.tenstorrent.tt-model.tt-metal":"v0.79.0-dev20260903-138-gfee9e0d359","org.tenstorrent.tt-model.weights":"Qwen/Qwen3.8-27B"},"ArgsEscaped":true},"created":"2026-10-03T05:11:03.366502764+09:00","history":[{"created":"2026-09-24T22:24:04.688908979Z","created_by":"/bin/sh -c #(nop) ARG RELEASE","empty_layer":true},{"created":"2026-09-24T22:24:04.723527376Z","created_by":"/bin/sh -c #(nop) ARG LAUNCHPAD_BUILD_ARCH","empty_layer":true},{"created":"2026-09-24T22:24:04.751455651Z","created_by":"/bin/sh -c #(nop) LABEL org.opencontainers.image.version=22.04","empty_layer":true},{"created":"2026-09-24T22:24:06.837737736Z","created_by":"/bin/sh -c #(nop) ADD file:b0bf3f64519bf10a51e00d4f8ab9c8693620659a0c4a4a5090a9765e754e1d1c in / "},{"created":"2026-09-24T22:24:07.239071539Z","created_by":"/bin/sh -c #(nop) CMD [\"/bin/bash\"]","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG OMPI_DIR=/opt/openmpi-v5.0.7-ulfm","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG EXTRA_MODELS_DIR=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG TT_MODEL_KIND","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_NAME","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_REPO","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_WEIGHTS","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_ARCH","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_PROFILES","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_TT_METAL_SHA","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_TT_METAL_DESCRIBE","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_PLUGIN_SHA","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 /bin/sh -c apt-get update \u0026\u0026 apt-get install -y --no-install-recommends libhwloc15 libnuma1 libatomic1 libudev1 libcap2 zlib1g libmpc3 libmpfr6 libgmp10 libzstd1 libevent-core-2.1-7 libevent-pthreads-2.1-7 libgl1 libsndfile1 ca-certificates \u0026\u0026 apt-get clean \u0026\u0026 rm -rf /var/lib/apt/lists/* # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:07:17.417705026+09:00","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 /bin/sh -c existing=\"$(getent passwd 1000 | cut -d: -f1)\" \u0026\u0026 if [ -n \"$existing\" ]; then userdel -r \"$existing\" 2\u003e/dev/null || userdel \"$existing\"; fi \u0026\u0026 useradd --uid 1000 --create-home --home-dir /home/tt --shell /bin/bash tt \u0026\u0026 mkdir -p /home/tt/work/logs /cache /opt/tt-metal \u0026\u0026 chown -R tt:tt /home/tt /cache /opt/tt-metal \u0026\u0026 chmod 1777 /home/tt /home/tt/work /home/tt/work/logs /cache # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:27.718820963+09:00","created_by":"COPY /opt/openmpi-v5.0.7-ulfm /opt/openmpi-v5.0.7-ulfm # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:30.575915892+09:00","created_by":"COPY /opt/tenstorrent /opt/tenstorrent # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:31.930224101+09:00","created_by":"COPY /usr/local/share/uv /usr/local/share/uv # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:49.821667975+09:00","created_by":"COPY /opt/tt-venv /opt/tt-venv # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:50.018305462+09:00","created_by":"COPY /opt/vllm /opt/vllm # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:50.833470062+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/runtime /opt/tt-metal/runtime # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:51.090162095+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/build_Release /opt/tt-metal/build_Release # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:51.342451117+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/build /opt/tt-metal/build # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.162323006+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/tt_metal /opt/tt-metal/tt_metal # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.561349473+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/ttnn /opt/tt-metal/ttnn # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.636881027+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/tools /opt/tt-metal/tools # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.679150816+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/setup.py /opt/tt-metal/pyproject.toml /opt/tt-metal/ # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.763642292+09:00","created_by":"COPY --chown=tt:tt code/ /opt/tt-metal/ # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.803147995+09:00","created_by":"COPY entrypoint.sh /usr/local/bin/entrypoint.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"COPY --chmod=0755 serve-default.sh /usr/local/bin/serve-default.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV VENV=/opt/tt-venv","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV VIRTUAL_ENV=/opt/tt-venv","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV PATH=/opt/tt-venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_METAL_RUNTIME_ROOT=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_METAL_HOME=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV PYTHONPATH=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV LD_LIBRARY_PATH=/opt/tt-metal/build/lib:/opt/openmpi-v5.0.7-ulfm/lib","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ARG TT_VLLM_BUILTIN_MODELS=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_VLLM_BUILTIN_MODELS=0","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_MODEL_KIND=vllm-plugin","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV HF_HOME=/hf","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_METAL_CACHE=/cache","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV HOME=/home/tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV USER=tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV LOGNAME=tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"USER tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.857467483+09:00","created_by":"WORKDIR /home/tt/work","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.893157111+09:00","created_by":"COPY verify.sh /ctx/verify.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:11:03.148431143+09:00","created_by":"RUN |12 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 TT_VLLM_BUILTIN_MODELS=0 /bin/sh -c bash /ctx/verify.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:11:03.148431143+09:00","created_by":"USER root","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"RUN |12 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 TT_VLLM_BUILTIN_MODELS=0 /bin/sh -c chmod -R a+rwX /home/tt # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"USER tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"LABEL org.tenstorrent.tt-model=qwen3.8-27b-p150x2 org.tenstorrent.tt-model.repo=changh95/qwen3.8-27b-p150x2 org.tenstorrent.tt-model.weights=Qwen/Qwen3.8-27B org.tenstorrent.tt-model.arch=blackhole org.tenstorrent.tt-model.kind=vllm-plugin org.tenstorrent.tt-model.profiles=tp2,p1d1,tp2-dflash2 org.opencontainers.image.revision=fee9e0d35948111be29083c4eb144023d86890d8 org.tenstorrent.tt-model.tt-metal=v0.79.0-dev20260903-138-gfee9e0d359 org.tenstorrent.tt-model.plugin=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"ENTRYPOINT [\"/usr/local/bin/entrypoint.sh\"]","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"CMD [\"/usr/local/bin/serve-default.sh\"]","comment":"buildkit.dockerfile.v0","empty_layer":true}],"os":"linux","rootfs":{"type":"layers","diff_ids":["sha256:cbaf391e670933f026c05c7dedec3faf8f194b6c2d83b0d5694ad68323c8077c","sha256:cfccc3cbc964861981c069fce1701ffa33f052675c244b7a422825765dd9a63f","sha256:d4f545a332922c93c1dc239e25127c4ae7271a58b97a7ab079d369fd9f4dbc35","sha256:e093e8ec226ffd10480fc5a2bcbedf619b6138801ed0534a40f5cd68955a73d6","sha256:78e44c64fa49347ea5c050c2fdab481b14c4b6dc164a1668c642a97bea1e7685","sha256:2137fc23bb3fe89686694fc573d254802babc0361b3dce51dff7fd83b2b3848c","sha256:22c669305206fddfb75fcec816f7eafc89909aca78e0d47f40040ee4b3216733","sha256:31113ed63efe6fe971ea06eeeda5f10a9e62523bb5fd69dc21221e771ef77739","sha256:8e82e1780c58e70a07294d5525970abd23e780f29a1868d0c27650a435475e31","sha256:2251b873b9d6a96305d4051751caa3f25918ff1c934fc8784a4669dd52c0eff4","sha256:a68c563178b2e86d35fc7b382791a345ad3ed51b0bd5ca5f098109dff40f0ddf","sha256:7870fda76a62b7f08224aa010e777daca52828e7dfd77dad81725784043152a3","sha256:05faaf4debc0728ec662239967cc8aa14ce5b905e0fd7abfc0bb2327126d1414","sha256:93ebc8896d0c5c68f7350d99df7a3fcabd59c1b9d77d518b858067b5f08ed26b","sha256:b73fff02849ac52fe40f953a7aaaa80a1b329d491ab5ebb7f139c1d4e58f18ac","sha256:050f9922591a6b58429eead751bfe7f435181720c78f6d5368c905e375e46f03","sha256:210aee8d627bdba2b75c940d0df6e1e66595d23d84a522165bcf6ce0900e4b30","sha256:55183c62ebd7eb1f6563f266b6e9793ecb66e65bbf420a9bb356836d54f951a7","sha256:5f70bf18a086007016e948b04aed3b82103a36bea41755b6cddfaf10ace3c6ef","sha256:38d87bfaad3b918b0332db9c70eabe949f28d527ed8bdb8ef9c96809ba7a174d","sha256:5eefbd7a03694f21269829dfb4e26281e6fbc6d4dd25688f6bcbd46bb3e0e726","sha256:41b77a65171ea10410caa54ffc892427a7c0adfecb0e3c1ee4cb5bae8e1e9f19"]}}
image/blobs/sha256/f08a1fdc23ba07b7b094692072f62ca467a03cbe754bc49130dabddf907da186 ADDED
@@ -0,0 +1 @@
 
 
1
+ {"architecture":"unknown","os":"unknown","config":{},"rootfs":{"type":"layers","diff_ids":["sha256:d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d"]}}
image/index.json CHANGED
@@ -1 +1 @@
1
- {"schemaVersion":2,"mediaType":"application/vnd.oci.image.index.v1+json","manifests":[{"mediaType":"application/vnd.oci.image.index.v1+json","digest":"sha256:85f009466e2e0c7e5791155bb660685283c06777b22100d4e47be8b255f7188b","size":856,"annotations":{"io.containerd.image.name":"docker.io/tt-model/qwen3.8-27b-p150x2:85f009466e2e","org.opencontainers.image.ref.name":"85f009466e2e"}}]}
 
1
+ {"schemaVersion":2,"mediaType":"application/vnd.oci.image.index.v1+json","manifests":[{"mediaType":"application/vnd.oci.image.index.v1+json","digest":"sha256:58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2","size":856,"annotations":{"io.containerd.image.name":"docker.io/tt-model/qwen3.8-27b-p150x2:58b3a5c0c045","org.opencontainers.image.ref.name":"58b3a5c0c045"}}]}
image/manifest.json CHANGED
@@ -1 +1 @@
1
- [{"Config":"blobs/sha256/c1cee0e98a22990f0c2d56303ec294b78d49553a845e4cee4c3076271697936f","RepoTags":["tt-model/qwen3.8-27b-p150x2:85f009466e2e"],"Layers":["blobs/sha256/20c3783cc497b5b0df1fc5f92bd64c3d2fbb24057c88692c8fda205b6ea8a2f2","blobs/sha256/29d6aca5d52e707d96a8b31350cec8fd7a385b12a201ff47e2d35ecda1978759","blobs/sha256/ce1de91dc55a512733910cbfe3b4835d3d3d809dc4195830d02307d75a631741","blobs/sha256/33e41fbc4483798384707211700774dadf2c4c981c7feef5972a770a667b1c5d","blobs/sha256/91a700a31080d7ebdfb65c55e1132c2959c093ad556e746be9aa9979ffb95392","blobs/sha256/8c1e7dc0fbf745fe7f209a12d6d7def5083c3148b899a119621c5cade9891aa0","blobs/sha256/3ba876c6dee543cab5fb8e1b618ab139729d3a3d3ec9b37d1918ce9b30d19e01","blobs/sha256/d918e45236eae17057b82dc8421abcc148363473e966331c434db9b198febdb3","blobs/sha256/ac2552c043d43ee56a6dac5332014faf8428d88212ee3dd8011330025f8635d4","blobs/sha256/073961ff7195d0384171aac589d793fab5d69cd059deec41e4b72b2480939633","blobs/sha256/5d097bd8880168351fe4d377b18500b6ddd0861f0d41a1b343e709fb1228a40a","blobs/sha256/de1f733e7dbf0d0e214d7c4e0404288f647239a9432076b9829c8a1ec3d36f36","blobs/sha256/9d72841732444f3ddf6aa6cc5e145f51ff1f30362871278f41a053452ad3c350","blobs/sha256/2a83bcc4b2ab59296a54ce4ba6e65ef50a28b9e216decad6c0c86bce640922aa","blobs/sha256/ed1ca5123f0fff20ff1657f68dec4658928cdf5683819c3f6d51a6c7b2d4615e","blobs/sha256/f855c9e5b07c842db49d8e2c51602cb89c8fed06c49ad3bbdc80c60f60435270","blobs/sha256/cfcf858b614621e601a00167fa438b45a1cdf2374d3e2e6de7f6380e70f45a06","blobs/sha256/c7d3001f30792f9a393e538fff87ecb237095b6f88e357b35b9503db373f1c73","blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1","blobs/sha256/bc2f4e5deb28e0758dbcde2c14a96791cddc1a0e9e5a110ac74366867ce3d10a","blobs/sha256/bbb31a4efb9b8ce28aea2a8782edee6a18197184034338f70f4bd008f27e2b35","blobs/sha256/f166e218d388b90ff8b58e8c9eb53a5748f81196d463ecebd449af09eee2a904"]}]
 
1
+ [{"Config":"blobs/sha256/ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7","RepoTags":["tt-model/qwen3.8-27b-p150x2:58b3a5c0c045"],"Layers":["blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67","blobs/sha256/6973a7aa7572623f5d0581fa621c6a91869112c028a42eb75fd5dad653cc7550","blobs/sha256/4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9","blobs/sha256/96ec90bda118ca3b0adc835b3984bab76d0463fd2fb8d18304fa14596868c1a9","blobs/sha256/a092adbd4319ea0e08680de7f4cadc2d2016ed6b7bb27721f985334ec3ddf72a","blobs/sha256/f37009d803089ea4de72f86b7b2504b33b0e660cd05baa5b5f0b701ab907154f","blobs/sha256/1de484bf746902ac142693f9c9c810133dec4e4dead0d07da4049c845dfd4b33","blobs/sha256/45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d","blobs/sha256/dbc3dfaa9d9011e424bbc5046b70645e8c6c1db180dd7e04178c962df6eb25d7","blobs/sha256/9f6d55da9fcb57e2f47984e8ee74c0e35973bee3b4f7428af5c4242ce4d5dcdb","blobs/sha256/bc3fc935ead0134d3ab22f25c6fabfa03f66fe05bc0af2c7f422d5aeb8e59374","blobs/sha256/6f8be5fbe44d20285f20ab7ddf0f118a155ae79f913784a83545e1eff48c1ef7","blobs/sha256/295aa94991e93759c4dfa8c455c7779e88483d53f19ae0eba70044ddae690081","blobs/sha256/b7f5500fbf7b7923cfcb92f65307d234f9cccd846060141ed2aa4b3e56328738","blobs/sha256/1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2","blobs/sha256/1f118d55ecce2afdad0bcb54c6d62e873bcb62f8be7afd631d4e8d59287c3c51","blobs/sha256/bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7","blobs/sha256/ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9","blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1","blobs/sha256/70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9","blobs/sha256/e808130579a129d44f7bb71d6c7b94b11ec446778b2f2004e540f7bf755f12de","blobs/sha256/e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691"]}]
requirements.lock CHANGED
@@ -4,26 +4,26 @@ aiohttp==3.14.3
4
  aiosignal==1.4.0
5
  annotated-doc==0.0.5
6
  annotated-types==0.8.0
7
- anthropic==1.8.0
8
  anyio==4.15.1
9
  apache-tvm-ffi==0.1.10
10
  astor==0.8.1
11
  attrs==26.1.0
12
- blake3==1.0.9
13
  cachetools==7.2.0
14
- cbor2==6.1.4
15
  certifi==2026.7.22
16
  cffi==2.1.1
17
  cfgv==3.5.0
18
- charset-normalizer==3.5.1
19
  click==8.5.0
20
  cloudpickle==3.1.2
21
  compressed-tensors==0.17.0
22
  contourpy==1.3.3
23
- cryptography==50.0.1
24
  cuda-bindings==13.4.3
25
  cuda-core==1.2.1
26
- cuda-pathfinder==1.8.2
27
  cuda-python==13.4.1
28
  cuda-tile==1.6.0
29
  cycler==0.12.1
@@ -43,12 +43,12 @@ fastapi-cli==0.0.32
43
  fastapi-cloud-cli==0.26.0
44
  fastar==0.12.0
45
  fastsafetensors==0.4.0
46
- filelock==4.0.3
47
  flashinfer-python==0.6.14
48
- fonttools==4.66.0
49
  frozenlist==1.8.0
50
  fsspec==2026.9.0
51
- googleapis-common-protos==1.75.4
52
  graphviz==0.21
53
  grpcio==1.84.0
54
  h11==0.16.0
@@ -60,7 +60,7 @@ httpx==0.28.1
60
  httpx2==2.13.1
61
  huggingface_hub==1.33.0
62
  humming-kernels==0.1.10
63
- identify==2.6.19
64
  idna==3.20
65
  ijson==3.5.1
66
  interegular==0.3.3
@@ -87,11 +87,11 @@ mistral_common==1.12.0
87
  ml_dtypes==0.5.4
88
  model-hosting-container-standards==0.1.16
89
  mpmath==1.3.0
90
- msgspec==0.21.1
91
  multidict==6.9.1
92
  networkx==3.7
93
  ninja==1.13.2
94
- nodeenv==1.10.0
95
  numba==0.65.0
96
  numpy==1.26.4
97
  nvidia-cuda-cccl==13.3.4.3.1
@@ -109,7 +109,7 @@ nvidia-cutlass-dsl-libs-cu13==4.6.0
109
  nvidia-ml-py==13.615.71
110
  nvidia-nvvm==13.4.92
111
  nvtx==0.2.15
112
- openai==3.19.2
113
  openai-harmony==0.0.8
114
  opencv-python-headless==4.11.0.86
115
  opentelemetry-api==1.45.0
@@ -128,7 +128,7 @@ packaging==26.3
128
  pandas==3.0.6
129
  partial-json-parser==0.2.1.1.post7
130
  pillow==12.3.0
131
- platformdirs==4.11.13
132
  pre_commit==4.6.2
133
  prometheus-fastapi-instrumentator==8.1.0
134
  prometheus_client==0.26.0
@@ -145,20 +145,20 @@ pydantic-settings==2.15.0
145
  pydantic_core==2.46.5
146
  pyelftools==0.33
147
  Pygments==2.21.0
148
- PyJWT==2.15.0
149
  pyluwen==0.9.0
150
  pynvvideocodec==2.0.4
151
  pyparsing==3.3.3
152
  python-dateutil==2.9.0.post0
153
  python-discovery==1.6.1
154
- python-dotenv==1.2.3
155
  python-json-logger==4.2.0
156
  python-multipart==0.0.32
157
  PyYAML==6.0.3
158
  pyzmq==27.2.0
159
  quack-kernels==0.6.3
160
  referencing==0.37.0
161
- regex==2026.9.10
162
  requests==2.34.2
163
  rich==15.0.0
164
  rich-toolkit==0.20.5
@@ -167,14 +167,14 @@ rpds-py==2026.6.3
167
  safetensors==0.8.0
168
  seaborn==0.13.2
169
  sentencepiece==0.2.2
170
- sentry-sdk==2.70.0
171
- setproctitle==1.3.7
172
  setuptools==80.10.2
173
  setuptools-scm==8.1.0
174
  shellingham==1.5.4
175
  six==1.17.0
176
  sniffio==1.3.1
177
- sse-starlette==3.4.11
178
  starlette==1.7.0
179
  supervisor==4.3.0
180
  sympy==1.14.0
@@ -189,26 +189,22 @@ tokenspeed-triton==3.8.10.post20260920
189
  tomli==2.4.1
190
  torch==2.11.0+cpu
191
  torch_c_dlpack_ext==0.1.5
192
- torchcodec==0.16.0+cpu
193
  torchvision==0.26.0+cpu
194
  tqdm==4.70.1
195
- transformers==5.17.0
196
  triton==3.8.0
197
  truststore==0.10.4
198
  tt-smi==6.6.0
199
  tt-tools-common==1.6.0
200
  tt-umd==0.9.11
201
- ttnn==0.65.2.dev9815
202
- ttnn==0.75.0rc10.dev1229+gf6deef232f7
203
  typer==0.27.2
204
  typing-inspection==0.4.4
205
  typing_extensions==4.16.0
206
  urllib3==2.8.0
207
  uvicorn==0.54.0
208
- uvloop==0.22.1
209
- virtualenv==21.12.1
210
- vllm==0.26.0+empty
211
- vllm-tt-plugin==0.1.0
212
  watchfiles==1.3.0
213
  websockets==17.1
214
  wheel==0.48.0
 
4
  aiosignal==1.4.0
5
  annotated-doc==0.0.5
6
  annotated-types==0.8.0
7
+ anthropic==1.11.0
8
  anyio==4.15.1
9
  apache-tvm-ffi==0.1.10
10
  astor==0.8.1
11
  attrs==26.1.0
12
+ blake3==1.0.10
13
  cachetools==7.2.0
14
+ cbor2==6.1.5
15
  certifi==2026.7.22
16
  cffi==2.1.1
17
  cfgv==3.5.0
18
+ charset-normalizer==3.5.2
19
  click==8.5.0
20
  cloudpickle==3.1.2
21
  compressed-tensors==0.17.0
22
  contourpy==1.3.3
23
+ cryptography==50.0.2
24
  cuda-bindings==13.4.3
25
  cuda-core==1.2.1
26
+ cuda-pathfinder==1.8.3
27
  cuda-python==13.4.1
28
  cuda-tile==1.6.0
29
  cycler==0.12.1
 
43
  fastapi-cloud-cli==0.26.0
44
  fastar==0.12.0
45
  fastsafetensors==0.4.0
46
+ filelock==4.0.9
47
  flashinfer-python==0.6.14
48
+ fonttools==4.66.1
49
  frozenlist==1.8.0
50
  fsspec==2026.9.0
51
+ googleapis-common-protos==1.75.5
52
  graphviz==0.21
53
  grpcio==1.84.0
54
  h11==0.16.0
 
60
  httpx2==2.13.1
61
  huggingface_hub==1.33.0
62
  humming-kernels==0.1.10
63
+ identify==2.6.20
64
  idna==3.20
65
  ijson==3.5.1
66
  interegular==0.3.3
 
87
  ml_dtypes==0.5.4
88
  model-hosting-container-standards==0.1.16
89
  mpmath==1.3.0
90
+ msgspec==0.22.0
91
  multidict==6.9.1
92
  networkx==3.7
93
  ninja==1.13.2
94
+ nodeenv==1.11.0
95
  numba==0.65.0
96
  numpy==1.26.4
97
  nvidia-cuda-cccl==13.3.4.3.1
 
109
  nvidia-ml-py==13.615.71
110
  nvidia-nvvm==13.4.92
111
  nvtx==0.2.15
112
+ openai==3.24.0
113
  openai-harmony==0.0.8
114
  opencv-python-headless==4.11.0.86
115
  opentelemetry-api==1.45.0
 
128
  pandas==3.0.6
129
  partial-json-parser==0.2.1.1.post7
130
  pillow==12.3.0
131
+ platformdirs==4.12.2
132
  pre_commit==4.6.2
133
  prometheus-fastapi-instrumentator==8.1.0
134
  prometheus_client==0.26.0
 
145
  pydantic_core==2.46.5
146
  pyelftools==0.33
147
  Pygments==2.21.0
148
+ PyJWT==2.15.1
149
  pyluwen==0.9.0
150
  pynvvideocodec==2.0.4
151
  pyparsing==3.3.3
152
  python-dateutil==2.9.0.post0
153
  python-discovery==1.6.1
154
+ python-dotenv==1.2.4
155
  python-json-logger==4.2.0
156
  python-multipart==0.0.32
157
  PyYAML==6.0.3
158
  pyzmq==27.2.0
159
  quack-kernels==0.6.3
160
  referencing==0.37.0
161
+ regex==2026.9.29
162
  requests==2.34.2
163
  rich==15.0.0
164
  rich-toolkit==0.20.5
 
167
  safetensors==0.8.0
168
  seaborn==0.13.2
169
  sentencepiece==0.2.2
170
+ sentry-sdk==2.71.0
171
+ setproctitle==1.3.8
172
  setuptools==80.10.2
173
  setuptools-scm==8.1.0
174
  shellingham==1.5.4
175
  six==1.17.0
176
  sniffio==1.3.1
177
+ sse-starlette==3.5.0
178
  starlette==1.7.0
179
  supervisor==4.3.0
180
  sympy==1.14.0
 
189
  tomli==2.4.1
190
  torch==2.11.0+cpu
191
  torch_c_dlpack_ext==0.1.5
192
+ torchcodec==0.17.0+cpu
193
  torchvision==0.26.0+cpu
194
  tqdm==4.70.1
195
+ transformers==5.18.0
196
  triton==3.8.0
197
  truststore==0.10.4
198
  tt-smi==6.6.0
199
  tt-tools-common==1.6.0
200
  tt-umd==0.9.11
 
 
201
  typer==0.27.2
202
  typing-inspection==0.4.4
203
  typing_extensions==4.16.0
204
  urllib3==2.8.0
205
  uvicorn==0.54.0
206
+ uvloop==0.23.0
207
+ virtualenv==21.14.5
 
 
208
  watchfiles==1.3.0
209
  websockets==17.1
210
  wheel==0.48.0
tt_kernel_manifest.json CHANGED
@@ -1,12 +1,12 @@
1
  {
2
  "schema_version": "5.1",
3
  "name": "qwen3.8-27b-p150x2",
4
- "tt_metal_version": "0.65.2.dev9815",
5
  "arch": "blackhole",
6
  "device_count": 2,
7
  "producer": {
8
  "tt_kernel_version": "0.1.0",
9
- "created_at": "2026-09-25T17:06:41.155149+00:00",
10
  "hostname": "tt-quietbox"
11
  },
12
  "weights": {
@@ -27,8 +27,8 @@
27
  "image": {
28
  "registry": "hf",
29
  "repository": "qwen3.8-27b-p150x2",
30
- "tag": "tt-model/qwen3.8-27b-p150x2:85f009466e2e",
31
- "digest": "sha256:85f009466e2e0c7e5791155bb660685283c06777b22100d4e47be8b255f7188b"
32
  },
33
  "kind": "vllm-plugin",
34
  "runtime": {
@@ -211,26 +211,26 @@
211
  "from pathlib import Path; assert Path('/opt/tt-metal/models/model_trace_region_sizes.yaml').is_file()"
212
  ],
213
  "built": {
214
- "image": "tt-model/qwen3.8-27b-p150x2:85f009466e2e",
215
  "repo": "changh95/qwen3.8-27b-p150x2",
216
  "tt_model_version": "0.1.0",
217
- "created_at": "2026-09-25T16:58:52+00:00",
218
  "tt_metal": {
219
- "sha": "aa20afaf172ac35370b8d112edc0f771ee6d2571",
220
- "describe": "v0.79.0-dev20260903-120-gaa20afaf17",
221
  "dirty": false,
222
- "scm_version": "0.65.2.dev9815",
223
  "mode": "local",
224
  "remote": "https://github.com/tenstorrent/tt-metal.git",
225
  "branch": "qwen36-p150x2",
226
  "pushed": true
227
  },
228
- "code_sha256": "5510f5995b0222608598a27000f92c6f02f0c357ef5c6c11d00450e5c616db79",
229
  "plugin": {
230
  "sha": "ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7",
231
  "repo": "https://github.com/changh95/vllm-tt-plugin"
232
  },
233
- "image_digest": "sha256:85f009466e2e0c7e5791155bb660685283c06777b22100d4e47be8b255f7188b"
234
  }
235
  }
236
  }
 
1
  {
2
  "schema_version": "5.1",
3
  "name": "qwen3.8-27b-p150x2",
4
+ "tt_metal_version": "0.65.2.dev9856",
5
  "arch": "blackhole",
6
  "device_count": 2,
7
  "producer": {
8
  "tt_kernel_version": "0.1.0",
9
+ "created_at": "2026-10-02T20:12:52.368580+00:00",
10
  "hostname": "tt-quietbox"
11
  },
12
  "weights": {
 
27
  "image": {
28
  "registry": "hf",
29
  "repository": "qwen3.8-27b-p150x2",
30
+ "tag": "tt-model/qwen3.8-27b-p150x2:58b3a5c0c045",
31
+ "digest": "sha256:58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2"
32
  },
33
  "kind": "vllm-plugin",
34
  "runtime": {
 
211
  "from pathlib import Path; assert Path('/opt/tt-metal/models/model_trace_region_sizes.yaml').is_file()"
212
  ],
213
  "built": {
214
+ "image": "tt-model/qwen3.8-27b-p150x2:58b3a5c0c045",
215
  "repo": "changh95/qwen3.8-27b-p150x2",
216
  "tt_model_version": "0.1.0",
217
+ "created_at": "2026-10-02T20:06:49+00:00",
218
  "tt_metal": {
219
+ "sha": "fee9e0d35948111be29083c4eb144023d86890d8",
220
+ "describe": "v0.79.0-dev20260903-138-gfee9e0d359",
221
  "dirty": false,
222
+ "scm_version": "0.65.2.dev9856",
223
  "mode": "local",
224
  "remote": "https://github.com/tenstorrent/tt-metal.git",
225
  "branch": "qwen36-p150x2",
226
  "pushed": true
227
  },
228
+ "code_sha256": "b0c865d431db3b437ab3e681b01835c22f2f116164a2fe2e30c635fd43e0d6b9",
229
  "plugin": {
230
  "sha": "ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7",
231
  "repo": "https://github.com/changh95/vllm-tt-plugin"
232
  },
233
+ "image_digest": "sha256:58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2"
234
  }
235
  }
236
  }