Add files using upload-large-folder tool
Browse files- README.md +31 -33
- code/models/demos/blackhole/qwen36/tt-model-p300x2.yaml +327 -0
- code/models/demos/blackhole/qwen36/tt-model.yaml +27 -28
- code/models/demos/blackhole/qwen36/tt/chunked_prefill.py +114 -0
- code/models/demos/blackhole/qwen36/tt/qwen36_vllm.py +127 -7
- code/models/demos/blackhole/qwen36/tt/qwen36_vllm_dflash.py +340 -22
- image/blobs/sha256/1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2 +0 -0
- image/blobs/sha256/45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d +0 -0
- image/blobs/sha256/4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9 +0 -0
- image/blobs/sha256/58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2 +28 -0
- image/blobs/sha256/70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9 +0 -0
- image/blobs/sha256/8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146 +121 -0
- image/blobs/sha256/aba4bce132b5625823266e9eb71834b271fbc91ad4b8693c3bd34ecf068eb41e +19 -0
- image/blobs/sha256/bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7 +0 -0
- image/blobs/sha256/d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d +1 -0
- image/blobs/sha256/e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691 +0 -0
- image/blobs/sha256/ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9 +0 -0
- image/blobs/sha256/ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7 +1 -0
- image/blobs/sha256/f08a1fdc23ba07b7b094692072f62ca467a03cbe754bc49130dabddf907da186 +1 -0
- image/index.json +1 -1
- image/manifest.json +1 -1
- requirements.lock +24 -28
- tt_kernel_manifest.json +11 -11
README.md
CHANGED
|
@@ -3,16 +3,15 @@ tags:
|
|
| 3 |
- blackhole
|
| 4 |
- p150x2
|
| 5 |
- tt-model-cache
|
| 6 |
-
- tt-model-catalog
|
| 7 |
- tt-model-container
|
| 8 |
- vllm-plugin
|
| 9 |
---
|
| 10 |
|
| 11 |
# qwen3.8-27b-p150x2
|
| 12 |
|
| 13 |
-
[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling
|
| 14 |
|
| 15 |
-
Runs on **p150x2** — see the serve profiles below.
|
| 16 |
|
| 17 |
Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).
|
| 18 |
|
|
@@ -46,30 +45,30 @@ curl -s http://localhost:20000/v1/chat/completions -H 'Content-Type: application
|
|
| 46 |
|
| 47 |
| profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
|
| 48 |
|---|---|---|---|---|---|---|
|
| 49 |
-
| `tp2` (default) | 32 | 64K | 622K tokens | 17.
|
| 50 |
| `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
|
| 51 |
-
| `tp2-dflash2` | 4 | 64K | 262K tokens |
|
| 52 |
|
| 53 |
-
* `tp2` is 1.
|
| 54 |
-
* `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.
|
| 55 |
|
| 56 |
### Performance
|
| 57 |
|
| 58 |
-
Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means.
|
| 59 |
|
| 60 |
#### `tp2`: throughput, t/s (t/s/u)
|
| 61 |
|
| 62 |
| ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
|
| 63 |
|---|---|---|---|---|---|---|
|
| 64 |
-
| 128 / 128 | 17 (17.
|
| 65 |
-
| 1,024 / 128 | 17 (
|
| 66 |
-
| 2,048 / 128 |
|
| 67 |
-
| 4,096 / 128 |
|
| 68 |
-
| 8,192 / 128 | 14 (
|
| 69 |
-
| 16,384 / 128 |
|
| 70 |
-
| 32,768 / 128 |
|
| 71 |
-
| 128 / 1,024 | 18 (17.
|
| 72 |
-
| 8,192 / 1,024 | 17 (
|
| 73 |
|
| 74 |
`@18`: the KV pool seats 18 prompts of that length.
|
| 75 |
|
|
@@ -77,11 +76,11 @@ Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exac
|
|
| 77 |
|
| 78 |
| ISL / OSL | 1 user | 8 users | 32 users |
|
| 79 |
|---|---|---|---|
|
| 80 |
-
| 128 / 128 |
|
| 81 |
-
| 2,048 / 128 |
|
| 82 |
-
| 8,192 / 128 | 1,
|
| 83 |
-
| 32,768 / 128 |
|
| 84 |
-
| 128 / 1,024 |
|
| 85 |
|
| 86 |
#### `p1d1`: throughput, t/s (t/s/u)
|
| 87 |
|
|
@@ -111,15 +110,14 @@ Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exac
|
|
| 111 |
|
| 112 |
| ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
|
| 113 |
|---|---|---|---|---|
|
| 114 |
-
| 128 / 128 | 17.
|
| 115 |
-
| 2,048 / 128 | 16.
|
| 116 |
-
| 8,192 / 128 |
|
| 117 |
-
| 32,768 / 128 |
|
| 118 |
-
| 128 / 1,024 | 17.
|
| 119 |
-
| 8,192 / 1,024 |
|
| 120 |
-
| code (SPEED-Bench, EOS-terminated) | 16.2 | 58.4 | 3.6x | 194 / 13 |
|
| 121 |
|
| 122 |
-
* 4 users:
|
| 123 |
|
| 124 |
### Notes
|
| 125 |
|
|
@@ -143,9 +141,9 @@ The exact sources the image was built from — `code/` in this repo is byte-iden
|
|
| 143 |
|
| 144 |
| component | built from |
|
| 145 |
| --- | --- |
|
| 146 |
-
| tt-metal | [`
|
| 147 |
| vLLM | [`v0.26.0`](https://github.com/vllm-project/vllm/releases/tag/v0.26.0) |
|
| 148 |
| vllm-tt-plugin | [`ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7`](https://github.com/changh95/vllm-tt-plugin/commit/ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7) |
|
| 149 |
-
| `code/` digest | `
|
| 150 |
-
| built | 2026-
|
| 151 |
|
|
|
|
| 3 |
- blackhole
|
| 4 |
- p150x2
|
| 5 |
- tt-model-cache
|
|
|
|
| 6 |
- tt-model-container
|
| 7 |
- vllm-plugin
|
| 8 |
---
|
| 9 |
|
| 10 |
# qwen3.8-27b-p150x2
|
| 11 |
|
| 12 |
+
[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling package for the 4-chip P300x2: [qwen3.8-27b-p300x2](https://huggingface.co/changh95/qwen3.8-27b-p300x2) (profiles `plain`, `batch8-dflash2`, `single-user-dflash2`).
|
| 13 |
|
| 14 |
+
Runs on **p150x2** or **p150x2** or **p150x2** — see the serve profiles below.
|
| 15 |
|
| 16 |
Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1).
|
| 17 |
|
|
|
|
| 45 |
|
| 46 |
| profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
|
| 47 |
|---|---|---|---|---|---|---|
|
| 48 |
+
| `tp2` (default) | 32 | 64K | 622K tokens | 17.5 | 0.2 s / 1.8 s / 6.6 s | speed, long prompts, many users, all sampling options |
|
| 49 |
| `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
|
| 50 |
+
| `tp2-dflash2` | 4 | 64K | 262K tokens | 41.9 | 0.15 s / 1.9 s / 6.8 s | 1-4 greedy users, code, prompts under ~16k tokens |
|
| 51 |
|
| 52 |
+
* `tp2` is 1.3x faster per user than `p1d1` at 1 user and has 3.8x the KV pool; `p1d1` keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where `tp2` at 8 users degrades to 134-255 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens `p1d1` delivers 84-90% of `tp2`'s aggregate throughput.
|
| 53 |
+
* `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.9x at 8k, 1.4x at 32k) and grows with answer length (3.1-3.2x at 1k-token answers).
|
| 54 |
|
| 55 |
### Performance
|
| 56 |
|
| 57 |
+
Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
|
| 58 |
|
| 59 |
#### `tp2`: throughput, t/s (t/s/u)
|
| 60 |
|
| 61 |
| ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
|
| 62 |
|---|---|---|---|---|---|---|
|
| 63 |
+
| 128 / 128 | 17 (17.5) | 33 (16.3) | 58 (14.4) | 100 (12.5) | 162 (10.1) | 245 (7.7) |
|
| 64 |
+
| 1,024 / 128 | 17 (17.1) | 32 (15.8) | 54 (13.5) | 90 (11.2) | 138 (8.6) | 192 (6.0) |
|
| 65 |
+
| 2,048 / 128 | 17 (16.8) | 30 (15.2) | 52 (13.0) | 80 (10.0) | 118 (7.4) | 154 (4.8) |
|
| 66 |
+
| 4,096 / 128 | 16 (15.9) | 28 (13.8) | 45 (11.2) | 64 (8.0) | 84 (5.3) | 103 (3.2) |
|
| 67 |
+
| 8,192 / 128 | 14 (14.3) | 23 (11.6) | 35 (8.7) | 45 (5.6) | 54 (3.4) | 61 (1.9) |
|
| 68 |
+
| 16,384 / 128 | 12 (12.3) | 18 (9.0) | 24 (6.0) | 28 (3.6) | 32 (2.0) | 34 (1.1) |
|
| 69 |
+
| 32,768 / 128 | 9 (9.2) | 12 (6.0) | 14 (3.5) | 15 (1.9) | 16 (1.0) | 16 (0.9) @18 |
|
| 70 |
+
| 128 / 1,024 | 18 (17.7) | 34 (17.1) | 62 (15.6) | 116 (14.5) | 205 (12.8) | 368 (11.5) |
|
| 71 |
+
| 8,192 / 1,024 | 17 (17.1) | 32 (16.0) | 59 (14.8) | 97 (12.1) | 151 (9.4) | 219 (6.8) |
|
| 72 |
|
| 73 |
`@18`: the KV pool seats 18 prompts of that length.
|
| 74 |
|
|
|
|
| 76 |
|
| 77 |
| ISL / OSL | 1 user | 8 users | 32 users |
|
| 78 |
|---|---|---|---|
|
| 79 |
+
| 128 / 128 | 197 / 56 | 1,436 / 70 | 6,272 / 82 |
|
| 80 |
+
| 2,048 / 128 | 501 / 56 | 3,600 / 73 | 15,447 / 88 |
|
| 81 |
+
| 8,192 / 128 | 1,778 / 57 | 12,352 / 81 | 34,181 / 260 |
|
| 82 |
+
| 32,768 / 128 | 6,631 / 58 | 34,573 / 255 | - |
|
| 83 |
+
| 128 / 1,024 | 197 / 56 | 1,444 / 68 | 6,273 / 81 |
|
| 84 |
|
| 85 |
#### `p1d1`: throughput, t/s (t/s/u)
|
| 86 |
|
|
|
|
| 110 |
|
| 111 |
| ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
|
| 112 |
|---|---|---|---|---|
|
| 113 |
+
| 128 / 128 | 17.5 | 41.9 | 2.39x | 149 / 23 |
|
| 114 |
+
| 2,048 / 128 | 16.8 | 33.7 | 2.01x | 517 / 26 |
|
| 115 |
+
| 8,192 / 128 | 14.3 | 26.6 | 1.86x | 1,874 / 23 |
|
| 116 |
+
| 32,768 / 128 | 9.2 | 12.7 | 1.38x | 6,778 / 26 |
|
| 117 |
+
| 128 / 1,024 | 17.7 | 56.8 | 3.21x | 149 / 18 |
|
| 118 |
+
| 8,192 / 1,024 | 17.1 | 52.6 | 3.08x | 1,767 / 17 |
|
|
|
|
| 119 |
|
| 120 |
+
* 4 users: 113 vs 58 t/s aggregate at 128/128 (1.9x), 151 vs 62 at 128/1,024 (2.4x), 47 vs 35 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
|
| 121 |
|
| 122 |
### Notes
|
| 123 |
|
|
|
|
| 141 |
|
| 142 |
| component | built from |
|
| 143 |
| --- | --- |
|
| 144 |
+
| tt-metal | [`fee9e0d35948111be29083c4eb144023d86890d8`](https://github.com/tenstorrent/tt-metal/commit/fee9e0d35948111be29083c4eb144023d86890d8) |
|
| 145 |
| vLLM | [`v0.26.0`](https://github.com/vllm-project/vllm/releases/tag/v0.26.0) |
|
| 146 |
| vllm-tt-plugin | [`ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7`](https://github.com/changh95/vllm-tt-plugin/commit/ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7) |
|
| 147 |
+
| `code/` digest | `b0c865d431db3b43` (sha256, first 16 hex digits) |
|
| 148 |
+
| built | 2026-10-02T20:06:49+00:00 by tt-model 0.1.0 |
|
| 149 |
|
code/models/demos/blackhole/qwen36/tt-model-p300x2.yaml
ADDED
|
@@ -0,0 +1,327 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# tt-model-p300x2.yaml — Qwen/Qwen3.8-27B on a P300x2 (2x P300 boards = 4 Blackhole chips, (1,4) mesh) via vLLM.
|
| 2 |
+
# ONE image, three serve profiles:
|
| 3 |
+
# plain plain batched decode, up to 32 users, full sampling (DEFAULT; was `batch32` of the old plain package)
|
| 4 |
+
# batch8-dflash2 DFlash2 speculative decoding, up to 8 users, two verify buckets, chunked prefill
|
| 5 |
+
# single-user-dflash2 DFlash2 speculative decoding, one user
|
| 6 |
+
# This file replaces both earlier P300x2 manifests: the plain package's models/demos/blackhole/qwen36/tt-model.yaml on
|
| 7 |
+
# branch qwen36-prefill-opt-package (image 0becf4834925, tt-metal 0e988de173f, plugin 751ec33) and
|
| 8 |
+
# tt-model-dflash2-p300x2.yaml (changh95/qwen3.8-27b-dflash2-p300x2, retired: tt-metal 8645cc34677, plugin 96d4416).
|
| 9 |
+
# tt-model.yaml on this branch is the p150x2 package (changh95/qwen3.8-27b-p150x2).
|
| 10 |
+
#
|
| 11 |
+
# tt-model package --container models/demos/blackhole/qwen36/tt-model-p300x2.yaml --out <builds>
|
| 12 |
+
# tt-model serve <builds>/qwen3.8-27b-p300x2/tt_kernel_manifest.json [--profile batch8-dflash2]
|
| 13 |
+
# tt-model push <builds>/qwen3.8-27b-p300x2 --private
|
| 14 |
+
schema: "5.1"
|
| 15 |
+
repo: changh95/qwen3.8-27b-p300x2
|
| 16 |
+
name: qwen3.8-27b-p300x2
|
| 17 |
+
weights:
|
| 18 |
+
repo: Qwen/Qwen3.8-27B
|
| 19 |
+
revision: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 # refs/main as validated, 2026-09-12
|
| 20 |
+
kind: vllm-plugin
|
| 21 |
+
arch: blackhole
|
| 22 |
+
|
| 23 |
+
source:
|
| 24 |
+
# Branch qwen36-p300x2-pkg = qwen36-p150x2 (round-5 integration a7a8aeb8d41: TP=2 prefill, tp2/TP=4 chunked prefill,
|
| 25 |
+
# fast GDN remap, long-prompt TTFT) + the SDPA width pin default-off commit + this manifest.
|
| 26 |
+
tt_metal: /home/ttuser/experiments/qwen36_27b/tt-metal-pkg5
|
| 27 |
+
code: # import closure of qwen36_vllm.py + the demo/spec-decode path
|
| 28 |
+
- models/common
|
| 29 |
+
- models/tt_transformers
|
| 30 |
+
- models/demos/blackhole/qwen36 # model, DFlash drafters, demo, tests, vllm_bundle
|
| 31 |
+
- models/demos/qwen3_vl/tt/common.py
|
| 32 |
+
- models/demos/utils
|
| 33 |
+
- models/experimental/gated_attention_gated_deltanet
|
| 34 |
+
- models/perf
|
| 35 |
+
- models/tt_dit/utils/tensor.py
|
| 36 |
+
- models/model_trace_region_sizes.yaml
|
| 37 |
+
- models/model_targets.yaml
|
| 38 |
+
ubuntu: "22.04"
|
| 39 |
+
python: "3.12"
|
| 40 |
+
|
| 41 |
+
runtime:
|
| 42 |
+
vllm: {version: "0.26.0"} # the validated serving venv: vllm 0.26.0 (empty target)
|
| 43 |
+
# vllm-tt-plugin branch pd-disagg: superset of both earlier pins (751ec33 of the plain package, 7ddf2c8 of the DFlash2
|
| 44 |
+
# v13 package): the block-output / batched / ragged contracts the DFlash2 profiles use, the TT chunk policy and the
|
| 45 |
+
# block-output chunked-prefill capability batch8-dflash2 uses, the generic TT_MODEL_CLASS_OVERRIDES the plain profile
|
| 46 |
+
# uses (a67ef80, already in 751ec33), plus PD connector, host sampler and scheduler fixes. Pinned to the pushed commit.
|
| 47 |
+
plugin: {repo: https://github.com/changh95/vllm-tt-plugin, ref: 96d4416d31f79973c9ff10b3b41e15d0b06e3430}
|
| 48 |
+
extra_models_dir: models/demos/blackhole/qwen36/vllm_bundle # main_class Qwen36DFlashForCausalLM (the DFlash2 profiles);
|
| 49 |
+
# `plain` selects Qwen36ForCausalLM via TT_MODEL_CLASS_OVERRIDES
|
| 50 |
+
|
| 51 |
+
serve: # common to every profile; a profile's env MERGES over this (it cannot unset a key), so
|
| 52 |
+
# nothing profile-specific lives here: no QWEN36_DRAFTER / DFLASH_* / QWEN36_MTP keys
|
| 53 |
+
hardware: p300x2
|
| 54 |
+
mesh_device: P150x4 # tt-metal's name for the (1,4) Blackhole mesh a P300x2 opens
|
| 55 |
+
port: 8000
|
| 56 |
+
max_model_len: 262144
|
| 57 |
+
max_num_seqs: 32 # overridden per profile below
|
| 58 |
+
block_size: 64
|
| 59 |
+
capabilities:
|
| 60 |
+
tool_parser: qwen3_coder
|
| 61 |
+
reasoning_parser: qwen3
|
| 62 |
+
additional_config:
|
| 63 |
+
tt:
|
| 64 |
+
fabric_config: FABRIC_1D
|
| 65 |
+
trace_region_size: 1073741824
|
| 66 |
+
l1_small_size: 24576
|
| 67 |
+
sample_on_device_mode: decode_only # plain: the ttnn device sampler (1x4); DFlash2: REQUIRED by the block-output contract
|
| 68 |
+
env:
|
| 69 |
+
ARCH_NAME: blackhole
|
| 70 |
+
TT_QWEN35_TEXT_VER: qwen36_blackhole
|
| 71 |
+
TT_MESH_GRAPH_DESC_PATH: /opt/tt-metal/tt_metal/fabric/mesh_graph_descriptors/p300_x2_mesh_graph_descriptor.textproto
|
| 72 |
+
QWEN36_MAX_TOKENS_ALL_USERS: "525312" # bf8 paged KV pool, every profile (= both earlier packages)
|
| 73 |
+
VLLM_RPC_TIMEOUT: "900000"
|
| 74 |
+
VLLM_CONFIGURE_LOGGING: "1"
|
| 75 |
+
TORCHDYNAMO_DISABLE: "1"
|
| 76 |
+
HF_HUB_OFFLINE: "1"
|
| 77 |
+
args:
|
| 78 |
+
- [--max-num-batched-tokens, "262144"]
|
| 79 |
+
|
| 80 |
+
serve_profiles:
|
| 81 |
+
- name: plain
|
| 82 |
+
description: Plain batched decode, up to 32 concurrent users, full sampling support (formerly `batch32`).
|
| 83 |
+
max_num_seqs: 32
|
| 84 |
+
env:
|
| 85 |
+
# The bundle's main_class is the DFlash class; the plugin registers the checkpoint arch under its TT-prefixed name
|
| 86 |
+
# (platform.py: "TT" + arch, the bundle too), so the override MUST be keyed on TTQwen3_5ForConditionalGeneration to
|
| 87 |
+
# beat it; the raw arch is kept for the plain lookup path (same string as the p150x2 tp2 profile, boot-verified there).
|
| 88 |
+
TT_MODEL_CLASS_OVERRIDES: "TTQwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM,Qwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM"
|
| 89 |
+
QWEN36_MTP: "0" # explicit: the plain class must not load the MTP drafter head / its paged KV cache
|
| 90 |
+
QWEN36_PLAIN_GDN_SLOT_FAST: "1" # fast bit-exact GDN slot hand-off at admission (code default: TP=2 only). Bit-exact at
|
| 91 |
+
# TP=4 (slot test 1x4, B=32, decode between admissions); ~130 ms less per admitted
|
| 92 |
+
# request. Set here, not as a TP=4 code default, so the DFlash2 profiles are unchanged
|
| 93 |
+
# (since 1bec7b836a8 QWEN36_DRAFTER=mtp alone would build it)
|
| 94 |
+
QWEN36_DRAFTER: mtp # carried over from batch32; the plain class never speculates. Kept as a backstop: if the
|
| 95 |
+
# override ever failed to apply, the DFlash class would serve plain (speculation OFF)
|
| 96 |
+
# instead of starting speculative serving at 32 slots
|
| 97 |
+
args:
|
| 98 |
+
- [--max-num-batched-tokens, "262144"]
|
| 99 |
+
- name: single-user-dflash2
|
| 100 |
+
description: One request at a time with lossless DFlash2 speculative decoding (block-output serving).
|
| 101 |
+
max_num_seqs: 1
|
| 102 |
+
env:
|
| 103 |
+
QWEN36_DRAFTER: dflash2
|
| 104 |
+
QWEN36_GDN_SPEC_FUSED: "1" # fused gdn_spec_step verify (M2): verify 49.8 -> 37.7 ms at 8 rows, 70 -> 39-41 at 32 rows
|
| 105 |
+
# DFlash2 speculative decoding (incoai/Qwen3.8-27B-DFlash2, block 8 -> 7 drafts per verify, lossless greedy). The
|
| 106 |
+
# drafter is a second checkpoint the single `weights` pointer cannot carry: it is resolved from the HF cache mounted
|
| 107 |
+
# at /hf, so run `hf download incoai/Qwen3.8-27B-DFlash2 --revision dedf8df68adfb1afeaf7b7480c0a0243108177b4` once
|
| 108 |
+
# before `tt-model serve` (quickstart below).
|
| 109 |
+
DFLASH_WEIGHTS: incoai/Qwen3.8-27B-DFlash2@dedf8df68adfb1afeaf7b7480c0a0243108177b4
|
| 110 |
+
QWEN36_DFLASH_BLOCK: "8"
|
| 111 |
+
QWEN36_DFLASH_MAX_PROMPT: "2048" # prompts up to this take the single masked-bucket prefill; longer ones the
|
| 112 |
+
# eager chunked spec prefill. EVERY text prompt speculates.
|
| 113 |
+
QWEN36_DFLASH_SERVE_BLOCK: "32" # tokens committed per vLLM step on a speculative request
|
| 114 |
+
args:
|
| 115 |
+
- [--max-num-batched-tokens, "262144"]
|
| 116 |
+
- --no-async-scheduling # the plugin's block-output contract is synchronous
|
| 117 |
+
- name: batch8-dflash2
|
| 118 |
+
description: Up to 8 concurrent requests with lossless DFlash2 speculative decoding in two verify buckets — K=7 drafts per step (4x8) while up to 4 requests are live, K=3 (8x4) at 5-8 — switched at runtime with no recapture (batched ragged block-output serving).
|
| 119 |
+
max_num_seqs: 8
|
| 120 |
+
env:
|
| 121 |
+
QWEN36_DRAFTER: dflash2
|
| 122 |
+
QWEN36_GDN_SPEC_FUSED: "1" # fused gdn_spec_step verify (M2): verify 49.8 -> 37.7 ms at 8 rows, 70 -> 39-41 at 32 rows
|
| 123 |
+
DFLASH_WEIGHTS: incoai/Qwen3.8-27B-DFlash2@dedf8df68adfb1afeaf7b7480c0a0243108177b4 # see single-user-dflash2
|
| 124 |
+
QWEN36_DFLASH_BLOCK: "8" # K_max = 7: the 4x8 bucket drafts the checkpoint's whole block; every bucket's K <= BLOCK-1
|
| 125 |
+
QWEN36_DFLASH_MAX_PROMPT: "2048"
|
| 126 |
+
QWEN36_DFLASH_SERVE_BLOCK: "32"
|
| 127 |
+
QWEN36_DFLASH_BUCKETS: "8x4,4x8" # verify geometries B x T (profiles/dual_bucket_spec.json): the decoder verifies in the
|
| 128 |
+
# smallest bucket that seats the live requests and reseeds them across a switch bit-exactly
|
| 129 |
+
QWEN36_DFLASH_RAGGED: "1" # one speculative iteration per decode step: a bucket switch stays inside one step
|
| 130 |
+
QWEN36_DFLASH_FOLD_SEED: "1" # (both verify traces + both buckets' draft/extend traces fit: TRACE region 137 MiB of 1 GiB)
|
| 131 |
+
QWEN36_DFLASH_CHUNKED_PREFILL: "1" # chunked prefill (qwen36_vllm_dflash.py declares supports_chunked_prefill +
|
| 132 |
+
# tt_prefill_chunk_tokens 2048 + tt_block_output_chunked_prefill; refused at
|
| 133 |
+
# TP != 4 and max_num_seqs 1). Evaluated OFF vs ON: profiles/tp4_chunked/EVAL.md
|
| 134 |
+
additional_config:
|
| 135 |
+
tt: # deep-merged over serve.additional_config.tt
|
| 136 |
+
prefill_chunk_tokens: 2048 # chunk size (a multiple of the model's tt_prefill_chunk_tokens and of block_size)
|
| 137 |
+
chunked_prefill_decode_steps: 4 # speculative decode steps between two chunk steps while requests decode
|
| 138 |
+
# Other chunk-policy keys keep their block-output defaults (vllm_tt_plugin/config.py
|
| 139 |
+
# resolve_tt_prefill_chunk_policy): only a prompt with >= 8192 tokens left (chunked_prefill_min_tokens) that
|
| 140 |
+
# arrives while an EARLIER request decodes (chunked_prefill_protect_prior_decoders) is chunked; bursts
|
| 141 |
+
# (burst_longs 2) and idle steps prefill whole, exactly as without chunking.
|
| 142 |
+
args:
|
| 143 |
+
- [--max-num-batched-tokens, "262144"]
|
| 144 |
+
- --no-async-scheduling # the block-output contract and the chunk cadence are synchronous
|
| 145 |
+
default_profile: plain
|
| 146 |
+
|
| 147 |
+
verify:
|
| 148 |
+
# plain: the class the override names imports, and the plugin parses the profile's override string to it (vllm first:
|
| 149 |
+
# vllm_tt_plugin.platform imported cold hits a vllm.platforms circular import)
|
| 150 |
+
- "from models.demos.blackhole.qwen36.tt.qwen36_vllm import Qwen36ForCausalLM; assert Qwen36ForCausalLM"
|
| 151 |
+
- "import os, vllm; os.environ['TT_MODEL_CLASS_OVERRIDES']='TTQwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM,Qwen3_5ForConditionalGeneration=models.demos.blackhole.qwen36.tt.qwen36_vllm:Qwen36ForCausalLM'; from vllm_tt_plugin.platform import _tt_model_class_overrides as f; o=f(); assert o['TTQwen3_5ForConditionalGeneration'].endswith('qwen36_vllm:Qwen36ForCausalLM') and len(o) == 2, o"
|
| 152 |
+
- "import os; os.environ['QWEN36_MTP']='0'; os.environ['QWEN36_DRAFTER']='mtp'; from models.demos.blackhole.qwen36.tt.model_config import mtp_requested_by_env; assert not mtp_requested_by_env()"
|
| 153 |
+
- "import os; os.environ['QWEN36_DRAFTER']='mtp'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; assert not C.model_capabilities['tt_adaptive_block_output'] and C.model_capabilities['output_tokens_per_step'] == 1, dict(C.model_capabilities)"
|
| 154 |
+
# DFlash2 profiles
|
| 155 |
+
- "import os; os.environ['QWEN36_DRAFTER']='dflash2'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; assert C.model_capabilities['tt_adaptive_block_output'] and C.model_capabilities['tt_adaptive_block_batched'] and C.model_capabilities['output_tokens_per_step'] == 32, C.model_capabilities"
|
| 156 |
+
- "from models.demos.blackhole.qwen36.tt.dflash2_serving import DFlash2ServingDecoder; assert DFlash2ServingDecoder"
|
| 157 |
+
- "from models.demos.blackhole.qwen36.tt.dflash2_serving import DFlash2DualBucketDecoder, BucketPlanner; assert DFlash2DualBucketDecoder and BucketPlanner"
|
| 158 |
+
- "import os; os.environ['QWEN36_DRAFTER']='dflash2'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import parse_buckets, check_buckets; assert check_buckets(parse_buckets('8x4,4x8'), 8, 7, True) == ((8, 4), (4, 8))"
|
| 159 |
+
- "import json; m=json.load(open('/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle/qwen36/vllm_metadata.json')); assert m['main_class'].endswith(':Qwen36DFlashForCausalLM'), m"
|
| 160 |
+
- "from vllm_tt_plugin.config import get_tt_block_output_kv_lookahead_tokens, is_tt_adaptive_block_output_model, is_tt_adaptive_block_batched, is_tt_adaptive_block_ragged"
|
| 161 |
+
- "import os; os.environ['QWEN36_DRAFTER']='dflash2'; os.environ['QWEN36_DFLASH_RAGGED']='1'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; assert C.model_capabilities['tt_adaptive_block_batched'] and C.model_capabilities['tt_adaptive_block_ragged']"
|
| 162 |
+
- "from models.demos.blackhole.qwen36.tt.dflash2_decode import DFlash2Decoder; from models.demos.blackhole.qwen36.tt.dflash2_tp import DFlash2DrafterTP; assert DFlash2Decoder and DFlash2DrafterTP"
|
| 163 |
+
# batch8-dflash2 chunked prefill: the model declares the block-output chunk contract only with the knob on (never with it
|
| 164 |
+
# off: single-user-dflash2), and the plugin carries the chunk policy that reads prefill_chunk_tokens / chunked_prefill_decode_steps.
|
| 165 |
+
- "import os; os.environ['QWEN36_DRAFTER']='dflash2'; os.environ['QWEN36_DFLASH_CHUNKED_PREFILL']='1'; from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; c=C.model_capabilities; assert c['supports_chunked_prefill'] and c['tt_block_output_chunked_prefill'] and c['tt_prefill_chunk_tokens'] == 2048, dict(c)"
|
| 166 |
+
- "import os; os.environ['QWEN36_DRAFTER']='dflash2'; os.environ.pop('QWEN36_DFLASH_CHUNKED_PREFILL', None); from models.demos.blackhole.qwen36.tt.qwen36_vllm_dflash import Qwen36DFlashForCausalLM as C; c=C.model_capabilities; assert not c.get('supports_chunked_prefill') and 'tt_block_output_chunked_prefill' not in c, dict(c)"
|
| 167 |
+
- "from vllm_tt_plugin.config import resolve_tt_prefill_chunk_policy, is_tt_block_output_chunked_prefill, get_tt_prefill_chunk_extras"
|
| 168 |
+
# common
|
| 169 |
+
- "import vllm.model_executor.models.qwen3_5"
|
| 170 |
+
- "import transformers.models.qwen3_5"
|
| 171 |
+
- "from pathlib import Path; assert Path('/opt/tt-metal/tt_metal/fabric/mesh_graph_descriptors/p300_x2_mesh_graph_descriptor.textproto').is_file()"
|
| 172 |
+
- "from pathlib import Path; assert Path('/opt/tt-metal/models/model_trace_region_sizes.yaml').is_file()"
|
| 173 |
+
- "from pathlib import Path; assert Path('/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle/qwen36/vllm_metadata.json').is_file()"
|
| 174 |
+
|
| 175 |
+
# Card: rendered by profiles/render_card.py p300x2 <this file> from profiles/card_p300x2_bench.md. Every {{...}} is a
|
| 176 |
+
# placeholder for a number measured on this build (lane B2 benchmarks); fill the source, then re-render.
|
| 177 |
+
card:
|
| 178 |
+
description: |
|
| 179 |
+
[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on a Tenstorrent P300x2 (2× P300 = 4 Blackhole chips, 4-way tensor parallel), 256K context. One image, three profiles: `plain` (default; plain batched decode for up to 32 users with every sampling option), `batch8-dflash2` (lossless [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding for 1 to 8 users, 1.2-1.8× `plain` per user) and `single-user-dflash2` (the fastest single stream: 60 tokens/s at short prompts and 83 with 1k-token answers, vs 32 for `plain`). It replaces [changh95/qwen3.8-27b-dflash2-p300x2](https://huggingface.co/changh95/qwen3.8-27b-dflash2-p300x2), whose profiles now live here.
|
| 180 |
+
quickstart: |
|
| 181 |
+
### Run
|
| 182 |
+
|
| 183 |
+
```bash
|
| 184 |
+
tt-model pull changh95/qwen3.8-27b-p300x2 --with-weights
|
| 185 |
+
tt-model serve changh95/qwen3.8-27b-p300x2 # plain (default)
|
| 186 |
+
tt-model serve changh95/qwen3.8-27b-p300x2 --profile batch8-dflash2 # needs the drafter weights below
|
| 187 |
+
tt-model serve changh95/qwen3.8-27b-p300x2 --profile single-user-dflash2 # needs the drafter weights below
|
| 188 |
+
tt-model profiles changh95/qwen3.8-27b-p300x2 # list the profiles
|
| 189 |
+
```
|
| 190 |
+
|
| 191 |
+
* `batch32` is now called `plain`: same configuration, new name. `--profile batch32` no longer exists; use `--profile plain` or no flag.
|
| 192 |
+
* First boot converts the weights and compiles the kernels (about 7 min); later boots take about 2-3 min.
|
| 193 |
+
|
| 194 |
+
### Drafter weights (DFlash2 profiles only, once)
|
| 195 |
+
|
| 196 |
+
```bash
|
| 197 |
+
hf download incoai/Qwen3.8-27B-DFlash2 --revision dedf8df68adfb1afeaf7b7480c0a0243108177b4
|
| 198 |
+
```
|
| 199 |
+
|
| 200 |
+
The drafter is a second checkpoint; the server reads it from your Hugging Face cache. `plain` does not need it.
|
| 201 |
+
|
| 202 |
+
### Query
|
| 203 |
+
|
| 204 |
+
OpenAI-compatible API on port 20000, model id `Qwen/Qwen3.8-27B`:
|
| 205 |
+
|
| 206 |
+
```bash
|
| 207 |
+
curl -s http://localhost:20000/v1/chat/completions -H 'Content-Type: application/json' \
|
| 208 |
+
-d '{"model": "Qwen/Qwen3.8-27B", "messages": [{"role": "user", "content": "Write a Python function that merges two sorted lists."}], "max_tokens": 512}'
|
| 209 |
+
```
|
| 210 |
+
|
| 211 |
+
* Streaming, tool calls (`qwen3_coder` parser) and reasoning content (`qwen3` parser) work as in vLLM on every profile; `"chat_template_kwargs": {"enable_thinking": false}` turns thinking off.
|
| 212 |
+
* `plain` supports every vLLM sampling option. The DFlash2 profiles return the target model's greedy trajectory: `temperature`/`top_p` affect only the first token; `logprobs`, structured outputs and images are not supported.
|
| 213 |
+
|
| 214 |
+
### Profiles at a glance
|
| 215 |
+
|
| 216 |
+
| profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
|
| 217 |
+
|---|---|---|---|---|---|---|
|
| 218 |
+
| `plain` (default) | 32 | 256K | 525K tokens | 31.6 | 0.2 s / 1.1 s / 4.0 s | many users, long prompts, all sampling options (temperature, top-p/top-k, penalties, logprobs, structured outputs) |
|
| 219 |
+
| `batch8-dflash2` | 8 | 256K | 525K tokens | 57.3 | 0.1 s / 1.1 s / 4.1 s | 1-8 greedy users |
|
| 220 |
+
| `single-user-dflash2` | 1 | 256K | 525K tokens | 60.1 | 0.1 s / 1.1 s / 4.1 s | exactly one greedy user, fastest per token |
|
| 221 |
+
|
| 222 |
+
* `plain` is the old `batch32` profile under a new name. At 1 user `batch8-dflash2` is 1.8x faster per user; at 8 users it still leads in aggregate throughput (271 vs 163 t/s at 128 / 128). Beyond 8 users, or for sampled output, use `plain` (up to 688 t/s at 32 users).
|
| 223 |
+
* The DFlash2 profiles are lossless: each request follows the target model's greedy trajectory (up to bf16 near-ties). Their gain is largest with long answers (1.6x at 1,024-token answers, 1 user) and fades with prompt length (1.19x at 32k tokens).
|
| 224 |
+
|
| 225 |
+
## Benchmarks
|
| 226 |
+
|
| 227 |
+
tt-inference-server `--workflow benchmarks` (standard P300X2 grid: random prompts of exactly ISL tokens, `ignore_eos`, greedy, streaming) on one P300x2, measured 2026-10-01/02 on this image, one profile at a time. `t/s` = aggregate output tokens per second, `(t/s/u)` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Points with several runs (the 1-user rows repeat in every run; `single-user-dflash2` ran three times) show the median.
|
| 228 |
+
|
| 229 |
+
### `plain` (default), t/s (t/s/u)
|
| 230 |
+
|
| 231 |
+
| ISL / OSL | 1 user | 2 users | 4 users | 8 users | 16 users | 32 users |
|
| 232 |
+
|---|---|---|---|---|---|---|
|
| 233 |
+
| 128 / 128 | 31.6 (31.6) | 50.9 (25.4) | 94.0 (23.5) | 163.4 (20.4) | 256.8 (16.1) | 359.0 (11.2) |
|
| 234 |
+
| 1,024 / 128 | 30.8 (30.8) | 48.9 (24.5) | 88.4 (22.1) | 146.1 (18.3) | 218.1 (13.6) | 283.9 (8.9) |
|
| 235 |
+
| 2,048 / 128 | 30.0 (30.0) | 47.2 (23.6) | 82.2 (20.6) | 131.4 (16.4) | 186.9 (11.7) | 235.1 (7.3) |
|
| 236 |
+
| 4,096 / 128 | 28.3 (28.3) | 43.3 (21.6) | 71.0 (17.8) | 104.4 (13.0) | 137.4 (8.6) | 160.8 (5.0) |
|
| 237 |
+
| 8,192 / 128 | 25.4 (25.4) | 36.8 (18.4) | 55.9 (14.0) | 73.5 (9.2) | 88.9 (5.6) | 98.3 (3.1) |
|
| 238 |
+
| 16,384 / 128 | 21.5 (21.5) | 28.9 (14.4) | 38.5 (9.6) | 47.0 (5.9) | 52.6 (3.3) | – |
|
| 239 |
+
| 32,768 / 128 | 16.0 (16.0) | 19.7 (9.8) | 23.4 (5.9) | 25.7 (3.2) | – | – |
|
| 240 |
+
| 65,536 / 128 | 9.8 (9.8) | 11.2 (5.6) | 12.3 (3.1) | 12.5 (1.6) | – | – |
|
| 241 |
+
| 131,072 / 128 | 4.8 (4.8) | 5.3 (2.6) | 5.5 (1.4) | – | – | – |
|
| 242 |
+
| 128 / 1,024 | 32.4 (32.4) | 59.7 (29.9) | 117.0 (29.3) | 224.1 (28.0) | 415.8 (26.0) | 687.7 (21.5) |
|
| 243 |
+
| 8,192 / 1,024 | 31.1 (31.1) | 56.0 (28.0) | 103.5 (25.9) | 179.7 (22.5) | 286.8 (17.9) | 393.8 (12.3) |
|
| 244 |
+
| 10,000 / 1,024 | 30.9 (30.9) | 54.9 (27.4) | 100.1 (25.0) | 170.6 (21.3) | 266.1 (16.6) | 356.1 (11.1) |
|
| 245 |
+
|
| 246 |
+
* A dash = the 525K-token KV pool does not seat that many prompts of that length (the workflow skips the point).
|
| 247 |
+
|
| 248 |
+
### `plain`, TTFT ms / TPOT ms
|
| 249 |
+
|
| 250 |
+
| ISL / OSL | 1 user | 8 users | 32 users |
|
| 251 |
+
|---|---|---|---|
|
| 252 |
+
| 128 / 128 | 157 / 30.7 | 1,183 / 40.0 | 5,019 / 50.3 |
|
| 253 |
+
| 1,024 / 128 | 245 / 30.8 | 1,810 / 40.9 | 7,920 / 51.2 |
|
| 254 |
+
| 2,048 / 128 | 345 / 30.9 | 2,495 / 41.7 | 10,772 / 52.3 |
|
| 255 |
+
| 4,096 / 128 | 587 / 30.9 | 4,216 / 44.1 | 18,361 / 55.9 |
|
| 256 |
+
| 8,192 / 128 | 1,103 / 31.0 | 7,776 / 48.4 | 33,750 / 62.3 |
|
| 257 |
+
| 16,384 / 128 | 1,982 / 31.3 | 14,673 / 56.2 | – |
|
| 258 |
+
| 32,768 / 128 | 4,003 / 31.6 | 30,395 / 74.2 | – |
|
| 259 |
+
| 65,536 / 128 | 9,024 / 32.3 | 53,087 / 224.8 | – |
|
| 260 |
+
| 131,072 / 128 | 22,314 / 33.9 | – | – |
|
| 261 |
+
| 128 / 1,024 | 158 / 30.7 | 1,202 / 34.6 | 5,169 / 41.5 |
|
| 262 |
+
| 8,192 / 1,024 | 1,113 / 31.1 | 7,753 / 37.0 | 33,766 / 48.3 |
|
| 263 |
+
| 10,000 / 1,024 | 1,255 / 31.1 | 9,347 / 37.8 | 36,777 / 54.0 |
|
| 264 |
+
|
| 265 |
+
### `batch8-dflash2`, t/s (t/s/u)
|
| 266 |
+
|
| 267 |
+
| ISL / OSL | 1 user | 2 users | 4 users | 8 users | `plain` 1 user t/s/u |
|
| 268 |
+
|---|---|---|---|---|---|
|
| 269 |
+
| 128 / 128 | 57.3 (57.3) | 100.6 (51.4) | 173.9 (45.8) | 270.5 (34.9) | 31.6 |
|
| 270 |
+
| 1,024 / 128 | 55.7 (55.7) | 101.2 (51.0) | 162.8 (43.5) | 245.6 (31.5) | 30.8 |
|
| 271 |
+
| 2,048 / 128 | 48.7 (48.7) | 87.1 (45.0) | 140.0 (36.0) | 195.1 (24.9) | 30.0 |
|
| 272 |
+
| 4,096 / 128 | 45.0 (45.0) | 73.2 (38.1) | 108.9 (28.1) | 141.9 (18.0) | 28.3 |
|
| 273 |
+
| 8,192 / 128 | 38.4 (38.4) | 54.2 (27.4) | 69.9 (18.7) | 85.1 (10.8) | 25.4 |
|
| 274 |
+
| 16,384 / 128 | 28.2 (28.2) | 37.8 (19.2) | 43.4 (11.4) | 51.9 (6.6) | 21.5 |
|
| 275 |
+
| 32,768 / 128 | 19.0 (19.0) | 23.4 (11.8) | 22.7 (6.2) | 25.5 (3.4) | 16.0 |
|
| 276 |
+
| 65,536 / 128 | 10.7 (10.7) | 12.1 (6.1) | 11.7 (3.1) | 12.3 (1.6) | 9.8 |
|
| 277 |
+
| 131,072 / 128 | 4.9 (4.9) | 5.4 (2.7) | 5.5 (1.4) | – | 4.8 |
|
| 278 |
+
| 128 / 1,024 | 53.2 (53.2) | 88.1 (44.6) | 162.2 (45.2) | 279.0 (39.5) | 32.4 |
|
| 279 |
+
| 8,192 / 1,024 | 43.2 (43.2) | 62.2 (32.7) | 121.1 (35.3) | 219.4 (30.9) | 31.1 |
|
| 280 |
+
| 10,000 / 1,024 | 52.7 (52.7) | 64.8 (36.8) | 153.6 (40.4) | 191.5 (29.4) | 30.9 |
|
| 281 |
+
|
| 282 |
+
* A dash = the KV pool does not seat that many users at that length. At 1 to 4 users the server drafts 7 tokens per step; at 5 to 8, 3 per step. Long prompts (8,192 tokens or more) that arrive while other users decode are prefilled in 2,048-token chunks, which keeps the other users' tokens flowing and raises that prompt's own TTFT.
|
| 283 |
+
|
| 284 |
+
### `batch8-dflash2`, TTFT ms / TPOT ms
|
| 285 |
+
|
| 286 |
+
| ISL / OSL | 1 user | 8 users |
|
| 287 |
+
|---|---|---|
|
| 288 |
+
| 128 / 128 | 138 / 16.5 | 331 / 26.3 |
|
| 289 |
+
| 1,024 / 128 | 190 / 16.6 | 565 / 27.5 |
|
| 290 |
+
| 2,048 / 128 | 288 / 18.4 | 890 / 33.5 |
|
| 291 |
+
| 4,096 / 128 | 550 / 18.0 | 1,840 / 41.6 |
|
| 292 |
+
| 8,192 / 128 | 1,079 / 17.7 | 5,603 / 49.3 |
|
| 293 |
+
| 16,384 / 128 | 2,003 / 20.0 | 10,244 / 73.1 |
|
| 294 |
+
| 32,768 / 128 | 4,085 / 20.9 | 30,933 / 55.6 |
|
| 295 |
+
| 65,536 / 128 | 9,208 / 22.2 | 54,090 / 206.9 |
|
| 296 |
+
| 131,072 / 128 | 22,733 / 27.3 | – |
|
| 297 |
+
| 128 / 1,024 | 139 / 18.7 | 376 / 25.0 |
|
| 298 |
+
| 8,192 / 1,024 | 1,080 / 22.1 | 4,411 / 28.0 |
|
| 299 |
+
| 10,000 / 1,024 | 1,231 / 17.8 | 5,891 / 28.3 |
|
| 300 |
+
|
| 301 |
+
### `single-user-dflash2`, one user
|
| 302 |
+
|
| 303 |
+
| ISL / OSL | TTFT ms | TPOT ms | t/s/u | `plain` t/s/u |
|
| 304 |
+
|---|---|---|---|---|
|
| 305 |
+
| 128 / 128 | 134 | 15.7 | 60.1 | 31.6 |
|
| 306 |
+
| 1,024 / 128 | 184 | 16.1 | 57.6 | 30.8 |
|
| 307 |
+
| 2,048 / 128 | 289 | 17.4 | 51.3 | 30.0 |
|
| 308 |
+
| 4,096 / 128 | 537 | 16.9 | 47.5 | 28.3 |
|
| 309 |
+
| 8,192 / 128 | 1,067 | 16.1 | 41.1 | 25.4 |
|
| 310 |
+
| 16,384 / 128 | 1,991 | 17.4 | 30.5 | 21.5 |
|
| 311 |
+
| 32,768 / 128 | 4,086 | 17.2 | 20.4 | 16.0 |
|
| 312 |
+
| 65,536 / 128 | 9,187 | 18.8 | 11.1 | 9.8 |
|
| 313 |
+
| 131,072 / 128 | 22,620 | 18.5 | 5.1 | 4.8 |
|
| 314 |
+
| 128 / 1,024 | 134 | 12.5 | 79.2 | 32.4 |
|
| 315 |
+
| 8,192 / 1,024 | 1,082 | 11.0 | 83.2 | 31.1 |
|
| 316 |
+
| 10,000 / 1,024 | 1,221 | 13.1 | 70.0 | 30.9 |
|
| 317 |
+
|
| 318 |
+
* Median of three runs per point.
|
| 319 |
+
|
| 320 |
+
### Changes in this release
|
| 321 |
+
|
| 322 |
+
* One package for the P300x2. The DFlash2 profiles moved here from [changh95/qwen3.8-27b-dflash2-p300x2](https://huggingface.co/changh95/qwen3.8-27b-dflash2-p300x2), which is retired (its page points here). The old plain profile `batch32` is now `plain`.
|
| 323 |
+
* `plain` now includes two correctness fixes. Earlier plain images could produce degenerate greedy text on short prompts, and their output could change from the second request on. The traced short-prompt prefill no longer trusts stale page-table rows, and each request's Gated-DeltaNet state is written into its decode slot exactly.
|
| 324 |
+
* `batch8-dflash2` prefills long prompts in chunks. A prompt of 8,192 tokens or more that arrives while other requests decode is prefilled in 2,048-token chunks between their steps, so they keep streaming (first shipped in the last DFlash2 release, 2026-09-30).
|
| 325 |
+
* Faster prefill of long prompts: TTFT at 32k tokens 5,149 -> 4,085 ms on `batch8-dflash2`, 5,077 -> 4,086 ms on `single-user-dflash2` and 4,457 -> 4,003 ms on `plain` (1 user).
|
| 326 |
+
* `plain` throughput: the exact slot write costs time per admitted request, so short prompts at many users are slower than with the previous plain image: 128 / 128 tokens at 32 users 491 -> 359 tokens/s (-27%), at 8 users 188 -> 163 (-13%), TTFT at 128 tokens 94 -> 157 ms. One user decodes at the same speed (32.3 -> 31.6 t/s/u); long answers at 32 users lose 7% (739 -> 688 t/s at 128 / 1,024). The previous image's higher numbers came from a slot write that was not exact.
|
| 327 |
+
* Intended output change: on the DFlash2 profiles, prompts of 2,048 tokens or more now prefill with exactly the numerics of the `plain` path. Greedy output for those prompts can differ from the previous DFlash2 package at bf16 near-ties. Shorter prompts give the same output as before.
|
code/models/demos/blackhole/qwen36/tt-model.yaml
CHANGED
|
@@ -194,7 +194,7 @@ verify:
|
|
| 194 |
# benchmarks land; the text below is the pre-benchmark draft (placeholders in {{...}}).
|
| 195 |
card:
|
| 196 |
description: |
|
| 197 |
-
[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling
|
| 198 |
quickstart: |
|
| 199 |
### Run
|
| 200 |
|
|
@@ -217,30 +217,30 @@ card:
|
|
| 217 |
|
| 218 |
| profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
|
| 219 |
|---|---|---|---|---|---|---|
|
| 220 |
-
| `tp2` (default) | 32 | 64K | 622K tokens | 17.
|
| 221 |
| `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
|
| 222 |
-
| `tp2-dflash2` | 4 | 64K | 262K tokens |
|
| 223 |
|
| 224 |
-
* `tp2` is 1.
|
| 225 |
-
* `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.
|
| 226 |
|
| 227 |
### Performance
|
| 228 |
|
| 229 |
-
Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means.
|
| 230 |
|
| 231 |
#### `tp2`: throughput, t/s (t/s/u)
|
| 232 |
|
| 233 |
| ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
|
| 234 |
|---|---|---|---|---|---|---|
|
| 235 |
-
| 128 / 128 | 17 (17.
|
| 236 |
-
| 1,024 / 128 | 17 (
|
| 237 |
-
| 2,048 / 128 |
|
| 238 |
-
| 4,096 / 128 |
|
| 239 |
-
| 8,192 / 128 | 14 (
|
| 240 |
-
| 16,384 / 128 |
|
| 241 |
-
| 32,768 / 128 |
|
| 242 |
-
| 128 / 1,024 | 18 (17.
|
| 243 |
-
| 8,192 / 1,024 | 17 (
|
| 244 |
|
| 245 |
`@18`: the KV pool seats 18 prompts of that length.
|
| 246 |
|
|
@@ -248,11 +248,11 @@ card:
|
|
| 248 |
|
| 249 |
| ISL / OSL | 1 user | 8 users | 32 users |
|
| 250 |
|---|---|---|---|
|
| 251 |
-
| 128 / 128 |
|
| 252 |
-
| 2,048 / 128 |
|
| 253 |
-
| 8,192 / 128 | 1,
|
| 254 |
-
| 32,768 / 128 |
|
| 255 |
-
| 128 / 1,024 |
|
| 256 |
|
| 257 |
#### `p1d1`: throughput, t/s (t/s/u)
|
| 258 |
|
|
@@ -282,15 +282,14 @@ card:
|
|
| 282 |
|
| 283 |
| ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
|
| 284 |
|---|---|---|---|---|
|
| 285 |
-
| 128 / 128 | 17.
|
| 286 |
-
| 2,048 / 128 | 16.
|
| 287 |
-
| 8,192 / 128 |
|
| 288 |
-
| 32,768 / 128 |
|
| 289 |
-
| 128 / 1,024 | 17.
|
| 290 |
-
| 8,192 / 1,024 |
|
| 291 |
-
| code (SPEED-Bench, EOS-terminated) | 16.2 | 58.4 | 3.6x | 194 / 13 |
|
| 292 |
|
| 293 |
-
* 4 users:
|
| 294 |
|
| 295 |
### Notes
|
| 296 |
|
|
|
|
| 194 |
# benchmarks land; the text below is the pre-benchmark draft (placeholders in {{...}}).
|
| 195 |
card:
|
| 196 |
description: |
|
| 197 |
+
[Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) served with vLLM on two Tenstorrent p150a cards linked by ethernet (2 Blackhole chips, 32 GB each). One image, three profiles: `tp2` (default, tensor-parallel over both chips), `p1d1` (one chip prefills, the other decodes), `tp2-dflash2` (`tp2` + [DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) speculative decoding). Sibling package for the 4-chip P300x2: [qwen3.8-27b-p300x2](https://huggingface.co/changh95/qwen3.8-27b-p300x2) (profiles `plain`, `batch8-dflash2`, `single-user-dflash2`).
|
| 198 |
quickstart: |
|
| 199 |
### Run
|
| 200 |
|
|
|
|
| 217 |
|
| 218 |
| profile | users | context | KV pool | 1 user tok/s | TTFT 128 / 8k / 32k tokens | pick it for |
|
| 219 |
|---|---|---|---|---|---|---|
|
| 220 |
+
| `tp2` (default) | 32 | 64K | 622K tokens | 17.5 | 0.2 s / 1.8 s / 6.6 s | speed, long prompts, many users, all sampling options |
|
| 221 |
| `p1d1` | 8 | 64K | 164K tokens | 13.7 | 0.3 s / 3.8 s / 16 s | steady per-token latency while long prompts keep arriving (decode never pauses for a prefill) |
|
| 222 |
+
| `tp2-dflash2` | 4 | 64K | 262K tokens | 41.9 | 0.15 s / 1.9 s / 6.8 s | 1-4 greedy users, code, prompts under ~16k tokens |
|
| 223 |
|
| 224 |
+
* `tp2` is 1.3x faster per user than `p1d1` at 1 user and has 3.8x the KV pool; `p1d1` keeps TPOT at 71-87 ms at 1-8 users across every prompt length, where `tp2` at 8 users degrades to 134-255 ms once prompts reach 16k-32k tokens. At 8 users on prompts up to 4k tokens `p1d1` delivers 84-90% of `tp2`'s aggregate throughput.
|
| 225 |
+
* `tp2-dflash2` is lossless (greedy trajectory of `tp2` up to bf16 near-ties); its gain fades with prompt length (2.4x at 128 tokens, 1.9x at 8k, 1.4x at 32k) and grows with answer length (3.1-3.2x at 1k-token answers).
|
| 226 |
|
| 227 |
### Performance
|
| 228 |
|
| 229 |
+
Method: tt-inference-server `--workflow benchmarks` grid (random prompts of exactly ISL tokens, OSL output tokens, streaming). `t/s` = aggregate output tokens/s, `t/s/u` = per user (= OSL / mean end-to-end latency); TTFT and TPOT are means. Measured on a P300x2 development box emulating the p150x2 (two chips on different cards over a 2-channel ethernet link), not on a p150a pair.
|
| 230 |
|
| 231 |
#### `tp2`: throughput, t/s (t/s/u)
|
| 232 |
|
| 233 |
| ISL / OSL | 1 user | 2 | 4 | 8 | 16 | 32 |
|
| 234 |
|---|---|---|---|---|---|---|
|
| 235 |
+
| 128 / 128 | 17 (17.5) | 33 (16.3) | 58 (14.4) | 100 (12.5) | 162 (10.1) | 245 (7.7) |
|
| 236 |
+
| 1,024 / 128 | 17 (17.1) | 32 (15.8) | 54 (13.5) | 90 (11.2) | 138 (8.6) | 192 (6.0) |
|
| 237 |
+
| 2,048 / 128 | 17 (16.8) | 30 (15.2) | 52 (13.0) | 80 (10.0) | 118 (7.4) | 154 (4.8) |
|
| 238 |
+
| 4,096 / 128 | 16 (15.9) | 28 (13.8) | 45 (11.2) | 64 (8.0) | 84 (5.3) | 103 (3.2) |
|
| 239 |
+
| 8,192 / 128 | 14 (14.3) | 23 (11.6) | 35 (8.7) | 45 (5.6) | 54 (3.4) | 61 (1.9) |
|
| 240 |
+
| 16,384 / 128 | 12 (12.3) | 18 (9.0) | 24 (6.0) | 28 (3.6) | 32 (2.0) | 34 (1.1) |
|
| 241 |
+
| 32,768 / 128 | 9 (9.2) | 12 (6.0) | 14 (3.5) | 15 (1.9) | 16 (1.0) | 16 (0.9) @18 |
|
| 242 |
+
| 128 / 1,024 | 18 (17.7) | 34 (17.1) | 62 (15.6) | 116 (14.5) | 205 (12.8) | 368 (11.5) |
|
| 243 |
+
| 8,192 / 1,024 | 17 (17.1) | 32 (16.0) | 59 (14.8) | 97 (12.1) | 151 (9.4) | 219 (6.8) |
|
| 244 |
|
| 245 |
`@18`: the KV pool seats 18 prompts of that length.
|
| 246 |
|
|
|
|
| 248 |
|
| 249 |
| ISL / OSL | 1 user | 8 users | 32 users |
|
| 250 |
|---|---|---|---|
|
| 251 |
+
| 128 / 128 | 197 / 56 | 1,436 / 70 | 6,272 / 82 |
|
| 252 |
+
| 2,048 / 128 | 501 / 56 | 3,600 / 73 | 15,447 / 88 |
|
| 253 |
+
| 8,192 / 128 | 1,778 / 57 | 12,352 / 81 | 34,181 / 260 |
|
| 254 |
+
| 32,768 / 128 | 6,631 / 58 | 34,573 / 255 | - |
|
| 255 |
+
| 128 / 1,024 | 197 / 56 | 1,444 / 68 | 6,273 / 81 |
|
| 256 |
|
| 257 |
#### `p1d1`: throughput, t/s (t/s/u)
|
| 258 |
|
|
|
|
| 282 |
|
| 283 |
| ISL / OSL | `tp2` t/s/u | `tp2-dflash2` t/s/u | speed-up | `tp2-dflash2` TTFT / TPOT ms |
|
| 284 |
|---|---|---|---|---|
|
| 285 |
+
| 128 / 128 | 17.5 | 41.9 | 2.39x | 149 / 23 |
|
| 286 |
+
| 2,048 / 128 | 16.8 | 33.7 | 2.01x | 517 / 26 |
|
| 287 |
+
| 8,192 / 128 | 14.3 | 26.6 | 1.86x | 1,874 / 23 |
|
| 288 |
+
| 32,768 / 128 | 9.2 | 12.7 | 1.38x | 6,778 / 26 |
|
| 289 |
+
| 128 / 1,024 | 17.7 | 56.8 | 3.21x | 149 / 18 |
|
| 290 |
+
| 8,192 / 1,024 | 17.1 | 52.6 | 3.08x | 1,767 / 17 |
|
|
|
|
| 291 |
|
| 292 |
+
* 4 users: 113 vs 58 t/s aggregate at 128/128 (1.9x), 151 vs 62 at 128/1,024 (2.4x), 47 vs 35 at 8k/128 (1.3x); no gain at 32k prompts with 4 users.
|
| 293 |
|
| 294 |
### Notes
|
| 295 |
|
code/models/demos/blackhole/qwen36/tt/chunked_prefill.py
ADDED
|
@@ -0,0 +1,114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# SPDX-FileCopyrightText: © 2026 Tenstorrent USA, Inc.
|
| 2 |
+
# SPDX-License-Identifier: Apache-2.0
|
| 3 |
+
"""Scheduler-driven chunked prefill for batched TP serving: the host-side row plan (no ttnn, host-testable).
|
| 4 |
+
|
| 5 |
+
vLLM (with the plugin's chunk policy, QWEN36_CHUNKED_PREFILL=1) may split one long prompt into 2048-aligned chunks that
|
| 6 |
+
run in separate prefill calls, with decode steps and other prompts' prefills in between. The per-sequence GDN state of
|
| 7 |
+
the partial prompt (recurrent state, conv taps, cross-chunk conv carry) lives in the persistent B=1 prefill scratch
|
| 8 |
+
between its calls (its decode slot row is NOT safe storage: the batched decode rewrites idle rows inside the pow2 bucket
|
| 9 |
+
and slot remaps gather every row). Any other prompt prefilled while a partial is in flight resets that scratch, so the
|
| 10 |
+
partial's state is first copied to a same-shape park buffer ("park") and copied back before its next chunk ("unpark").
|
| 11 |
+
|
| 12 |
+
``ChunkedPrefillPlanner.plan`` turns one prefill call's rows into an execution plan and the scratch owner that results:
|
| 13 |
+
|
| 14 |
+
* a row is a RESUME row only when the caller flags it (``resume_mask``; the plugin sets it for a scheduler chunk
|
| 15 |
+
continuation). A resume row must continue the current owner exactly: same first KV block, ``start == next_pos``,
|
| 16 |
+
``start`` a positive multiple of the chunk size. Anything else raises (never silently re-prefill with a wrong state).
|
| 17 |
+
* every other row re-prefills from position 0 (``start`` ignored), which is the pre-chunking behaviour for any row.
|
| 18 |
+
* an intermediate row (``final_mask`` False) must end on a chunk boundary; it becomes the scratch owner, writes no
|
| 19 |
+
decode slot and reads back no logits.
|
| 20 |
+
* the resume row runs first. Before any row that resets the scratch (every non-resume row) an unparked owner is parked;
|
| 21 |
+
before a resume row a parked owner is unparked. So a call that carries only short prompts between two chunks of the
|
| 22 |
+
partial parks it too (review B1), and a stale owner left by an aborted/preempted partial costs one park at most.
|
| 23 |
+
"""
|
| 24 |
+
|
| 25 |
+
from dataclasses import dataclass, replace
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
@dataclass(frozen=True)
|
| 29 |
+
class ScratchOwner:
|
| 30 |
+
"""The partial prompt whose GDN state the B=1 prefill scratch (or its park buffer) holds."""
|
| 31 |
+
|
| 32 |
+
first_block: int # page_table[row, 0] of the request: its identity for the continuation check
|
| 33 |
+
next_pos: int # tokens [0, next_pos) are in the paged KV and in the state; the next chunk must start here
|
| 34 |
+
parked: bool # True: the state is in the park buffer (the scratch was reused since)
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
@dataclass(frozen=True)
|
| 38 |
+
class RowPlan:
|
| 39 |
+
row: int # index into the call's rows (logits are returned in call order)
|
| 40 |
+
start: int # absolute first position this call computes (0 unless resume)
|
| 41 |
+
end: int # exclusive end (the chunk end / prompt length)
|
| 42 |
+
resume: bool
|
| 43 |
+
final: bool # False = intermediate chunk: no slot write, no logits readback
|
| 44 |
+
park_before: bool
|
| 45 |
+
unpark_before: bool
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
class ChunkedPrefillPlanner:
|
| 49 |
+
def __init__(self, chunk_tokens):
|
| 50 |
+
if int(chunk_tokens) <= 0:
|
| 51 |
+
raise ValueError(f"chunk_tokens must be positive, got {chunk_tokens}")
|
| 52 |
+
self.chunk = int(chunk_tokens)
|
| 53 |
+
self.owner = None # ScratchOwner | None
|
| 54 |
+
|
| 55 |
+
def plan(self, starts, ends, resume_mask, final_mask, first_blocks):
|
| 56 |
+
"""Return (list[RowPlan] in execution order, owner after the call). Does not modify ``self.owner``."""
|
| 57 |
+
n = len(ends)
|
| 58 |
+
for name, seq in (("starts", starts), ("resume_mask", resume_mask), ("final_mask", final_mask)):
|
| 59 |
+
if len(seq) != n:
|
| 60 |
+
raise ValueError(f"chunked prefill: {name} has {len(seq)} entries for {n} rows")
|
| 61 |
+
if len(first_blocks) != n:
|
| 62 |
+
raise ValueError(f"chunked prefill: first_blocks has {len(first_blocks)} entries for {n} rows")
|
| 63 |
+
C = self.chunk
|
| 64 |
+
resume_rows = [u for u in range(n) if resume_mask[u]]
|
| 65 |
+
if len(resume_rows) > 1:
|
| 66 |
+
raise ValueError(f"chunked prefill: {len(resume_rows)} resume rows {resume_rows} in one call (max 1)")
|
| 67 |
+
new_partials = [u for u in range(n) if not resume_mask[u] and not final_mask[u]]
|
| 68 |
+
if len(new_partials) + (1 if resume_rows and not final_mask[resume_rows[0]] else 0) > 1:
|
| 69 |
+
raise ValueError(
|
| 70 |
+
f"chunked prefill: more than one partial prompt in one call (resume rows {resume_rows}, "
|
| 71 |
+
f"new intermediate rows {new_partials}); the scheduler admits one long prompt at a time"
|
| 72 |
+
)
|
| 73 |
+
for u in range(n):
|
| 74 |
+
if int(ends[u]) < 1:
|
| 75 |
+
raise ValueError(f"chunked prefill: row {u} has end {ends[u]}")
|
| 76 |
+
if not final_mask[u] and int(ends[u]) % C != 0:
|
| 77 |
+
raise ValueError(
|
| 78 |
+
f"chunked prefill: intermediate row {u} ends at {ends[u]}, not a multiple of the chunk {C}"
|
| 79 |
+
)
|
| 80 |
+
owner = self.owner
|
| 81 |
+
if resume_rows:
|
| 82 |
+
u = resume_rows[0]
|
| 83 |
+
s = int(starts[u])
|
| 84 |
+
if s <= 0 or s % C != 0 or s >= int(ends[u]):
|
| 85 |
+
raise ValueError(
|
| 86 |
+
f"chunked prefill: resume row {u} has start {s}, end {ends[u]}: the start must be a positive "
|
| 87 |
+
f"multiple of {C} below the end"
|
| 88 |
+
)
|
| 89 |
+
if owner is None or owner.first_block != int(first_blocks[u]) or owner.next_pos != s:
|
| 90 |
+
raise ValueError(
|
| 91 |
+
f"chunked prefill: resume row {u} (first_block={int(first_blocks[u])}, start={s}) does not continue "
|
| 92 |
+
f"the scratch owner {owner}"
|
| 93 |
+
)
|
| 94 |
+
order = resume_rows + [u for u in range(n) if u not in resume_rows]
|
| 95 |
+
plans = []
|
| 96 |
+
for u in order:
|
| 97 |
+
resume = bool(resume_mask[u])
|
| 98 |
+
final = bool(final_mask[u])
|
| 99 |
+
park = unpark = False
|
| 100 |
+
if resume:
|
| 101 |
+
unpark = owner.parked
|
| 102 |
+
start = int(starts[u])
|
| 103 |
+
else:
|
| 104 |
+
park = owner is not None and not owner.parked
|
| 105 |
+
if park:
|
| 106 |
+
owner = replace(owner, parked=True)
|
| 107 |
+
start = 0
|
| 108 |
+
plans.append(RowPlan(u, start, int(ends[u]), resume, final, park, unpark))
|
| 109 |
+
if final:
|
| 110 |
+
if resume:
|
| 111 |
+
owner = None
|
| 112 |
+
else:
|
| 113 |
+
owner = ScratchOwner(int(first_blocks[u]), int(ends[u]), parked=False)
|
| 114 |
+
return plans, owner
|
code/models/demos/blackhole/qwen36/tt/qwen36_vllm.py
CHANGED
|
@@ -84,19 +84,51 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 84 |
# so no device-resident token/position chain is ever assumed; the plugin then overlaps only the
|
| 85 |
# engine's scheduling/output work with the device step (its steady-decode fast path needs device sampling).
|
| 86 |
# The gate is read when the platform queries the capability (_ModelCapabilities), not at import time.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 87 |
class _ModelCapabilities(dict):
|
| 88 |
-
"""dict whose ``supports_async_decode``
|
|
|
|
| 89 |
|
| 90 |
_ASYNC = "supports_async_decode"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
|
| 92 |
def __getitem__(self, key):
|
| 93 |
if key == self._ASYNC:
|
| 94 |
return os.environ.get("QWEN36_ASYNC_DECODE_OK", "0") == "1"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 95 |
return super().__getitem__(key)
|
| 96 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 97 |
def get(self, key, default=None):
|
| 98 |
if key == self._ASYNC:
|
| 99 |
return self[key]
|
|
|
|
|
|
|
| 100 |
return super().get(key, default)
|
| 101 |
|
| 102 |
model_capabilities = _ModelCapabilities(
|
|
@@ -351,7 +383,26 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 351 |
"batched (max_num_seqs>1) serving is text-only; multimodal is single-sequence "
|
| 352 |
"(max_concurrency=1). Run the model at max_num_seqs=1 for image/video requests."
|
| 353 |
)
|
| 354 |
-
return self._prefill_forward_tp_batched(
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 355 |
vision_tokens = self._compute_vision_tokens(model, kwargs)
|
| 356 |
if model.use_tp:
|
| 357 |
return self._prefill_forward_tp(model, tokens, page_table, prompt_lens, vision_tokens=vision_tokens)
|
|
@@ -417,7 +468,9 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 417 |
logger.info(f"Finished prefill up to {T} tokens, starting decode...")
|
| 418 |
return logits, torch.zeros(1, dtype=torch.long)
|
| 419 |
|
| 420 |
-
def _prefill_forward_tp_batched(
|
|
|
|
|
|
|
| 421 |
"""TP batched (max_num_seqs>1) prefill: prefill each request in this step into its decode slot.
|
| 422 |
|
| 423 |
vLLM prefills new requests while other slots decode, so each user's B=1 state is written into
|
|
@@ -428,6 +481,10 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 428 |
page_table: torch [N, max_blocks] — row u = request u's blocks.
|
| 429 |
prompt_lens: per-request real lengths (row u trimmed to prompt_lens[u]).
|
| 430 |
empty_slots: per-request decode slot; defaults to range(N) (mirrors Generator.prefill_forward_text).
|
|
|
|
|
|
|
|
|
|
|
|
|
| 431 |
Returns ([N, 1, vocab] host logits, [N] zero rope_deltas — text M-RoPE delta is 0, applied model-side).
|
| 432 |
"""
|
| 433 |
N = tokens.shape[0]
|
|
@@ -437,8 +494,27 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 437 |
empty_slots = [int(s) for s in empty_slots]
|
| 438 |
token_ids_list = [tokens[u : u + 1, : plens[u]].to(torch.int32) for u in range(N)]
|
| 439 |
pt = page_table if isinstance(page_table, torch.Tensor) else ttnn.to_torch(page_table)
|
| 440 |
-
|
| 441 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 442 |
logits = torch.cat([hl.reshape(1, 1, -1) for hl in host_logits], dim=0) # [N, 1, vocab]
|
| 443 |
logger.info(f"Finished batched prefill of {N} user(s), starting decode...")
|
| 444 |
return logits, torch.zeros(N, dtype=torch.long)
|
|
@@ -454,10 +530,26 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 454 |
# BEFORE the decode trace reads it. The plugin remaps its own buffers (and the seed RNG via
|
| 455 |
# super().decode_forward), but GDN state is model-internal, so mirror the same reindex here.
|
| 456 |
# slot_remap is passed through unchanged so the seed-RNG remap inside super() still runs.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 457 |
if model.use_tp and model.args.max_batch_size > 1:
|
| 458 |
slot_remap = kwargs.get("slot_remap")
|
| 459 |
if slot_remap is not None:
|
| 460 |
-
model._remap_gdn_slots(slot_remap)
|
|
|
|
|
|
|
|
|
|
| 461 |
# Decode bucketing (default on; TT_DECODE_BUCKETING=0 off): slice host inputs to the
|
| 462 |
# smallest power-of-2 width >= active prefix [0:num_active) before the base forward.
|
| 463 |
# No runner edit / output re-pad — plugin reads unpadded_batch_size in slot order.
|
|
@@ -509,7 +601,23 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 509 |
_sm = getattr(_m, "sampling", None)
|
| 510 |
if _sm is not None and hasattr(_sm, "set_trace_bucket"):
|
| 511 |
_sm.set_trace_bucket(B)
|
| 512 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 513 |
|
| 514 |
def warmup_model_prefill(self, kv_cache, enable_trace, *args, **kwargs):
|
| 515 |
# Capture the chunk-prefill trace + warm the masked-bucket set so requests only replay
|
|
@@ -540,6 +648,9 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 540 |
)
|
| 541 |
prev = model._bind_gdn_prefill_scratch() if batched else None
|
| 542 |
try:
|
|
|
|
|
|
|
|
|
|
| 543 |
model.capture_prefill_trace_chunked(
|
| 544 |
self.mesh_device, page_table, chunk_size=_PREFILL_WARMUP_CHUNK, capture_chunk_trace=True
|
| 545 |
)
|
|
@@ -550,11 +661,20 @@ class Qwen36ForCausalLM(Generator, SupportsMultiModal):
|
|
| 550 |
# Compile the device-side slot-write programs (QWEN36_GDN_SLOT_DEVICE_COPY=2: fill_cache + masked where) and
|
| 551 |
# upload the per-slot row masks now, so the first real request does not pay ~450 ms for it.
|
| 552 |
model.warmup_gdn_slot_write()
|
|
|
|
|
|
|
|
|
|
| 553 |
# Steady-state view: weights + KV pool + GDN slot state + the persistent prefill buffers are all allocated
|
| 554 |
# (the decode traces are captured earlier by warmup_model_decode). The free DRAM here, minus a margin for
|
| 555 |
# the transient prefill activations, is the headroom QWEN36_MAX_TOKENS_ALL_USERS can grow into.
|
| 556 |
_log_device_memory(self.mesh_device, "after prefill warmup")
|
| 557 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 558 |
def warmup_model_decode(self, *args, **kwargs):
|
| 559 |
# Defer to WarmupForwardMixin, which warms the paged-SDPA + GDN decode path at pos 0.
|
| 560 |
# Drop stale `non_greedy_decoding_on_device` from the old vLLM plugin; no-op for Qwen.
|
|
|
|
| 84 |
# so no device-resident token/position chain is ever assumed; the plugin then overlaps only the
|
| 85 |
# engine's scheduling/output work with the device step (its steady-decode fast path needs device sampling).
|
| 86 |
# The gate is read when the platform queries the capability (_ModelCapabilities), not at import time.
|
| 87 |
+
# QWEN36_CHUNKED_PREFILL=1 (default off; batched TP text serving) declares scheduler-driven chunked prefill:
|
| 88 |
+
# ``supports_chunked_prefill`` True plus ``tt_prefill_chunk_tokens`` (the 2048 chunk-trace unit; the plugin's chunk
|
| 89 |
+
# size must be a multiple of it), consumed by the plugin's chunk policy. A resumed chunk continues from the GDN
|
| 90 |
+
# state parked in the B=1 prefill scratch (model.prefill_paged_slots, tt/chunked_prefill.py). With the knob off
|
| 91 |
+
# both keys read as absent (``in`` False, ``.get`` returns the caller's default, ``[]`` of the chunk size raises),
|
| 92 |
+
# so every profile resolves exactly as before. The two keys are DYNAMIC: they live only in the accessors, not in
|
| 93 |
+
# the dict storage, so a copy (``dict(caps)``, ``{**caps}``) or ``.items()`` / iteration sees only the static
|
| 94 |
+
# entries. Consumers must query the class attribute through ``[]`` / ``in`` / ``.get`` (the plugin does), and a
|
| 95 |
+
# subclass that copies the dict (the DFlash class) must state its own chunked-prefill entries explicitly.
|
| 96 |
+
# The knob is model-side only: with QWEN36_CHUNKED_PREFILL=1 the park buffer (~147 MiB/device at TP=2) is
|
| 97 |
+
# allocated at warmup even when the plugin later leaves the policy off (async scheduling, kv_transfer_config /
|
| 98 |
+
# p1d1, --no-enable-chunked-prefill), which only costs KV-pool headroom; set the knob only on a chunking profile.
|
| 99 |
class _ModelCapabilities(dict):
|
| 100 |
+
"""dict whose ``supports_async_decode`` (QWEN36_ASYNC_DECODE_OK) and chunked-prefill (QWEN36_CHUNKED_PREFILL)
|
| 101 |
+
entries are read from the environment at access time."""
|
| 102 |
|
| 103 |
_ASYNC = "supports_async_decode"
|
| 104 |
+
_CHUNKED = "supports_chunked_prefill"
|
| 105 |
+
_CHUNK_TOKENS = "tt_prefill_chunk_tokens"
|
| 106 |
+
|
| 107 |
+
@staticmethod
|
| 108 |
+
def _chunked_on():
|
| 109 |
+
return os.environ.get("QWEN36_CHUNKED_PREFILL", "0") == "1"
|
| 110 |
|
| 111 |
def __getitem__(self, key):
|
| 112 |
if key == self._ASYNC:
|
| 113 |
return os.environ.get("QWEN36_ASYNC_DECODE_OK", "0") == "1"
|
| 114 |
+
if key == self._CHUNKED:
|
| 115 |
+
return self._chunked_on()
|
| 116 |
+
if key == self._CHUNK_TOKENS:
|
| 117 |
+
if not self._chunked_on():
|
| 118 |
+
raise KeyError(key)
|
| 119 |
+
return _PREFILL_WARMUP_CHUNK
|
| 120 |
return super().__getitem__(key)
|
| 121 |
|
| 122 |
+
def __contains__(self, key):
|
| 123 |
+
if key in (self._CHUNK_TOKENS, self._CHUNKED):
|
| 124 |
+
return self._chunked_on()
|
| 125 |
+
return super().__contains__(key)
|
| 126 |
+
|
| 127 |
def get(self, key, default=None):
|
| 128 |
if key == self._ASYNC:
|
| 129 |
return self[key]
|
| 130 |
+
if key in (self._CHUNKED, self._CHUNK_TOKENS):
|
| 131 |
+
return self[key] if self._chunked_on() else default
|
| 132 |
return super().get(key, default)
|
| 133 |
|
| 134 |
model_capabilities = _ModelCapabilities(
|
|
|
|
| 383 |
"batched (max_num_seqs>1) serving is text-only; multimodal is single-sequence "
|
| 384 |
"(max_concurrency=1). Run the model at max_num_seqs=1 for image/video requests."
|
| 385 |
)
|
| 386 |
+
return self._prefill_forward_tp_batched(
|
| 387 |
+
model,
|
| 388 |
+
tokens,
|
| 389 |
+
page_table,
|
| 390 |
+
prompt_lens,
|
| 391 |
+
kwargs.get("empty_slots"),
|
| 392 |
+
start_pos=kwargs.get("start_pos"),
|
| 393 |
+
resume_mask=kwargs.get("prefill_resume_mask"),
|
| 394 |
+
final_mask=kwargs.get("prefill_final_mask"),
|
| 395 |
+
)
|
| 396 |
+
resume_mask = kwargs.get("prefill_resume_mask")
|
| 397 |
+
final_mask = kwargs.get("prefill_final_mask")
|
| 398 |
+
if (resume_mask is not None and any(bool(m) for m in resume_mask)) or (
|
| 399 |
+
final_mask is not None and not all(bool(m) for m in final_mask)
|
| 400 |
+
):
|
| 401 |
+
# Chunks are only split while other requests decode, which needs max_num_seqs > 1 (the batched path).
|
| 402 |
+
raise ValueError(
|
| 403 |
+
f"chunked prefill needs the batched TP path (use_tp={model.use_tp}, "
|
| 404 |
+
f"max_batch_size={model.args.max_batch_size}); resume_mask={resume_mask} final_mask={final_mask}"
|
| 405 |
+
)
|
| 406 |
vision_tokens = self._compute_vision_tokens(model, kwargs)
|
| 407 |
if model.use_tp:
|
| 408 |
return self._prefill_forward_tp(model, tokens, page_table, prompt_lens, vision_tokens=vision_tokens)
|
|
|
|
| 468 |
logger.info(f"Finished prefill up to {T} tokens, starting decode...")
|
| 469 |
return logits, torch.zeros(1, dtype=torch.long)
|
| 470 |
|
| 471 |
+
def _prefill_forward_tp_batched(
|
| 472 |
+
self, model, tokens, page_table, prompt_lens, empty_slots, start_pos=None, resume_mask=None, final_mask=None
|
| 473 |
+
):
|
| 474 |
"""TP batched (max_num_seqs>1) prefill: prefill each request in this step into its decode slot.
|
| 475 |
|
| 476 |
vLLM prefills new requests while other slots decode, so each user's B=1 state is written into
|
|
|
|
| 481 |
page_table: torch [N, max_blocks] — row u = request u's blocks.
|
| 482 |
prompt_lens: per-request real lengths (row u trimmed to prompt_lens[u]).
|
| 483 |
empty_slots: per-request decode slot; defaults to range(N) (mirrors Generator.prefill_forward_text).
|
| 484 |
+
start_pos / resume_mask / final_mask: scheduler-driven chunked prefill (QWEN36_CHUNKED_PREFILL=1): a row with
|
| 485 |
+
resume_mask True continues its prompt at start_pos (the chunk start); final_mask False marks an
|
| 486 |
+
intermediate chunk (prompt_lens = the chunk end). Rows are always passed from token 0. Without the
|
| 487 |
+
masks start_pos is ignored and every row prefills from 0 (the pre-chunking behaviour).
|
| 488 |
Returns ([N, 1, vocab] host logits, [N] zero rope_deltas — text M-RoPE delta is 0, applied model-side).
|
| 489 |
"""
|
| 490 |
N = tokens.shape[0]
|
|
|
|
| 494 |
empty_slots = [int(s) for s in empty_slots]
|
| 495 |
token_ids_list = [tokens[u : u + 1, : plens[u]].to(torch.int32) for u in range(N)]
|
| 496 |
pt = page_table if isinstance(page_table, torch.Tensor) else ttnn.to_torch(page_table)
|
| 497 |
+
chunked = resume_mask is not None or final_mask is not None
|
| 498 |
+
chunk_desc = ""
|
| 499 |
+
if chunked and (
|
| 500 |
+
(resume_mask is not None and any(resume_mask)) or (final_mask is not None and not all(final_mask))
|
| 501 |
+
):
|
| 502 |
+
# (start, end, resume, final) per row: the one line that shows a split prefill in the server log
|
| 503 |
+
res = [bool(resume_mask[u]) if resume_mask is not None else False for u in range(N)]
|
| 504 |
+
fin = [bool(final_mask[u]) if final_mask is not None else True for u in range(N)]
|
| 505 |
+
chunk_desc = " chunked rows " + str(
|
| 506 |
+
[(int(start_pos[u]) if res[u] else 0, plens[u], res[u], fin[u]) for u in range(N)]
|
| 507 |
+
)
|
| 508 |
+
logger.info(f"Prefilling {N} user(s) into slots {empty_slots} (TP batched masked-bucket){chunk_desc}")
|
| 509 |
+
host_logits = model.prefill_paged_slots(
|
| 510 |
+
token_ids_list,
|
| 511 |
+
pt,
|
| 512 |
+
empty_slots,
|
| 513 |
+
valid_lens=plens,
|
| 514 |
+
start_positions=[int(start_pos[u]) for u in range(N)] if chunked and start_pos is not None else None,
|
| 515 |
+
resume_mask=[bool(resume_mask[u]) for u in range(N)] if resume_mask is not None else None,
|
| 516 |
+
final_mask=[bool(final_mask[u]) for u in range(N)] if final_mask is not None else None,
|
| 517 |
+
)
|
| 518 |
logits = torch.cat([hl.reshape(1, 1, -1) for hl in host_logits], dim=0) # [N, 1, vocab]
|
| 519 |
logger.info(f"Finished batched prefill of {N} user(s), starting decode...")
|
| 520 |
return logits, torch.zeros(N, dtype=torch.long)
|
|
|
|
| 530 |
# BEFORE the decode trace reads it. The plugin remaps its own buffers (and the seed RNG via
|
| 531 |
# super().decode_forward), but GDN state is model-internal, so mirror the same reindex here.
|
| 532 |
# slot_remap is passed through unchanged so the seed-RNG remap inside super() still runs.
|
| 533 |
+
# QWEN36_DECODE_SLOW_LOG=1 (perf triage, off by default): drain the device FIRST, so neither the step time
|
| 534 |
+
# nor the remap time includes work still queued from the previous step (a prefill's slot-write copies, the
|
| 535 |
+
# conv-history sync), then time the GDN slot remap phase by phase (model._remap_gdn_slots timing=: device
|
| 536 |
+
# synced between phases) and log every remap step and every other step slower than 150 ms. The syncs slow
|
| 537 |
+
# the logged steps a little; do not quote TPOT from a run with it on.
|
| 538 |
+
_slow_log = os.environ.get("QWEN36_DECODE_SLOW_LOG", "0") == "1"
|
| 539 |
+
if _slow_log:
|
| 540 |
+
_tq0 = time.perf_counter()
|
| 541 |
+
ttnn.synchronize_device(self.mesh_device)
|
| 542 |
+
_t_queued = time.perf_counter() - _tq0
|
| 543 |
+
_td0 = time.perf_counter() if _slow_log else 0.0
|
| 544 |
+
_t_remap = 0.0
|
| 545 |
+
_remap_timing = {} if _slow_log else None
|
| 546 |
if model.use_tp and model.args.max_batch_size > 1:
|
| 547 |
slot_remap = kwargs.get("slot_remap")
|
| 548 |
if slot_remap is not None:
|
| 549 |
+
model._remap_gdn_slots(slot_remap, timing=_remap_timing)
|
| 550 |
+
if _slow_log:
|
| 551 |
+
ttnn.synchronize_device(self.mesh_device)
|
| 552 |
+
_t_remap = time.perf_counter() - _td0
|
| 553 |
# Decode bucketing (default on; TT_DECODE_BUCKETING=0 off): slice host inputs to the
|
| 554 |
# smallest power-of-2 width >= active prefix [0:num_active) before the base forward.
|
| 555 |
# No runner edit / output re-pad — plugin reads unpadded_batch_size in slot order.
|
|
|
|
| 601 |
_sm = getattr(_m, "sampling", None)
|
| 602 |
if _sm is not None and hasattr(_sm, "set_trace_bucket"):
|
| 603 |
_sm.set_trace_bucket(B)
|
| 604 |
+
if not _slow_log:
|
| 605 |
+
return super().decode_forward(*args, **kwargs)
|
| 606 |
+
out = super().decode_forward(*args, **kwargs)
|
| 607 |
+
_dt = time.perf_counter() - _td0
|
| 608 |
+
if _dt > 0.15 or kwargs.get("slot_remap") is not None:
|
| 609 |
+
_rt = _remap_timing or {}
|
| 610 |
+
_phases = " ".join(
|
| 611 |
+
f"{k}={1e3 * _rt[k]:.0f}" for k in ("taps_s", "rec_s", "conv_s", "packed_s", "hist_sync_s") if k in _rt
|
| 612 |
+
)
|
| 613 |
+
logger.info(
|
| 614 |
+
f"[DECODE_SLOW] {1e3 * _dt:.0f} ms B={int(tokens.shape[0]) if tokens is not None else None} "
|
| 615 |
+
f"remap={'yes' if kwargs.get('slot_remap') is not None else 'no'} remap_ms={1e3 * _t_remap:.0f} "
|
| 616 |
+
f"queued_ms={1e3 * _t_queued:.0f} moved={_rt.get('moved', 0)} cross_parity={_rt.get('cross_parity', 0)} "
|
| 617 |
+
f"packed={'yes' if _rt.get('packed') else 'no'} phases_ms[{_phases}] "
|
| 618 |
+
f"kw={sorted(k for k in kwargs if kwargs[k] is not None and k not in ('tokens', 'page_table', 'kv_cache'))}"
|
| 619 |
+
)
|
| 620 |
+
return out
|
| 621 |
|
| 622 |
def warmup_model_prefill(self, kv_cache, enable_trace, *args, **kwargs):
|
| 623 |
# Capture the chunk-prefill trace + warm the masked-bucket set so requests only replay
|
|
|
|
| 648 |
)
|
| 649 |
prev = model._bind_gdn_prefill_scratch() if batched else None
|
| 650 |
try:
|
| 651 |
+
if batched and self.model_capabilities.get("supports_chunked_prefill", False):
|
| 652 |
+
# Chunked prefill parks a partial prompt's scratch state: allocate + compile it before any trace.
|
| 653 |
+
model.ensure_gdn_park_buffer()
|
| 654 |
model.capture_prefill_trace_chunked(
|
| 655 |
self.mesh_device, page_table, chunk_size=_PREFILL_WARMUP_CHUNK, capture_chunk_trace=True
|
| 656 |
)
|
|
|
|
| 661 |
# Compile the device-side slot-write programs (QWEN36_GDN_SLOT_DEVICE_COPY=2: fill_cache + masked where) and
|
| 662 |
# upload the per-slot row masks now, so the first real request does not pay ~450 ms for it.
|
| 663 |
model.warmup_gdn_slot_write()
|
| 664 |
+
# Same for the fast GDN slot remap (QWEN36_GDN_REMAP_FAST): a first-seen remap program costs ~300 ms of JIT.
|
| 665 |
+
if self._warm_gdn_remap():
|
| 666 |
+
model.warmup_gdn_remap()
|
| 667 |
# Steady-state view: weights + KV pool + GDN slot state + the persistent prefill buffers are all allocated
|
| 668 |
# (the decode traces are captured earlier by warmup_model_decode). The free DRAM here, minus a margin for
|
| 669 |
# the transient prefill activations, is the headroom QWEN36_MAX_TOKENS_ALL_USERS can grow into.
|
| 670 |
_log_device_memory(self.mesh_device, "after prefill warmup")
|
| 671 |
|
| 672 |
+
def _warm_gdn_remap(self):
|
| 673 |
+
"""Whether warmup_model_prefill compiles the fast GDN slot-remap programs: yes for every class whose decode
|
| 674 |
+
applies the plugin's slot_remap to the device GDN state (this one). Subclasses that never remap on device
|
| 675 |
+
(the speculative DFlash decode composes slot_remap into a row indirection) return False."""
|
| 676 |
+
return True
|
| 677 |
+
|
| 678 |
def warmup_model_decode(self, *args, **kwargs):
|
| 679 |
# Defer to WarmupForwardMixin, which warms the paged-SDPA + GDN decode path at pos 0.
|
| 680 |
# Drop stale `non_greedy_decoding_on_device` from the old vLLM plugin; no-op for Qwen.
|
code/models/demos/blackhole/qwen36/tt/qwen36_vllm_dflash.py
CHANGED
|
@@ -123,6 +123,63 @@ _PREFILL_CHUNK = 2048
|
|
| 123 |
# verify's reach past the block (W committed positions + K+1 candidate rows).
|
| 124 |
_MAX_DRAFT = 15
|
| 125 |
_DEBUG = os.environ.get("QWEN36_DFLASH_DEBUG", "0") == "1"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 126 |
|
| 127 |
|
| 128 |
def parse_buckets(spec):
|
|
@@ -206,24 +263,36 @@ def bucket_id(bucket):
|
|
| 206 |
class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
| 207 |
"""Qwen36ForCausalLM + model-internal multi-user DFlash2 speculation on decode steps (see module doc)."""
|
| 208 |
|
| 209 |
-
model_capabilities =
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
|
| 217 |
-
|
| 218 |
-
|
| 219 |
-
|
| 220 |
-
|
| 221 |
-
|
| 222 |
-
|
| 223 |
-
|
| 224 |
-
|
| 225 |
-
|
| 226 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 227 |
|
| 228 |
def __init__(self, *args, **kwargs):
|
| 229 |
super().__init__(*args, **kwargs)
|
|
@@ -247,6 +316,22 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 247 |
self._buckets = () # (B_cfg, T_cfg) verify geometries, largest B == max_num_seqs (module doc)
|
| 248 |
self._multi_bucket = False
|
| 249 |
self._last_bucket = None # bucket id plan() last put in force (logged on change)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 250 |
if _W > 1:
|
| 251 |
if not model.use_tp:
|
| 252 |
raise RuntimeError(
|
|
@@ -262,6 +347,11 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 262 |
f"{_BUCKETS_ENV}={os.environ.get(_BUCKETS_ENV)!r} needs the multi-bucket decoder "
|
| 263 |
"(dflash2_serving.DFlash2DualBucketDecoder), which this tree does not have"
|
| 264 |
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 265 |
logger.info(
|
| 266 |
f"Qwen36DFlash serving: slots={B} block W={_W} tokens/step{' (ragged: one iteration/step)' if _RAGGED else ''}, "
|
| 267 |
f"buckets={','.join(bucket_id(bt) for bt in self._buckets)} "
|
|
@@ -416,10 +506,15 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 416 |
return pt[:, :nb].contiguous()
|
| 417 |
return torch.cat([pt, torch.zeros(1, nb - pt.shape[1], dtype=torch.int32)], dim=1)
|
| 418 |
|
| 419 |
-
def _spec_prefill(self, model, dec, phys, prompt, T, pt_row):
|
| 420 |
"""Eager tap-capturing prefill of ONE request into physical slot ``phys`` (masked bucket for a
|
| 421 |
short prompt, 2048-token chunks + masked tail for a long one), each chunk's taps ingested into the
|
| 422 |
-
drafter's ring for that slot. Returns host logits [1, vocab] (float).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 423 |
|
| 424 |
def on_chunk(hidden, chunk_start, valid_len):
|
| 425 |
taps = model.take_dflash_eager_taps()
|
|
@@ -427,12 +522,34 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 427 |
raise RuntimeError("Qwen36DFlash: eager prefill captured no drafter taps (bucket trace gate on?)")
|
| 428 |
dec.ingest_prompt(phys, taps, chunk_start + valid_len, chunk_start=chunk_start)
|
| 429 |
|
| 430 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 431 |
model._dflash_tap = True
|
| 432 |
try:
|
| 433 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 434 |
finally:
|
| 435 |
model._dflash_tap = False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 436 |
if _FAST_LOGITS:
|
| 437 |
# Replicated, tile-padded to 32 rows: untilize on device (31 padding rows dropped, 16 MB -> 0.5 MB per
|
| 438 |
# device) and read device 0 only -- the QWEN36_PREFILL_LOGITS_FAST path of prefill_paged_slots. The
|
|
@@ -469,6 +586,14 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 469 |
self._in_warmup = False
|
| 470 |
return out
|
| 471 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 472 |
def warmup_model_prefill(self, *args, **kwargs):
|
| 473 |
self._in_warmup = True
|
| 474 |
try:
|
|
@@ -522,6 +647,11 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 522 |
S0 = min(buckets[0], _PREFILL_CHUNK - 1)
|
| 523 |
for u in range(1, B):
|
| 524 |
self._spec_prefill(model, dec, u, dummy_prompt(S0, seed=100 + u), S0, rows[u])
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 525 |
ttnn.synchronize_device(model.mesh_device)
|
| 526 |
logger.info(f"Qwen36DFlash phase-1 warmup (alloc + compile) done in {time.perf_counter() - t0:.1f}s")
|
| 527 |
self._warm_rows = rows
|
|
@@ -706,6 +836,9 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 706 |
dec.set_bucket(ids[bt])
|
| 707 |
dec.step()
|
| 708 |
dec.end(0)
|
|
|
|
|
|
|
|
|
|
| 709 |
ttnn.synchronize_device(dev)
|
| 710 |
n1 = int(count())
|
| 711 |
if n1 != n0:
|
|
@@ -721,6 +854,7 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 721 |
logger.info(
|
| 722 |
f"Qwen36DFlash post-capture guard: program cache unchanged at {n0} entries over the warm sweep + one "
|
| 723 |
f"traced step per bucket ({','.join(ids[bt] for bt in self._buckets)})"
|
|
|
|
| 724 |
)
|
| 725 |
|
| 726 |
def _spec_capture(self):
|
|
@@ -762,6 +896,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 762 |
self._reset_bucket_policy(dec)
|
| 763 |
self._spec = dec
|
| 764 |
self._spec_pre = None
|
|
|
|
|
|
|
|
|
|
|
|
|
| 765 |
logger.info(f"Qwen36DFlash phase-2 warmup (captures) done in {time.perf_counter() - t0:.1f}s: {stats} (W={_W})")
|
| 766 |
|
| 767 |
def _reset_bucket_policy(self, dec):
|
|
@@ -818,6 +956,14 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 818 |
|
| 819 |
def prefill_forward(self, tokens, page_table, kv_cache, prompt_lens, **kwargs):
|
| 820 |
if _W <= 1 or not self._spec_ready():
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 821 |
return super().prefill_forward(tokens, page_table, kv_cache, prompt_lens, **kwargs)
|
| 822 |
model = self.model[0]
|
| 823 |
dec = self._spec
|
|
@@ -830,6 +976,41 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 830 |
empty_slots = kwargs.get("empty_slots")
|
| 831 |
logical = [int(s) for s in empty_slots] if empty_slots is not None else list(range(N))
|
| 832 |
pt = torch.as_tensor(page_table)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 833 |
out = []
|
| 834 |
for u in range(N):
|
| 835 |
phys = self._phys[logical[u]]
|
|
@@ -847,6 +1028,114 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 847 |
logger.info(f"Finished prefill of {N} request(s), starting decode...")
|
| 848 |
return torch.cat(out, dim=0), torch.zeros(N, dtype=torch.long)
|
| 849 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 850 |
# ------------------------------------------------------------------ decode: one block per step, all live slots
|
| 851 |
def _no_device_sampler(self):
|
| 852 |
"""True when no model on the mesh has the ttnn device sampler (TP=2: 124,160 logits/device > its 64K cap)."""
|
|
@@ -854,6 +1143,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 854 |
|
| 855 |
def decode_forward(self, *args, **kwargs):
|
| 856 |
if _W <= 1 or not self._spec_ready():
|
|
|
|
|
|
|
|
|
|
|
|
|
| 857 |
if _W > 1 and kwargs.get("sampling_params") is not None and self._no_device_sampler():
|
| 858 |
# Pre-arm plain step on a mesh without a device sampler: host logits (the runner samples).
|
| 859 |
kwargs = dict(kwargs, sampling_params=None)
|
|
@@ -891,6 +1184,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 891 |
# user), so the begins below land in the geometry in force; a mis-ordering is a decoder assert, never a
|
| 892 |
# wrong-geometry seed. Single bucket: a host no-op.
|
| 893 |
plan = getattr(dec, "plan", None)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 894 |
if plan is not None:
|
| 895 |
live_after = {phys for phys in range(B) if dec.active[phys]}
|
| 896 |
for i in range(Bp):
|
|
@@ -900,6 +1197,10 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 900 |
if bucket != self._last_bucket:
|
| 901 |
logger.info(f"Qwen36DFlash: verify bucket {bucket} in force ({len(live_after)} live slot(s))")
|
| 902 |
self._last_bucket = bucket
|
|
|
|
|
|
|
|
|
|
|
|
|
| 903 |
live_rows = []
|
| 904 |
nosession_rows = [] # live rows without a session: EOS so the request ends (ragged: [EOS, -1, ...])
|
| 905 |
for i in range(Bp):
|
|
@@ -934,9 +1235,21 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 934 |
live_rows.append((i, phys))
|
| 935 |
t0 = time.perf_counter() if _DEBUG else 0.0
|
| 936 |
if _RAGGED:
|
|
|
|
|
|
|
|
|
|
| 937 |
out = self._decode_ragged(dec, live_rows, Bp, t0)
|
| 938 |
for i in nosession_rows:
|
| 939 |
out[i, 0] = self._eos_fill
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 940 |
return out
|
| 941 |
# Step until every live row can fill its block. A row whose carry already holds a stop token
|
| 942 |
# needs nothing: it emits through the stop this step (latency for a normal request); the
|
|
@@ -1068,6 +1381,11 @@ class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
|
| 1068 |
self._spec.end(phys)
|
| 1069 |
self._pending[phys] = None
|
| 1070 |
self._carry[phys], self._stopped[phys], self._prev_tail[phys] = [], False, None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1071 |
|
| 1072 |
def release_persistent_capture(self) -> None:
|
| 1073 |
if self._spec is None:
|
|
|
|
| 123 |
# verify's reach past the block (W committed positions + K+1 candidate rows).
|
| 124 |
_MAX_DRAFT = 15
|
| 125 |
_DEBUG = os.environ.get("QWEN36_DFLASH_DEBUG", "0") == "1"
|
| 126 |
+
# Scheduler-driven chunked prefill (opt-in; profiles/tp4_chunked/DESIGN.md): with the knob on, the class declares
|
| 127 |
+
# supports_chunked_prefill + tt_prefill_chunk_tokens (the 2048 eager unit) + tt_block_output_chunked_prefill, the plugin's
|
| 128 |
+
# TT chunk policy may split a long prompt into 2048-aligned chunks interleaved with the other users' speculative decode
|
| 129 |
+
# steps, and prefill_forward resumes a partial prompt through the eager prefill_for_spec path (GDN state parked in the
|
| 130 |
+
# B=1 scratch between calls, tt/chunked_prefill.py; drafter context in the slot's own ring). Off (default): the
|
| 131 |
+
# capability dict, prefill_forward, the warm-up and the guards are exactly today's.
|
| 132 |
+
_CP_ENV = "QWEN36_DFLASH_CHUNKED_PREFILL"
|
| 133 |
+
# Per-row prefill timing of the chunked path (device-synchronised; measurement only, off by default).
|
| 134 |
+
_CP_TIMING = os.environ.get("QWEN36_DFLASH_CP_TIMING", "0") == "1"
|
| 135 |
+
# Drained per-step timers in decode_forward (plan/switch, begins, step; measurement only, off by default).
|
| 136 |
+
_STEP_LOG = os.environ.get("QWEN36_DFLASH_STEP_LOG", "0") == "1"
|
| 137 |
+
|
| 138 |
+
|
| 139 |
+
def dflash_chunked_prefill_on():
|
| 140 |
+
"""The chunked-prefill knob, read at access time (speculation must be on: a plain-serving DFlash class is the base
|
| 141 |
+
class's behaviour and never declares the block-output chunk capability)."""
|
| 142 |
+
return _W > 1 and os.environ.get(_CP_ENV, "0") == "1"
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
class _DFlashCapabilities(dict):
|
| 146 |
+
"""The DFlash class's capability dict: the static entries are exactly today's (a ``dict()`` / ``{**}`` copy, ``items()``
|
| 147 |
+
and iteration see only them, ``supports_chunked_prefill`` present and False); with QWEN36_DFLASH_CHUNKED_PREFILL=1
|
| 148 |
+
(read at access time) ``[]`` / ``in`` / ``.get`` report chunked prefill on, the 2048-token chunk unit and the
|
| 149 |
+
block-output chunk contract. Only these three keys are dynamic (``supports_async_decode`` stays the static False).
|
| 150 |
+
"""
|
| 151 |
+
|
| 152 |
+
_CHUNKED = "supports_chunked_prefill"
|
| 153 |
+
_DYNAMIC_ONLY = ("tt_prefill_chunk_tokens", "tt_block_output_chunked_prefill")
|
| 154 |
+
|
| 155 |
+
@staticmethod
|
| 156 |
+
def _on_values():
|
| 157 |
+
return {
|
| 158 |
+
"supports_chunked_prefill": True,
|
| 159 |
+
"tt_prefill_chunk_tokens": _PREFILL_CHUNK,
|
| 160 |
+
"tt_block_output_chunked_prefill": True,
|
| 161 |
+
}
|
| 162 |
+
|
| 163 |
+
def __getitem__(self, key):
|
| 164 |
+
if key == self._CHUNKED or key in self._DYNAMIC_ONLY:
|
| 165 |
+
if dflash_chunked_prefill_on():
|
| 166 |
+
return self._on_values()[key]
|
| 167 |
+
if key in self._DYNAMIC_ONLY:
|
| 168 |
+
raise KeyError(key)
|
| 169 |
+
return super().__getitem__(key)
|
| 170 |
+
|
| 171 |
+
def __contains__(self, key):
|
| 172 |
+
if key in self._DYNAMIC_ONLY:
|
| 173 |
+
return dflash_chunked_prefill_on()
|
| 174 |
+
return super().__contains__(key)
|
| 175 |
+
|
| 176 |
+
def get(self, key, default=None):
|
| 177 |
+
if key == self._CHUNKED or key in self._DYNAMIC_ONLY:
|
| 178 |
+
if dflash_chunked_prefill_on():
|
| 179 |
+
return self._on_values()[key]
|
| 180 |
+
if key in self._DYNAMIC_ONLY:
|
| 181 |
+
return default
|
| 182 |
+
return super().get(key, default)
|
| 183 |
|
| 184 |
|
| 185 |
def parse_buckets(spec):
|
|
|
|
| 263 |
class Qwen36DFlashForCausalLM(Qwen36ForCausalLM):
|
| 264 |
"""Qwen36ForCausalLM + model-internal multi-user DFlash2 speculation on decode steps (see module doc)."""
|
| 265 |
|
| 266 |
+
model_capabilities = _DFlashCapabilities(
|
| 267 |
+
{
|
| 268 |
+
**Qwen36ForCausalLM.model_capabilities,
|
| 269 |
+
# No scheduler-driven chunked prefill by default (QWEN36_CHUNKED_PREFILL is the plain class's knob and must not
|
| 270 |
+
# reach this class): the plain class reads its knob at access time; this dict copies only its static entries, so
|
| 271 |
+
# state the default explicitly (with _W == 1 the block-output refusals in the plugin would not catch it).
|
| 272 |
+
# QWEN36_DFLASH_CHUNKED_PREFILL=1 turns it on through _DFlashCapabilities' accessors (module constants).
|
| 273 |
+
"supports_chunked_prefill": False,
|
| 274 |
+
# A decode step commits exactly _W tokens per request (EOS-filled at a stop).
|
| 275 |
+
"output_tokens_per_step": _W,
|
| 276 |
+
# Block on decode steps; prefill anchors are plain width-1 rows.
|
| 277 |
+
"tt_adaptive_block_output": _W > 1,
|
| 278 |
+
# ...for EVERY request of the step (one multi-user speculative step), not only when solo.
|
| 279 |
+
"tt_adaptive_block_batched": _W > 1,
|
| 280 |
+
# Ragged rows: one speculative iteration per step, each row 1..W real ids then -1 padding
|
| 281 |
+
# (QWEN36_DFLASH_RAGGED=1; the plugin must implement the same contract).
|
| 282 |
+
"tt_adaptive_block_ragged": _W > 1 and _RAGGED,
|
| 283 |
+
# EVERY text prompt speculates (0 = no prompt-length frontier): the plain decode / chunk-prefill
|
| 284 |
+
# traces never run on the request path (a spec replay after them hangs the device).
|
| 285 |
+
"tt_adaptive_block_max_prompt_tokens": 0,
|
| 286 |
+
# The block step writes the W committed positions AND the last verify's K+1 candidate rows into
|
| 287 |
+
# the paged KV inside one step: have the scheduler allocate that reach up front.
|
| 288 |
+
"tt_block_output_kv_lookahead_tokens": (_W + _MAX_DRAFT + 1) if _W > 1 else 0,
|
| 289 |
+
}
|
| 290 |
+
)
|
| 291 |
+
|
| 292 |
+
# Chunked-prefill state (set per instance in __init__; class defaults for instances built without it, e.g. host tests)
|
| 293 |
+
_cp_on = False
|
| 294 |
+
_cp_owner_phys = None
|
| 295 |
+
_forbid_plain = False
|
| 296 |
|
| 297 |
def __init__(self, *args, **kwargs):
|
| 298 |
super().__init__(*args, **kwargs)
|
|
|
|
| 316 |
self._buckets = () # (B_cfg, T_cfg) verify geometries, largest B == max_num_seqs (module doc)
|
| 317 |
self._multi_bucket = False
|
| 318 |
self._last_bucket = None # bucket id plan() last put in force (logged on change)
|
| 319 |
+
# Chunked prefill (QWEN36_DFLASH_CHUNKED_PREFILL=1): the physical slot whose partial prompt the B=1 scratch (or its
|
| 320 |
+
# park buffer) holds, the third ownership check next to the planner's first block and next position.
|
| 321 |
+
self._cp_on = dflash_chunked_prefill_on()
|
| 322 |
+
self._cp_owner_phys = None
|
| 323 |
+
self._forbid_plain = False # armed at the end of the spec capture under the knob (plain-trace tripwire)
|
| 324 |
+
if self._cp_on:
|
| 325 |
+
if B <= 1:
|
| 326 |
+
raise RuntimeError(
|
| 327 |
+
f"{_CP_ENV}=1 needs the batched B=1 prefill scratch (--max-num-seqs > 1); got max_num_seqs={B} "
|
| 328 |
+
"(single-user-dflash2 has nothing to interleave a chunk with)"
|
| 329 |
+
)
|
| 330 |
+
if int(model.num_devices) != 4:
|
| 331 |
+
raise RuntimeError(
|
| 332 |
+
f"{_CP_ENV}=1 is validated at TP=4 (P150x4 / P300x2) only; got {int(model.num_devices)} device(s) "
|
| 333 |
+
"(the TP=2 spec prefill path is changing under lane Q; unset the knob)"
|
| 334 |
+
)
|
| 335 |
if _W > 1:
|
| 336 |
if not model.use_tp:
|
| 337 |
raise RuntimeError(
|
|
|
|
| 347 |
f"{_BUCKETS_ENV}={os.environ.get(_BUCKETS_ENV)!r} needs the multi-bucket decoder "
|
| 348 |
"(dflash2_serving.DFlash2DualBucketDecoder), which this tree does not have"
|
| 349 |
)
|
| 350 |
+
if self._cp_on:
|
| 351 |
+
logger.info(
|
| 352 |
+
f"Qwen36DFlash serving: scheduler-driven chunked prefill ON ({_CP_ENV}=1, chunk unit "
|
| 353 |
+
f"{_PREFILL_CHUNK}): partial prompts resume through the eager spec prefill"
|
| 354 |
+
)
|
| 355 |
logger.info(
|
| 356 |
f"Qwen36DFlash serving: slots={B} block W={_W} tokens/step{' (ragged: one iteration/step)' if _RAGGED else ''}, "
|
| 357 |
f"buckets={','.join(bucket_id(bt) for bt in self._buckets)} "
|
|
|
|
| 506 |
return pt[:, :nb].contiguous()
|
| 507 |
return torch.cat([pt, torch.zeros(1, nb - pt.shape[1], dtype=torch.int32)], dim=1)
|
| 508 |
|
| 509 |
+
def _spec_prefill(self, model, dec, phys, prompt, T, pt_row, start=0, final=True):
|
| 510 |
"""Eager tap-capturing prefill of ONE request into physical slot ``phys`` (masked bucket for a
|
| 511 |
short prompt, 2048-token chunks + masked tail for a long one), each chunk's taps ingested into the
|
| 512 |
+
drafter's ring for that slot. Returns host logits [1, vocab] (float).
|
| 513 |
+
|
| 514 |
+
Chunked prefill: ``start`` > 0 resumes the slot's partial prompt at ``start`` (its drafter frontier
|
| 515 |
+
``dec.ctx_len[phys]`` must equal ``start``; the GDN state must be in the bound scratch, see prefill_for_spec);
|
| 516 |
+
``final=False`` is an intermediate chunk ending at ``T``: no logits (returns None), the drafter frontier is
|
| 517 |
+
left at ``T``."""
|
| 518 |
|
| 519 |
def on_chunk(hidden, chunk_start, valid_len):
|
| 520 |
taps = model.take_dflash_eager_taps()
|
|
|
|
| 522 |
raise RuntimeError("Qwen36DFlash: eager prefill captured no drafter taps (bucket trace gate on?)")
|
| 523 |
dec.ingest_prompt(phys, taps, chunk_start + valid_len, chunk_start=chunk_start)
|
| 524 |
|
| 525 |
+
chunked = bool(start) or not final
|
| 526 |
+
if not start:
|
| 527 |
+
dec.ctx_len[phys] = 0
|
| 528 |
+
elif dec.ctx_len[phys] != start:
|
| 529 |
+
raise RuntimeError(
|
| 530 |
+
f"Qwen36DFlash chunked prefill: slot {phys} resumes at {start} but its drafter context ends at "
|
| 531 |
+
f"{dec.ctx_len[phys]} (the continuation moved slots or a chunk was lost)"
|
| 532 |
+
)
|
| 533 |
model._dflash_tap = True
|
| 534 |
try:
|
| 535 |
+
if chunked:
|
| 536 |
+
logits_dev = model.prefill_for_spec(
|
| 537 |
+
prompt, self._spec_pref_pt(model, pt_row), T, on_chunk, slot=phys, start=int(start), final=final
|
| 538 |
+
)
|
| 539 |
+
else:
|
| 540 |
+
logits_dev = model.prefill_for_spec(prompt, self._spec_pref_pt(model, pt_row), T, on_chunk, slot=phys)
|
| 541 |
finally:
|
| 542 |
model._dflash_tap = False
|
| 543 |
+
if not final:
|
| 544 |
+
if logits_dev is not None:
|
| 545 |
+
ttnn.deallocate(logits_dev)
|
| 546 |
+
raise RuntimeError("Qwen36DFlash chunked prefill: an intermediate chunk produced logits")
|
| 547 |
+
if dec.ctx_len[phys] != T:
|
| 548 |
+
raise RuntimeError(
|
| 549 |
+
f"Qwen36DFlash chunked prefill: intermediate chunk [{start}, {T}) left slot {phys}'s drafter "
|
| 550 |
+
f"context at {dec.ctx_len[phys]}"
|
| 551 |
+
)
|
| 552 |
+
return None
|
| 553 |
if _FAST_LOGITS:
|
| 554 |
# Replicated, tile-padded to 32 rows: untilize on device (31 padding rows dropped, 16 MB -> 0.5 MB per
|
| 555 |
# device) and read device 0 only -- the QWEN36_PREFILL_LOGITS_FAST path of prefill_paged_slots. The
|
|
|
|
| 586 |
self._in_warmup = False
|
| 587 |
return out
|
| 588 |
|
| 589 |
+
def _warm_gdn_remap(self):
|
| 590 |
+
"""Speculative serving (_W > 1) never runs remap_slots: decode_forward composes the plugin's slot_remap into
|
| 591 |
+
_phys, and plain decode is not used once the spec traces exist (the _forbid_plain tripwire). So
|
| 592 |
+
warmup_model_prefill must not compile the fast-remap programs either: at TP=2 (tp2-dflash2, B=4) that would
|
| 593 |
+
run 4 remaps on the B=4 GDN state after the spec traces were captured, shapes no test or served run exercised
|
| 594 |
+
(lane R review R1). _W == 1 is the plain decode path, which does remap."""
|
| 595 |
+
return _W <= 1
|
| 596 |
+
|
| 597 |
def warmup_model_prefill(self, *args, **kwargs):
|
| 598 |
self._in_warmup = True
|
| 599 |
try:
|
|
|
|
| 647 |
S0 = min(buckets[0], _PREFILL_CHUNK - 1)
|
| 648 |
for u in range(1, B):
|
| 649 |
self._spec_prefill(model, dec, u, dummy_prompt(S0, seed=100 + u), S0, rows[u])
|
| 650 |
+
if self._cp_on:
|
| 651 |
+
# 3) Chunked prefill: the park buffer (before ANY trace is captured; the base prefill warm-up's own call is
|
| 652 |
+
# then a no-op) and every resume / park program, through the chunk policy's own orchestration.
|
| 653 |
+
model.ensure_gdn_park_buffer()
|
| 654 |
+
self._cp_warm_sequence(model, dec, rows, S0, "phase-1 warm-up")
|
| 655 |
ttnn.synchronize_device(model.mesh_device)
|
| 656 |
logger.info(f"Qwen36DFlash phase-1 warmup (alloc + compile) done in {time.perf_counter() - t0:.1f}s")
|
| 657 |
self._warm_rows = rows
|
|
|
|
| 836 |
dec.set_bucket(ids[bt])
|
| 837 |
dec.step()
|
| 838 |
dec.end(0)
|
| 839 |
+
if self._cp_on:
|
| 840 |
+
# ...and one chunked prompt with a rider between its chunks (park, unpark, resumed chunk and tail).
|
| 841 |
+
self._cp_warm_sequence(model, dec, rows, S0, "post-capture guard")
|
| 842 |
ttnn.synchronize_device(dev)
|
| 843 |
n1 = int(count())
|
| 844 |
if n1 != n0:
|
|
|
|
| 854 |
logger.info(
|
| 855 |
f"Qwen36DFlash post-capture guard: program cache unchanged at {n0} entries over the warm sweep + one "
|
| 856 |
f"traced step per bucket ({','.join(ids[bt] for bt in self._buckets)})"
|
| 857 |
+
+ (" + a chunked prefill with a rider" if self._cp_on else "")
|
| 858 |
)
|
| 859 |
|
| 860 |
def _spec_capture(self):
|
|
|
|
| 896 |
self._reset_bucket_policy(dec)
|
| 897 |
self._spec = dec
|
| 898 |
self._spec_pre = None
|
| 899 |
+
if self._cp_on:
|
| 900 |
+
# Tripwire (hang class): from here on a plain trace replay raises instead of hanging the device.
|
| 901 |
+
model._forbid_plain_traces = True
|
| 902 |
+
self._forbid_plain = True
|
| 903 |
logger.info(f"Qwen36DFlash phase-2 warmup (captures) done in {time.perf_counter() - t0:.1f}s: {stats} (W={_W})")
|
| 904 |
|
| 905 |
def _reset_bucket_policy(self, dec):
|
|
|
|
| 956 |
|
| 957 |
def prefill_forward(self, tokens, page_table, kv_cache, prompt_lens, **kwargs):
|
| 958 |
if _W <= 1 or not self._spec_ready():
|
| 959 |
+
# Reachable with the tripwire armed only after release_persistent_capture (_spec None) while a plugin
|
| 960 |
+
# warm-up runs again (_in_warmup; otherwise _spec_ready() raises first): a plain forward would replay plain
|
| 961 |
+
# traces in a process that captured spec traces. On the request path _spec_ready() is True whenever the
|
| 962 |
+
# tripwire is armed; there the guard is model._check_plain_trace_allowed at the plain prefill replay sites.
|
| 963 |
+
if self._forbid_plain:
|
| 964 |
+
raise RuntimeError(
|
| 965 |
+
"Qwen36DFlash: plain prefill after the speculative traces were captured (hang class)"
|
| 966 |
+
)
|
| 967 |
return super().prefill_forward(tokens, page_table, kv_cache, prompt_lens, **kwargs)
|
| 968 |
model = self.model[0]
|
| 969 |
dec = self._spec
|
|
|
|
| 976 |
empty_slots = kwargs.get("empty_slots")
|
| 977 |
logical = [int(s) for s in empty_slots] if empty_slots is not None else list(range(N))
|
| 978 |
pt = torch.as_tensor(page_table)
|
| 979 |
+
resume_mask = kwargs.get("prefill_resume_mask")
|
| 980 |
+
final_mask = kwargs.get("prefill_final_mask")
|
| 981 |
+
if resume_mask is not None or final_mask is not None:
|
| 982 |
+
# The plugin's TT chunk policy (tt_block_output_chunked_prefill) passes both masks on EVERY prefill step.
|
| 983 |
+
resume_mask = [bool(x) for x in (resume_mask if resume_mask is not None else [False] * N)]
|
| 984 |
+
final_mask = [bool(x) for x in (final_mask if final_mask is not None else [True] * N)]
|
| 985 |
+
if len(resume_mask) != N or len(final_mask) != N:
|
| 986 |
+
raise RuntimeError(
|
| 987 |
+
f"Qwen36DFlash: prefill masks of {len(resume_mask)}/{len(final_mask)} rows for {N} prompt rows"
|
| 988 |
+
)
|
| 989 |
+
trivial = not any(resume_mask) and all(final_mask)
|
| 990 |
+
if not trivial and not self._cp_on:
|
| 991 |
+
raise RuntimeError(
|
| 992 |
+
f"Qwen36DFlash: the runner asked for a chunked prefill (resume={resume_mask}, final={final_mask}) "
|
| 993 |
+
f"but {_CP_ENV} is off"
|
| 994 |
+
)
|
| 995 |
+
# Whole prompts with no partial held in the scratch take today's loop below, byte for byte. Anything else
|
| 996 |
+
# (a resume, an intermediate chunk, or whole prompts while a partial is held: it must be parked first)
|
| 997 |
+
# goes through the planner.
|
| 998 |
+
if not trivial or model._chunked_prefill_planner().owner is not None:
|
| 999 |
+
starts = kwargs.get("start_pos")
|
| 1000 |
+
starts = [int(starts[u]) for u in range(N)] if starts is not None else [0] * N
|
| 1001 |
+
out = self._prefill_planned(
|
| 1002 |
+
model,
|
| 1003 |
+
dec,
|
| 1004 |
+
[torch.as_tensor(tokens)[u : u + 1, : plens[u]].to(torch.int32) for u in range(N)],
|
| 1005 |
+
[pt[u].reshape(-1).clone() for u in range(N)],
|
| 1006 |
+
[self._phys[logical[u]] for u in range(N)],
|
| 1007 |
+
starts,
|
| 1008 |
+
plens,
|
| 1009 |
+
resume_mask,
|
| 1010 |
+
final_mask,
|
| 1011 |
+
)
|
| 1012 |
+
logger.info(f"Finished prefill of {N} request(s), starting decode...")
|
| 1013 |
+
return torch.cat(out, dim=0), torch.zeros(N, dtype=torch.long)
|
| 1014 |
out = []
|
| 1015 |
for u in range(N):
|
| 1016 |
phys = self._phys[logical[u]]
|
|
|
|
| 1028 |
logger.info(f"Finished prefill of {N} request(s), starting decode...")
|
| 1029 |
return torch.cat(out, dim=0), torch.zeros(N, dtype=torch.long)
|
| 1030 |
|
| 1031 |
+
def _prefill_planned(self, model, dec, prompts, rows, phys_of, starts, ends, resume_mask, final_mask, seat=True):
|
| 1032 |
+
"""One prefill call under the chunk policy: the rows in ChunkedPrefillPlanner order (a resume row first; park
|
| 1033 |
+
before any row that resets the scratch while a partial is held unparked, unpark before the resume), each through
|
| 1034 |
+
the eager spec prefill. A final row returns its host logits [1, 1, vocab] and (``seat``) marks the slot for
|
| 1035 |
+
begin() at its first decode step; an intermediate row returns zero logits, writes no slot row and becomes the
|
| 1036 |
+
scratch owner. The planner's owner is committed only after every row succeeded (tt/chunked_prefill.py).
|
| 1037 |
+
Returns the logits in call order."""
|
| 1038 |
+
N = len(prompts)
|
| 1039 |
+
planner = model._chunked_prefill_planner()
|
| 1040 |
+
first_blocks = [int(r.reshape(-1)[0]) for r in rows]
|
| 1041 |
+
plans, owner_after = planner.plan(starts, ends, resume_mask, final_mask, first_blocks)
|
| 1042 |
+
out = [None] * N
|
| 1043 |
+
owner_phys = self._cp_owner_phys
|
| 1044 |
+
done = []
|
| 1045 |
+
for p in plans:
|
| 1046 |
+
u = p.row
|
| 1047 |
+
phys = int(phys_of[u])
|
| 1048 |
+
t0 = None
|
| 1049 |
+
if _CP_TIMING:
|
| 1050 |
+
ttnn.synchronize_device(model.mesh_device)
|
| 1051 |
+
t0 = time.perf_counter()
|
| 1052 |
+
if p.park_before:
|
| 1053 |
+
model._park_gdn_scratch()
|
| 1054 |
+
if p.unpark_before:
|
| 1055 |
+
model._unpark_gdn_scratch()
|
| 1056 |
+
t1 = None
|
| 1057 |
+
if _CP_TIMING:
|
| 1058 |
+
ttnn.synchronize_device(model.mesh_device)
|
| 1059 |
+
t1 = time.perf_counter()
|
| 1060 |
+
if p.resume:
|
| 1061 |
+
if phys != self._cp_owner_phys:
|
| 1062 |
+
raise RuntimeError(
|
| 1063 |
+
f"Qwen36DFlash chunked prefill: resume row {u} is on slot {phys}, the partial prompt is in slot "
|
| 1064 |
+
f"{self._cp_owner_phys} (the plugin must keep a continuation on its state slot)"
|
| 1065 |
+
)
|
| 1066 |
+
if dec.active[phys] or self._pending[phys] is not None:
|
| 1067 |
+
raise RuntimeError(
|
| 1068 |
+
f"Qwen36DFlash chunked prefill: resume slot {phys} has a live or pending session"
|
| 1069 |
+
)
|
| 1070 |
+
else:
|
| 1071 |
+
if dec.active[phys]:
|
| 1072 |
+
# vLLM released the previous occupant before reusing its slot; close its session if not.
|
| 1073 |
+
dec.end(phys)
|
| 1074 |
+
self._carry[phys], self._stopped[phys], self._prev_tail[phys] = [], False, None
|
| 1075 |
+
self._pending[phys] = None
|
| 1076 |
+
if owner_phys == phys:
|
| 1077 |
+
owner_phys = None # a new prompt reuses the stale partial's slot (the planner parked its state)
|
| 1078 |
+
T = int(p.end)
|
| 1079 |
+
if p.final:
|
| 1080 |
+
logger.info(
|
| 1081 |
+
f"Prefilling slot {phys} up to {T} tokens (TP eager spec prefill"
|
| 1082 |
+
+ (f", resumed at {p.start})" if p.resume else ")")
|
| 1083 |
+
)
|
| 1084 |
+
else:
|
| 1085 |
+
logger.info(
|
| 1086 |
+
f"Prefilling slot {phys} tokens [{p.start}, {T}) (TP eager spec prefill, intermediate chunk)"
|
| 1087 |
+
)
|
| 1088 |
+
lt = self._spec_prefill(model, dec, phys, prompts[u][:, :T], T, rows[u], start=p.start, final=p.final)
|
| 1089 |
+
if p.final:
|
| 1090 |
+
out[u] = lt.view(1, 1, -1)
|
| 1091 |
+
if seat:
|
| 1092 |
+
self._pending[phys] = (T, rows[u])
|
| 1093 |
+
if p.resume:
|
| 1094 |
+
owner_phys = None
|
| 1095 |
+
else:
|
| 1096 |
+
out[u] = torch.zeros(1, 1, model.vocab_size, dtype=torch.float32)
|
| 1097 |
+
owner_phys = phys
|
| 1098 |
+
if _CP_TIMING:
|
| 1099 |
+
ttnn.synchronize_device(model.mesh_device)
|
| 1100 |
+
t2 = time.perf_counter()
|
| 1101 |
+
logger.info(
|
| 1102 |
+
f"[DFLASH_CP] phys={phys} start={p.start} end={T} resume={int(p.resume)} final={int(p.final)} "
|
| 1103 |
+
f"park={int(p.park_before)} unpark={int(p.unpark_before)} park_ms={(t1 - t0) * 1e3:.1f} "
|
| 1104 |
+
f"prefill_ms={(t2 - t1) * 1e3:.1f}"
|
| 1105 |
+
)
|
| 1106 |
+
done.append((phys, p.start, T, int(p.resume), int(p.final), int(p.park_before), int(p.unpark_before)))
|
| 1107 |
+
planner.owner = owner_after
|
| 1108 |
+
self._cp_owner_phys = owner_phys if owner_after is not None else None
|
| 1109 |
+
logger.info(f"Qwen36DFlash chunked rows [(phys, start, end, resume, final, park, unpark)]: {done}")
|
| 1110 |
+
return out
|
| 1111 |
+
|
| 1112 |
+
def _cp_warm_sequence(self, model, dec, rows, S0, tag):
|
| 1113 |
+
"""Exercise every program of the chunked path (park / unpark, a resumed chunk with and without logits, a
|
| 1114 |
+
tail-only resume, an intermediate chunk's no-logits exit) with the chunk policy's own orchestration: slot 0
|
| 1115 |
+
takes a long prompt in chunks, slot 1 a short rider between them. Leaves no seated slot and no scratch owner."""
|
| 1116 |
+
L0 = _PREFILL_CHUNK + S0
|
| 1117 |
+
L1 = 2 * _PREFILL_CHUNK
|
| 1118 |
+
p0 = dummy_prompt(max(L0, L1), seed=4242)
|
| 1119 |
+
rider = dummy_prompt(S0, seed=4343)
|
| 1120 |
+
seqs = [
|
| 1121 |
+
# [0, C) intermediate; a rider (parks the partial); [C, C+S0) final: tail-only resume after an unpark
|
| 1122 |
+
([p0], [rows[0]], [0], [0], [_PREFILL_CHUNK], [False], [False]),
|
| 1123 |
+
([rider], [rows[1]], [1], [0], [S0], [False], [True]),
|
| 1124 |
+
([p0], [rows[0]], [0], [_PREFILL_CHUNK], [L0], [True], [True]),
|
| 1125 |
+
# [0, C) intermediate; [C, 2C) final together with a rider: exact-multiple resume (logits from the chunk)
|
| 1126 |
+
([p0], [rows[0]], [0], [0], [_PREFILL_CHUNK], [False], [False]),
|
| 1127 |
+
([p0, rider], [rows[0], rows[1]], [0, 1], [_PREFILL_CHUNK, 0], [L1, S0], [True, False], [True, True]),
|
| 1128 |
+
]
|
| 1129 |
+
for prompts, rws, phys, st, en, rs, fn in seqs:
|
| 1130 |
+
self._prefill_planned(model, dec, prompts, rws, phys, st, en, rs, fn, seat=False)
|
| 1131 |
+
ttnn.synchronize_device(model.mesh_device)
|
| 1132 |
+
owner = model._chunked_prefill_planner().owner
|
| 1133 |
+
if owner is not None or self._cp_owner_phys is not None or any(p is not None for p in self._pending):
|
| 1134 |
+
raise RuntimeError(
|
| 1135 |
+
f"Qwen36DFlash chunked-prefill {tag}: left scratch owner {owner} / slot {self._cp_owner_phys} / "
|
| 1136 |
+
f"pending {self._pending}"
|
| 1137 |
+
)
|
| 1138 |
+
|
| 1139 |
# ------------------------------------------------------------------ decode: one block per step, all live slots
|
| 1140 |
def _no_device_sampler(self):
|
| 1141 |
"""True when no model on the mesh has the ttnn device sampler (TP=2: 124,160 logits/device > its 64K cap)."""
|
|
|
|
| 1143 |
|
| 1144 |
def decode_forward(self, *args, **kwargs):
|
| 1145 |
if _W <= 1 or not self._spec_ready():
|
| 1146 |
+
# Same reachability as in prefill_forward (a re-warm-up after release_persistent_capture); the plain decode
|
| 1147 |
+
# trace replay lives in the generator base class and has no model-side guard of its own.
|
| 1148 |
+
if self._forbid_plain:
|
| 1149 |
+
raise RuntimeError("Qwen36DFlash: plain decode after the speculative traces were captured (hang class)")
|
| 1150 |
if _W > 1 and kwargs.get("sampling_params") is not None and self._no_device_sampler():
|
| 1151 |
# Pre-arm plain step on a mesh without a device sampler: host logits (the runner samples).
|
| 1152 |
kwargs = dict(kwargs, sampling_params=None)
|
|
|
|
| 1184 |
# user), so the begins below land in the geometry in force; a mis-ordering is a decoder assert, never a
|
| 1185 |
# wrong-geometry seed. Single bucket: a host no-op.
|
| 1186 |
plan = getattr(dec, "plan", None)
|
| 1187 |
+
if _STEP_LOG:
|
| 1188 |
+
ttnn.synchronize_device(self.model[0].mesh_device)
|
| 1189 |
+
_ts = [time.perf_counter()]
|
| 1190 |
+
_bucket_before = getattr(dec, "cur_id", None)
|
| 1191 |
if plan is not None:
|
| 1192 |
live_after = {phys for phys in range(B) if dec.active[phys]}
|
| 1193 |
for i in range(Bp):
|
|
|
|
| 1197 |
if bucket != self._last_bucket:
|
| 1198 |
logger.info(f"Qwen36DFlash: verify bucket {bucket} in force ({len(live_after)} live slot(s))")
|
| 1199 |
self._last_bucket = bucket
|
| 1200 |
+
if _STEP_LOG:
|
| 1201 |
+
ttnn.synchronize_device(self.model[0].mesh_device)
|
| 1202 |
+
_ts.append(time.perf_counter())
|
| 1203 |
+
_nbegin = sum(1 for i in range(Bp) if int(poss[i]) >= 0 and self._pending[self._phys[i]] is not None)
|
| 1204 |
live_rows = []
|
| 1205 |
nosession_rows = [] # live rows without a session: EOS so the request ends (ragged: [EOS, -1, ...])
|
| 1206 |
for i in range(Bp):
|
|
|
|
| 1235 |
live_rows.append((i, phys))
|
| 1236 |
t0 = time.perf_counter() if _DEBUG else 0.0
|
| 1237 |
if _RAGGED:
|
| 1238 |
+
if _STEP_LOG:
|
| 1239 |
+
ttnn.synchronize_device(self.model[0].mesh_device)
|
| 1240 |
+
_ts.append(time.perf_counter())
|
| 1241 |
out = self._decode_ragged(dec, live_rows, Bp, t0)
|
| 1242 |
for i in nosession_rows:
|
| 1243 |
out[i, 0] = self._eos_fill
|
| 1244 |
+
if _STEP_LOG:
|
| 1245 |
+
ttnn.synchronize_device(self.model[0].mesh_device)
|
| 1246 |
+
_ts.append(time.perf_counter())
|
| 1247 |
+
logger.info(
|
| 1248 |
+
f"[DFLASH_STEP] rows={len(live_rows)} begins={_nbegin} bucket={_bucket_before}->"
|
| 1249 |
+
f"{getattr(dec, 'cur_id', None)} plan_ms={(_ts[1] - _ts[0]) * 1e3:.1f} "
|
| 1250 |
+
f"begin_ms={(_ts[2] - _ts[1]) * 1e3:.1f} step_ms={(_ts[3] - _ts[2]) * 1e3:.1f} "
|
| 1251 |
+
f"tok={int((out >= 0).sum())}"
|
| 1252 |
+
)
|
| 1253 |
return out
|
| 1254 |
# Step until every live row can fill its block. A row whose carry already holds a stop token
|
| 1255 |
# needs nothing: it emits through the stop this step (latency for a normal request); the
|
|
|
|
| 1381 |
self._spec.end(phys)
|
| 1382 |
self._pending[phys] = None
|
| 1383 |
self._carry[phys], self._stopped[phys], self._prev_tail[phys] = [], False, None
|
| 1384 |
+
if self._cp_owner_phys == phys:
|
| 1385 |
+
# An aborted / preempted partial prompt: its scratch state is dead; drop the ownership so the next prompt
|
| 1386 |
+
# does not park it (a preempted request re-prefills from 0).
|
| 1387 |
+
self._cp_owner_phys = None
|
| 1388 |
+
self.model[0]._chunked_prefill_planner().owner = None
|
| 1389 |
|
| 1390 |
def release_persistent_capture(self) -> None:
|
| 1391 |
if self._spec is None:
|
image/blobs/sha256/1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2
ADDED
|
Binary file (6.62 kB). View file
|
|
|
image/blobs/sha256/45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d
ADDED
|
Binary file (115 Bytes). View file
|
|
|
image/blobs/sha256/4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9
ADDED
|
Binary file (4.39 kB). View file
|
|
|
image/blobs/sha256/58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schemaVersion": 2,
|
| 3 |
+
"mediaType": "application/vnd.oci.image.index.v1+json",
|
| 4 |
+
"manifests": [
|
| 5 |
+
{
|
| 6 |
+
"mediaType": "application/vnd.oci.image.manifest.v1+json",
|
| 7 |
+
"digest": "sha256:8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146",
|
| 8 |
+
"size": 4501,
|
| 9 |
+
"platform": {
|
| 10 |
+
"architecture": "amd64",
|
| 11 |
+
"os": "linux"
|
| 12 |
+
}
|
| 13 |
+
},
|
| 14 |
+
{
|
| 15 |
+
"mediaType": "application/vnd.oci.image.manifest.v1+json",
|
| 16 |
+
"digest": "sha256:aba4bce132b5625823266e9eb71834b271fbc91ad4b8693c3bd34ecf068eb41e",
|
| 17 |
+
"size": 566,
|
| 18 |
+
"annotations": {
|
| 19 |
+
"vnd.docker.reference.digest": "sha256:8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146",
|
| 20 |
+
"vnd.docker.reference.type": "attestation-manifest"
|
| 21 |
+
},
|
| 22 |
+
"platform": {
|
| 23 |
+
"architecture": "unknown",
|
| 24 |
+
"os": "unknown"
|
| 25 |
+
}
|
| 26 |
+
}
|
| 27 |
+
]
|
| 28 |
+
}
|
image/blobs/sha256/70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9
ADDED
|
Binary file (1.52 kB). View file
|
|
|
image/blobs/sha256/8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schemaVersion": 2,
|
| 3 |
+
"mediaType": "application/vnd.oci.image.manifest.v1+json",
|
| 4 |
+
"config": {
|
| 5 |
+
"mediaType": "application/vnd.oci.image.config.v1+json",
|
| 6 |
+
"digest": "sha256:ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7",
|
| 7 |
+
"size": 15070
|
| 8 |
+
},
|
| 9 |
+
"layers": [
|
| 10 |
+
{
|
| 11 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 12 |
+
"digest": "sha256:98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67",
|
| 13 |
+
"size": 29751627
|
| 14 |
+
},
|
| 15 |
+
{
|
| 16 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 17 |
+
"digest": "sha256:6973a7aa7572623f5d0581fa621c6a91869112c028a42eb75fd5dad653cc7550",
|
| 18 |
+
"size": 69482149
|
| 19 |
+
},
|
| 20 |
+
{
|
| 21 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 22 |
+
"digest": "sha256:4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9",
|
| 23 |
+
"size": 4393
|
| 24 |
+
},
|
| 25 |
+
{
|
| 26 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 27 |
+
"digest": "sha256:96ec90bda118ca3b0adc835b3984bab76d0463fd2fb8d18304fa14596868c1a9",
|
| 28 |
+
"size": 9565649
|
| 29 |
+
},
|
| 30 |
+
{
|
| 31 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 32 |
+
"digest": "sha256:a092adbd4319ea0e08680de7f4cadc2d2016ed6b7bb27721f985334ec3ddf72a",
|
| 33 |
+
"size": 138948153
|
| 34 |
+
},
|
| 35 |
+
{
|
| 36 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 37 |
+
"digest": "sha256:f37009d803089ea4de72f86b7b2504b33b0e660cd05baa5b5f0b701ab907154f",
|
| 38 |
+
"size": 66293828
|
| 39 |
+
},
|
| 40 |
+
{
|
| 41 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 42 |
+
"digest": "sha256:1de484bf746902ac142693f9c9c810133dec4e4dead0d07da4049c845dfd4b33",
|
| 43 |
+
"size": 1354630576
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 47 |
+
"digest": "sha256:45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d",
|
| 48 |
+
"size": 115
|
| 49 |
+
},
|
| 50 |
+
{
|
| 51 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 52 |
+
"digest": "sha256:dbc3dfaa9d9011e424bbc5046b70645e8c6c1db180dd7e04178c962df6eb25d7",
|
| 53 |
+
"size": 139072443
|
| 54 |
+
},
|
| 55 |
+
{
|
| 56 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 57 |
+
"digest": "sha256:9f6d55da9fcb57e2f47984e8ee74c0e35973bee3b4f7428af5c4242ce4d5dcdb",
|
| 58 |
+
"size": 38513788
|
| 59 |
+
},
|
| 60 |
+
{
|
| 61 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 62 |
+
"digest": "sha256:bc3fc935ead0134d3ab22f25c6fabfa03f66fe05bc0af2c7f422d5aeb8e59374",
|
| 63 |
+
"size": 38513668
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 67 |
+
"digest": "sha256:6f8be5fbe44d20285f20ab7ddf0f118a155ae79f913784a83545e1eff48c1ef7",
|
| 68 |
+
"size": 103118556
|
| 69 |
+
},
|
| 70 |
+
{
|
| 71 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 72 |
+
"digest": "sha256:295aa94991e93759c4dfa8c455c7779e88483d53f19ae0eba70044ddae690081",
|
| 73 |
+
"size": 11906048
|
| 74 |
+
},
|
| 75 |
+
{
|
| 76 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 77 |
+
"digest": "sha256:b7f5500fbf7b7923cfcb92f65307d234f9cccd846060141ed2aa4b3e56328738",
|
| 78 |
+
"size": 1783263
|
| 79 |
+
},
|
| 80 |
+
{
|
| 81 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 82 |
+
"digest": "sha256:1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2",
|
| 83 |
+
"size": 6618
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 87 |
+
"digest": "sha256:1f118d55ecce2afdad0bcb54c6d62e873bcb62f8be7afd631d4e8d59287c3c51",
|
| 88 |
+
"size": 4030916
|
| 89 |
+
},
|
| 90 |
+
{
|
| 91 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 92 |
+
"digest": "sha256:bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7",
|
| 93 |
+
"size": 1076
|
| 94 |
+
},
|
| 95 |
+
{
|
| 96 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 97 |
+
"digest": "sha256:ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9",
|
| 98 |
+
"size": 1379
|
| 99 |
+
},
|
| 100 |
+
{
|
| 101 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 102 |
+
"digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1",
|
| 103 |
+
"size": 32
|
| 104 |
+
},
|
| 105 |
+
{
|
| 106 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 107 |
+
"digest": "sha256:70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9",
|
| 108 |
+
"size": 1520
|
| 109 |
+
},
|
| 110 |
+
{
|
| 111 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 112 |
+
"digest": "sha256:e808130579a129d44f7bb71d6c7b94b11ec446778b2f2004e540f7bf755f12de",
|
| 113 |
+
"size": 773549
|
| 114 |
+
},
|
| 115 |
+
{
|
| 116 |
+
"mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
|
| 117 |
+
"digest": "sha256:e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691",
|
| 118 |
+
"size": 4245
|
| 119 |
+
}
|
| 120 |
+
]
|
| 121 |
+
}
|
image/blobs/sha256/aba4bce132b5625823266e9eb71834b271fbc91ad4b8693c3bd34ecf068eb41e
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schemaVersion": 2,
|
| 3 |
+
"mediaType": "application/vnd.oci.image.manifest.v1+json",
|
| 4 |
+
"config": {
|
| 5 |
+
"mediaType": "application/vnd.oci.image.config.v1+json",
|
| 6 |
+
"digest": "sha256:f08a1fdc23ba07b7b094692072f62ca467a03cbe754bc49130dabddf907da186",
|
| 7 |
+
"size": 167
|
| 8 |
+
},
|
| 9 |
+
"layers": [
|
| 10 |
+
{
|
| 11 |
+
"mediaType": "application/vnd.in-toto+json",
|
| 12 |
+
"digest": "sha256:d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d",
|
| 13 |
+
"size": 1569,
|
| 14 |
+
"annotations": {
|
| 15 |
+
"in-toto.io/predicate-type": "https://slsa.dev/provenance/v0.2"
|
| 16 |
+
}
|
| 17 |
+
}
|
| 18 |
+
]
|
| 19 |
+
}
|
image/blobs/sha256/bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7
ADDED
|
Binary file (1.08 kB). View file
|
|
|
image/blobs/sha256/d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"_type":"https://in-toto.io/Statement/v0.1","predicateType":"https://slsa.dev/provenance/v0.2","subject":[{"name":"pkg:docker/tt-model/qwen3.8-27b-p150x2@build-fee9e0d35?platform=linux%2Famd64","digest":{"sha256":"8910100bf0830e36e4a1b584464742494b299868383e81f0182b73623fb5b146"}}],"predicate":{"builder":{"id":""},"buildType":"https://mobyproject.org/buildkit@v1","materials":[{"uri":"pkg:docker/docker/dockerfile@1.7-labs","digest":{"sha256":"b99fecfe00268a8b556fad7d9c37ee25d716ae08a5d7320e6d51c4dd83246894"}},{"uri":"pkg:docker/ubuntu@22.04?platform=linux%2Famd64","digest":{"sha256":"b1066385161d28ddf6bc7e7b28a9170eec11484c821d1a5150d176cbde41d7f7"}},{"uri":"pkg:docker/ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64@latest?platform=linux%2Famd64","digest":{"sha256":"82d91fd88b21f297844406897cee5bd6c9ded6e39447ff6dda698c83b33404d2"}}],"invocation":{"configSource":{"entryPoint":"Dockerfile"},"parameters":{"frontend":"gateway.v0","args":{"cmdline":"docker/dockerfile:1.7-labs","context:metalsrc":"local:metalsrc","frontend.caps":"moby.buildkit.frontend.contexts+forward","source":"docker/dockerfile:1.7-labs"},"locals":[{"name":"context"},{"name":"dockerfile"},{"name":"metalsrc"}]},"environment":{"platform":"linux/amd64"}},"metadata":{"buildInvocationID":"jrcwl84ug3518kcuik2a5bv47","buildStartedOn":"2026-10-03T05:06:49.883974165+09:00","buildFinishedOn":"2026-10-03T05:12:19.066557014+09:00","completeness":{"parameters":false,"environment":true,"materials":false},"reproducible":false,"https://mobyproject.org/buildkit@v1#metadata":{}}}}
|
image/blobs/sha256/e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691
ADDED
|
Binary file (4.25 kB). View file
|
|
|
image/blobs/sha256/ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9
ADDED
|
Binary file (1.38 kB). View file
|
|
|
image/blobs/sha256/ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"architecture":"amd64","config":{"User":"tt","Env":["PATH=/opt/tt-venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","VENV=/opt/tt-venv","VIRTUAL_ENV=/opt/tt-venv","TT_METAL_RUNTIME_ROOT=/opt/tt-metal","TT_METAL_HOME=/opt/tt-metal","PYTHONPATH=/opt/tt-metal","LD_LIBRARY_PATH=/opt/tt-metal/build/lib:/opt/openmpi-v5.0.7-ulfm/lib","EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle","TT_VLLM_BUILTIN_MODELS=0","TT_MODEL_KIND=vllm-plugin","HF_HOME=/hf","TT_METAL_CACHE=/cache","HOME=/home/tt","USER=tt","LOGNAME=tt"],"Entrypoint":["/usr/local/bin/entrypoint.sh"],"Cmd":["/usr/local/bin/serve-default.sh"],"WorkingDir":"/home/tt/work","Labels":{"org.opencontainers.image.revision":"fee9e0d35948111be29083c4eb144023d86890d8","org.opencontainers.image.version":"22.04","org.tenstorrent.tt-model":"qwen3.8-27b-p150x2","org.tenstorrent.tt-model.arch":"blackhole","org.tenstorrent.tt-model.kind":"vllm-plugin","org.tenstorrent.tt-model.plugin":"ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7","org.tenstorrent.tt-model.profiles":"tp2,p1d1,tp2-dflash2","org.tenstorrent.tt-model.repo":"changh95/qwen3.8-27b-p150x2","org.tenstorrent.tt-model.tt-metal":"v0.79.0-dev20260903-138-gfee9e0d359","org.tenstorrent.tt-model.weights":"Qwen/Qwen3.8-27B"},"ArgsEscaped":true},"created":"2026-10-03T05:11:03.366502764+09:00","history":[{"created":"2026-09-24T22:24:04.688908979Z","created_by":"/bin/sh -c #(nop) ARG RELEASE","empty_layer":true},{"created":"2026-09-24T22:24:04.723527376Z","created_by":"/bin/sh -c #(nop) ARG LAUNCHPAD_BUILD_ARCH","empty_layer":true},{"created":"2026-09-24T22:24:04.751455651Z","created_by":"/bin/sh -c #(nop) LABEL org.opencontainers.image.version=22.04","empty_layer":true},{"created":"2026-09-24T22:24:06.837737736Z","created_by":"/bin/sh -c #(nop) ADD file:b0bf3f64519bf10a51e00d4f8ab9c8693620659a0c4a4a5090a9765e754e1d1c in / "},{"created":"2026-09-24T22:24:07.239071539Z","created_by":"/bin/sh -c #(nop) CMD [\"/bin/bash\"]","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG OMPI_DIR=/opt/openmpi-v5.0.7-ulfm","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG EXTRA_MODELS_DIR=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG TT_MODEL_KIND","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_NAME","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_REPO","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_WEIGHTS","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_ARCH","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_PROFILES","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_TT_METAL_SHA","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_TT_METAL_DESCRIBE","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"ARG MODEL_PLUGIN_SHA","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:07:17.310281888+09:00","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 /bin/sh -c apt-get update \u0026\u0026 apt-get install -y --no-install-recommends libhwloc15 libnuma1 libatomic1 libudev1 libcap2 zlib1g libmpc3 libmpfr6 libgmp10 libzstd1 libevent-core-2.1-7 libevent-pthreads-2.1-7 libgl1 libsndfile1 ca-certificates \u0026\u0026 apt-get clean \u0026\u0026 rm -rf /var/lib/apt/lists/* # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:07:17.417705026+09:00","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 /bin/sh -c existing=\"$(getent passwd 1000 | cut -d: -f1)\" \u0026\u0026 if [ -n \"$existing\" ]; then userdel -r \"$existing\" 2\u003e/dev/null || userdel \"$existing\"; fi \u0026\u0026 useradd --uid 1000 --create-home --home-dir /home/tt --shell /bin/bash tt \u0026\u0026 mkdir -p /home/tt/work/logs /cache /opt/tt-metal \u0026\u0026 chown -R tt:tt /home/tt /cache /opt/tt-metal \u0026\u0026 chmod 1777 /home/tt /home/tt/work /home/tt/work/logs /cache # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:27.718820963+09:00","created_by":"COPY /opt/openmpi-v5.0.7-ulfm /opt/openmpi-v5.0.7-ulfm # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:30.575915892+09:00","created_by":"COPY /opt/tenstorrent /opt/tenstorrent # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:31.930224101+09:00","created_by":"COPY /usr/local/share/uv /usr/local/share/uv # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:49.821667975+09:00","created_by":"COPY /opt/tt-venv /opt/tt-venv # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:50.018305462+09:00","created_by":"COPY /opt/vllm /opt/vllm # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:50.833470062+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/runtime /opt/tt-metal/runtime # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:51.090162095+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/build_Release /opt/tt-metal/build_Release # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:51.342451117+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/build /opt/tt-metal/build # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.162323006+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/tt_metal /opt/tt-metal/tt_metal # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.561349473+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/ttnn /opt/tt-metal/ttnn # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.636881027+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/tools /opt/tt-metal/tools # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.679150816+09:00","created_by":"COPY --chown=tt:tt /opt/tt-metal/setup.py /opt/tt-metal/pyproject.toml /opt/tt-metal/ # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.763642292+09:00","created_by":"COPY --chown=tt:tt code/ /opt/tt-metal/ # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.803147995+09:00","created_by":"COPY entrypoint.sh /usr/local/bin/entrypoint.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"COPY --chmod=0755 serve-default.sh /usr/local/bin/serve-default.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV VENV=/opt/tt-venv","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV VIRTUAL_ENV=/opt/tt-venv","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV PATH=/opt/tt-venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_METAL_RUNTIME_ROOT=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_METAL_HOME=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV PYTHONPATH=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV LD_LIBRARY_PATH=/opt/tt-metal/build/lib:/opt/openmpi-v5.0.7-ulfm/lib","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ARG TT_VLLM_BUILTIN_MODELS=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_VLLM_BUILTIN_MODELS=0","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_MODEL_KIND=vllm-plugin","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV HF_HOME=/hf","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV TT_METAL_CACHE=/cache","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV HOME=/home/tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV USER=tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"ENV LOGNAME=tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.820001794+09:00","created_by":"USER tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:09:52.857467483+09:00","created_by":"WORKDIR /home/tt/work","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:09:52.893157111+09:00","created_by":"COPY verify.sh /ctx/verify.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:11:03.148431143+09:00","created_by":"RUN |12 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 TT_VLLM_BUILTIN_MODELS=0 /bin/sh -c bash /ctx/verify.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:11:03.148431143+09:00","created_by":"USER root","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"RUN |12 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR=/opt/tt-metal/models/demos/blackhole/qwen36/vllm_bundle TT_MODEL_KIND=vllm-plugin MODEL_NAME=qwen3.8-27b-p150x2 MODEL_REPO=changh95/qwen3.8-27b-p150x2 MODEL_WEIGHTS=Qwen/Qwen3.8-27B MODEL_ARCH=blackhole MODEL_PROFILES=tp2,p1d1,tp2-dflash2 MODEL_TT_METAL_SHA=fee9e0d35948111be29083c4eb144023d86890d8 MODEL_TT_METAL_DESCRIBE=v0.79.0-dev20260903-138-gfee9e0d359 MODEL_PLUGIN_SHA=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7 TT_VLLM_BUILTIN_MODELS=0 /bin/sh -c chmod -R a+rwX /home/tt # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"USER tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"LABEL org.tenstorrent.tt-model=qwen3.8-27b-p150x2 org.tenstorrent.tt-model.repo=changh95/qwen3.8-27b-p150x2 org.tenstorrent.tt-model.weights=Qwen/Qwen3.8-27B org.tenstorrent.tt-model.arch=blackhole org.tenstorrent.tt-model.kind=vllm-plugin org.tenstorrent.tt-model.profiles=tp2,p1d1,tp2-dflash2 org.opencontainers.image.revision=fee9e0d35948111be29083c4eb144023d86890d8 org.tenstorrent.tt-model.tt-metal=v0.79.0-dev20260903-138-gfee9e0d359 org.tenstorrent.tt-model.plugin=ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"ENTRYPOINT [\"/usr/local/bin/entrypoint.sh\"]","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-03T05:11:03.366502764+09:00","created_by":"CMD [\"/usr/local/bin/serve-default.sh\"]","comment":"buildkit.dockerfile.v0","empty_layer":true}],"os":"linux","rootfs":{"type":"layers","diff_ids":["sha256:cbaf391e670933f026c05c7dedec3faf8f194b6c2d83b0d5694ad68323c8077c","sha256:cfccc3cbc964861981c069fce1701ffa33f052675c244b7a422825765dd9a63f","sha256:d4f545a332922c93c1dc239e25127c4ae7271a58b97a7ab079d369fd9f4dbc35","sha256:e093e8ec226ffd10480fc5a2bcbedf619b6138801ed0534a40f5cd68955a73d6","sha256:78e44c64fa49347ea5c050c2fdab481b14c4b6dc164a1668c642a97bea1e7685","sha256:2137fc23bb3fe89686694fc573d254802babc0361b3dce51dff7fd83b2b3848c","sha256:22c669305206fddfb75fcec816f7eafc89909aca78e0d47f40040ee4b3216733","sha256:31113ed63efe6fe971ea06eeeda5f10a9e62523bb5fd69dc21221e771ef77739","sha256:8e82e1780c58e70a07294d5525970abd23e780f29a1868d0c27650a435475e31","sha256:2251b873b9d6a96305d4051751caa3f25918ff1c934fc8784a4669dd52c0eff4","sha256:a68c563178b2e86d35fc7b382791a345ad3ed51b0bd5ca5f098109dff40f0ddf","sha256:7870fda76a62b7f08224aa010e777daca52828e7dfd77dad81725784043152a3","sha256:05faaf4debc0728ec662239967cc8aa14ce5b905e0fd7abfc0bb2327126d1414","sha256:93ebc8896d0c5c68f7350d99df7a3fcabd59c1b9d77d518b858067b5f08ed26b","sha256:b73fff02849ac52fe40f953a7aaaa80a1b329d491ab5ebb7f139c1d4e58f18ac","sha256:050f9922591a6b58429eead751bfe7f435181720c78f6d5368c905e375e46f03","sha256:210aee8d627bdba2b75c940d0df6e1e66595d23d84a522165bcf6ce0900e4b30","sha256:55183c62ebd7eb1f6563f266b6e9793ecb66e65bbf420a9bb356836d54f951a7","sha256:5f70bf18a086007016e948b04aed3b82103a36bea41755b6cddfaf10ace3c6ef","sha256:38d87bfaad3b918b0332db9c70eabe949f28d527ed8bdb8ef9c96809ba7a174d","sha256:5eefbd7a03694f21269829dfb4e26281e6fbc6d4dd25688f6bcbd46bb3e0e726","sha256:41b77a65171ea10410caa54ffc892427a7c0adfecb0e3c1ee4cb5bae8e1e9f19"]}}
|
image/blobs/sha256/f08a1fdc23ba07b7b094692072f62ca467a03cbe754bc49130dabddf907da186
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"architecture":"unknown","os":"unknown","config":{},"rootfs":{"type":"layers","diff_ids":["sha256:d5e318f14136d66c70711a17d2240573eec1850778ef2a323e4095b57c292e8d"]}}
|
image/index.json
CHANGED
|
@@ -1 +1 @@
|
|
| 1 |
-
{"schemaVersion":2,"mediaType":"application/vnd.oci.image.index.v1+json","manifests":[{"mediaType":"application/vnd.oci.image.index.v1+json","digest":"sha256:
|
|
|
|
| 1 |
+
{"schemaVersion":2,"mediaType":"application/vnd.oci.image.index.v1+json","manifests":[{"mediaType":"application/vnd.oci.image.index.v1+json","digest":"sha256:58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2","size":856,"annotations":{"io.containerd.image.name":"docker.io/tt-model/qwen3.8-27b-p150x2:58b3a5c0c045","org.opencontainers.image.ref.name":"58b3a5c0c045"}}]}
|
image/manifest.json
CHANGED
|
@@ -1 +1 @@
|
|
| 1 |
-
[{"Config":"blobs/sha256/
|
|
|
|
| 1 |
+
[{"Config":"blobs/sha256/ef21e6b1c8ba706956e6088db50adc82a55b90940ea91abe335970baa4ae91c7","RepoTags":["tt-model/qwen3.8-27b-p150x2:58b3a5c0c045"],"Layers":["blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67","blobs/sha256/6973a7aa7572623f5d0581fa621c6a91869112c028a42eb75fd5dad653cc7550","blobs/sha256/4c34c031a3d45b0e39e28a17b4c8e305961b05b44af5a12ca2c22472d49bd4f9","blobs/sha256/96ec90bda118ca3b0adc835b3984bab76d0463fd2fb8d18304fa14596868c1a9","blobs/sha256/a092adbd4319ea0e08680de7f4cadc2d2016ed6b7bb27721f985334ec3ddf72a","blobs/sha256/f37009d803089ea4de72f86b7b2504b33b0e660cd05baa5b5f0b701ab907154f","blobs/sha256/1de484bf746902ac142693f9c9c810133dec4e4dead0d07da4049c845dfd4b33","blobs/sha256/45a1df39c7735549366a789147dd306a2637fe3c1102a7866da751c15fd3568d","blobs/sha256/dbc3dfaa9d9011e424bbc5046b70645e8c6c1db180dd7e04178c962df6eb25d7","blobs/sha256/9f6d55da9fcb57e2f47984e8ee74c0e35973bee3b4f7428af5c4242ce4d5dcdb","blobs/sha256/bc3fc935ead0134d3ab22f25c6fabfa03f66fe05bc0af2c7f422d5aeb8e59374","blobs/sha256/6f8be5fbe44d20285f20ab7ddf0f118a155ae79f913784a83545e1eff48c1ef7","blobs/sha256/295aa94991e93759c4dfa8c455c7779e88483d53f19ae0eba70044ddae690081","blobs/sha256/b7f5500fbf7b7923cfcb92f65307d234f9cccd846060141ed2aa4b3e56328738","blobs/sha256/1f69e5d6a3141583d5a919514f57a0a73faba99582478112a330dde35015a1b2","blobs/sha256/1f118d55ecce2afdad0bcb54c6d62e873bcb62f8be7afd631d4e8d59287c3c51","blobs/sha256/bfc63b12aed80f89c4ba5491d9f0dfe9c0f4a5ac369f7355de66b5154b7635f7","blobs/sha256/ebd6edcfe3338809e3f5e8909b2f9fdb14482e23c30d6c402a75003bc506e7d9","blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1","blobs/sha256/70715e53bc38e6baf9b7e5372157a0ed81d2496314c10da6f660defba305b1a9","blobs/sha256/e808130579a129d44f7bb71d6c7b94b11ec446778b2f2004e540f7bf755f12de","blobs/sha256/e1f43de461348342e308d2bf3ce941d508ab55c6b4050f4c1b1012fa5df02691"]}]
|
requirements.lock
CHANGED
|
@@ -4,26 +4,26 @@ aiohttp==3.14.3
|
|
| 4 |
aiosignal==1.4.0
|
| 5 |
annotated-doc==0.0.5
|
| 6 |
annotated-types==0.8.0
|
| 7 |
-
anthropic==1.
|
| 8 |
anyio==4.15.1
|
| 9 |
apache-tvm-ffi==0.1.10
|
| 10 |
astor==0.8.1
|
| 11 |
attrs==26.1.0
|
| 12 |
-
blake3==1.0.
|
| 13 |
cachetools==7.2.0
|
| 14 |
-
cbor2==6.1.
|
| 15 |
certifi==2026.7.22
|
| 16 |
cffi==2.1.1
|
| 17 |
cfgv==3.5.0
|
| 18 |
-
charset-normalizer==3.5.
|
| 19 |
click==8.5.0
|
| 20 |
cloudpickle==3.1.2
|
| 21 |
compressed-tensors==0.17.0
|
| 22 |
contourpy==1.3.3
|
| 23 |
-
cryptography==50.0.
|
| 24 |
cuda-bindings==13.4.3
|
| 25 |
cuda-core==1.2.1
|
| 26 |
-
cuda-pathfinder==1.8.
|
| 27 |
cuda-python==13.4.1
|
| 28 |
cuda-tile==1.6.0
|
| 29 |
cycler==0.12.1
|
|
@@ -43,12 +43,12 @@ fastapi-cli==0.0.32
|
|
| 43 |
fastapi-cloud-cli==0.26.0
|
| 44 |
fastar==0.12.0
|
| 45 |
fastsafetensors==0.4.0
|
| 46 |
-
filelock==4.0.
|
| 47 |
flashinfer-python==0.6.14
|
| 48 |
-
fonttools==4.66.
|
| 49 |
frozenlist==1.8.0
|
| 50 |
fsspec==2026.9.0
|
| 51 |
-
googleapis-common-protos==1.75.
|
| 52 |
graphviz==0.21
|
| 53 |
grpcio==1.84.0
|
| 54 |
h11==0.16.0
|
|
@@ -60,7 +60,7 @@ httpx==0.28.1
|
|
| 60 |
httpx2==2.13.1
|
| 61 |
huggingface_hub==1.33.0
|
| 62 |
humming-kernels==0.1.10
|
| 63 |
-
identify==2.6.
|
| 64 |
idna==3.20
|
| 65 |
ijson==3.5.1
|
| 66 |
interegular==0.3.3
|
|
@@ -87,11 +87,11 @@ mistral_common==1.12.0
|
|
| 87 |
ml_dtypes==0.5.4
|
| 88 |
model-hosting-container-standards==0.1.16
|
| 89 |
mpmath==1.3.0
|
| 90 |
-
msgspec==0.
|
| 91 |
multidict==6.9.1
|
| 92 |
networkx==3.7
|
| 93 |
ninja==1.13.2
|
| 94 |
-
nodeenv==1.
|
| 95 |
numba==0.65.0
|
| 96 |
numpy==1.26.4
|
| 97 |
nvidia-cuda-cccl==13.3.4.3.1
|
|
@@ -109,7 +109,7 @@ nvidia-cutlass-dsl-libs-cu13==4.6.0
|
|
| 109 |
nvidia-ml-py==13.615.71
|
| 110 |
nvidia-nvvm==13.4.92
|
| 111 |
nvtx==0.2.15
|
| 112 |
-
openai==3.
|
| 113 |
openai-harmony==0.0.8
|
| 114 |
opencv-python-headless==4.11.0.86
|
| 115 |
opentelemetry-api==1.45.0
|
|
@@ -128,7 +128,7 @@ packaging==26.3
|
|
| 128 |
pandas==3.0.6
|
| 129 |
partial-json-parser==0.2.1.1.post7
|
| 130 |
pillow==12.3.0
|
| 131 |
-
platformdirs==4.
|
| 132 |
pre_commit==4.6.2
|
| 133 |
prometheus-fastapi-instrumentator==8.1.0
|
| 134 |
prometheus_client==0.26.0
|
|
@@ -145,20 +145,20 @@ pydantic-settings==2.15.0
|
|
| 145 |
pydantic_core==2.46.5
|
| 146 |
pyelftools==0.33
|
| 147 |
Pygments==2.21.0
|
| 148 |
-
PyJWT==2.15.
|
| 149 |
pyluwen==0.9.0
|
| 150 |
pynvvideocodec==2.0.4
|
| 151 |
pyparsing==3.3.3
|
| 152 |
python-dateutil==2.9.0.post0
|
| 153 |
python-discovery==1.6.1
|
| 154 |
-
python-dotenv==1.2.
|
| 155 |
python-json-logger==4.2.0
|
| 156 |
python-multipart==0.0.32
|
| 157 |
PyYAML==6.0.3
|
| 158 |
pyzmq==27.2.0
|
| 159 |
quack-kernels==0.6.3
|
| 160 |
referencing==0.37.0
|
| 161 |
-
regex==2026.9.
|
| 162 |
requests==2.34.2
|
| 163 |
rich==15.0.0
|
| 164 |
rich-toolkit==0.20.5
|
|
@@ -167,14 +167,14 @@ rpds-py==2026.6.3
|
|
| 167 |
safetensors==0.8.0
|
| 168 |
seaborn==0.13.2
|
| 169 |
sentencepiece==0.2.2
|
| 170 |
-
sentry-sdk==2.
|
| 171 |
-
setproctitle==1.3.
|
| 172 |
setuptools==80.10.2
|
| 173 |
setuptools-scm==8.1.0
|
| 174 |
shellingham==1.5.4
|
| 175 |
six==1.17.0
|
| 176 |
sniffio==1.3.1
|
| 177 |
-
sse-starlette==3.
|
| 178 |
starlette==1.7.0
|
| 179 |
supervisor==4.3.0
|
| 180 |
sympy==1.14.0
|
|
@@ -189,26 +189,22 @@ tokenspeed-triton==3.8.10.post20260920
|
|
| 189 |
tomli==2.4.1
|
| 190 |
torch==2.11.0+cpu
|
| 191 |
torch_c_dlpack_ext==0.1.5
|
| 192 |
-
torchcodec==0.
|
| 193 |
torchvision==0.26.0+cpu
|
| 194 |
tqdm==4.70.1
|
| 195 |
-
transformers==5.
|
| 196 |
triton==3.8.0
|
| 197 |
truststore==0.10.4
|
| 198 |
tt-smi==6.6.0
|
| 199 |
tt-tools-common==1.6.0
|
| 200 |
tt-umd==0.9.11
|
| 201 |
-
ttnn==0.65.2.dev9815
|
| 202 |
-
ttnn==0.75.0rc10.dev1229+gf6deef232f7
|
| 203 |
typer==0.27.2
|
| 204 |
typing-inspection==0.4.4
|
| 205 |
typing_extensions==4.16.0
|
| 206 |
urllib3==2.8.0
|
| 207 |
uvicorn==0.54.0
|
| 208 |
-
uvloop==0.
|
| 209 |
-
virtualenv==21.
|
| 210 |
-
vllm==0.26.0+empty
|
| 211 |
-
vllm-tt-plugin==0.1.0
|
| 212 |
watchfiles==1.3.0
|
| 213 |
websockets==17.1
|
| 214 |
wheel==0.48.0
|
|
|
|
| 4 |
aiosignal==1.4.0
|
| 5 |
annotated-doc==0.0.5
|
| 6 |
annotated-types==0.8.0
|
| 7 |
+
anthropic==1.11.0
|
| 8 |
anyio==4.15.1
|
| 9 |
apache-tvm-ffi==0.1.10
|
| 10 |
astor==0.8.1
|
| 11 |
attrs==26.1.0
|
| 12 |
+
blake3==1.0.10
|
| 13 |
cachetools==7.2.0
|
| 14 |
+
cbor2==6.1.5
|
| 15 |
certifi==2026.7.22
|
| 16 |
cffi==2.1.1
|
| 17 |
cfgv==3.5.0
|
| 18 |
+
charset-normalizer==3.5.2
|
| 19 |
click==8.5.0
|
| 20 |
cloudpickle==3.1.2
|
| 21 |
compressed-tensors==0.17.0
|
| 22 |
contourpy==1.3.3
|
| 23 |
+
cryptography==50.0.2
|
| 24 |
cuda-bindings==13.4.3
|
| 25 |
cuda-core==1.2.1
|
| 26 |
+
cuda-pathfinder==1.8.3
|
| 27 |
cuda-python==13.4.1
|
| 28 |
cuda-tile==1.6.0
|
| 29 |
cycler==0.12.1
|
|
|
|
| 43 |
fastapi-cloud-cli==0.26.0
|
| 44 |
fastar==0.12.0
|
| 45 |
fastsafetensors==0.4.0
|
| 46 |
+
filelock==4.0.9
|
| 47 |
flashinfer-python==0.6.14
|
| 48 |
+
fonttools==4.66.1
|
| 49 |
frozenlist==1.8.0
|
| 50 |
fsspec==2026.9.0
|
| 51 |
+
googleapis-common-protos==1.75.5
|
| 52 |
graphviz==0.21
|
| 53 |
grpcio==1.84.0
|
| 54 |
h11==0.16.0
|
|
|
|
| 60 |
httpx2==2.13.1
|
| 61 |
huggingface_hub==1.33.0
|
| 62 |
humming-kernels==0.1.10
|
| 63 |
+
identify==2.6.20
|
| 64 |
idna==3.20
|
| 65 |
ijson==3.5.1
|
| 66 |
interegular==0.3.3
|
|
|
|
| 87 |
ml_dtypes==0.5.4
|
| 88 |
model-hosting-container-standards==0.1.16
|
| 89 |
mpmath==1.3.0
|
| 90 |
+
msgspec==0.22.0
|
| 91 |
multidict==6.9.1
|
| 92 |
networkx==3.7
|
| 93 |
ninja==1.13.2
|
| 94 |
+
nodeenv==1.11.0
|
| 95 |
numba==0.65.0
|
| 96 |
numpy==1.26.4
|
| 97 |
nvidia-cuda-cccl==13.3.4.3.1
|
|
|
|
| 109 |
nvidia-ml-py==13.615.71
|
| 110 |
nvidia-nvvm==13.4.92
|
| 111 |
nvtx==0.2.15
|
| 112 |
+
openai==3.24.0
|
| 113 |
openai-harmony==0.0.8
|
| 114 |
opencv-python-headless==4.11.0.86
|
| 115 |
opentelemetry-api==1.45.0
|
|
|
|
| 128 |
pandas==3.0.6
|
| 129 |
partial-json-parser==0.2.1.1.post7
|
| 130 |
pillow==12.3.0
|
| 131 |
+
platformdirs==4.12.2
|
| 132 |
pre_commit==4.6.2
|
| 133 |
prometheus-fastapi-instrumentator==8.1.0
|
| 134 |
prometheus_client==0.26.0
|
|
|
|
| 145 |
pydantic_core==2.46.5
|
| 146 |
pyelftools==0.33
|
| 147 |
Pygments==2.21.0
|
| 148 |
+
PyJWT==2.15.1
|
| 149 |
pyluwen==0.9.0
|
| 150 |
pynvvideocodec==2.0.4
|
| 151 |
pyparsing==3.3.3
|
| 152 |
python-dateutil==2.9.0.post0
|
| 153 |
python-discovery==1.6.1
|
| 154 |
+
python-dotenv==1.2.4
|
| 155 |
python-json-logger==4.2.0
|
| 156 |
python-multipart==0.0.32
|
| 157 |
PyYAML==6.0.3
|
| 158 |
pyzmq==27.2.0
|
| 159 |
quack-kernels==0.6.3
|
| 160 |
referencing==0.37.0
|
| 161 |
+
regex==2026.9.29
|
| 162 |
requests==2.34.2
|
| 163 |
rich==15.0.0
|
| 164 |
rich-toolkit==0.20.5
|
|
|
|
| 167 |
safetensors==0.8.0
|
| 168 |
seaborn==0.13.2
|
| 169 |
sentencepiece==0.2.2
|
| 170 |
+
sentry-sdk==2.71.0
|
| 171 |
+
setproctitle==1.3.8
|
| 172 |
setuptools==80.10.2
|
| 173 |
setuptools-scm==8.1.0
|
| 174 |
shellingham==1.5.4
|
| 175 |
six==1.17.0
|
| 176 |
sniffio==1.3.1
|
| 177 |
+
sse-starlette==3.5.0
|
| 178 |
starlette==1.7.0
|
| 179 |
supervisor==4.3.0
|
| 180 |
sympy==1.14.0
|
|
|
|
| 189 |
tomli==2.4.1
|
| 190 |
torch==2.11.0+cpu
|
| 191 |
torch_c_dlpack_ext==0.1.5
|
| 192 |
+
torchcodec==0.17.0+cpu
|
| 193 |
torchvision==0.26.0+cpu
|
| 194 |
tqdm==4.70.1
|
| 195 |
+
transformers==5.18.0
|
| 196 |
triton==3.8.0
|
| 197 |
truststore==0.10.4
|
| 198 |
tt-smi==6.6.0
|
| 199 |
tt-tools-common==1.6.0
|
| 200 |
tt-umd==0.9.11
|
|
|
|
|
|
|
| 201 |
typer==0.27.2
|
| 202 |
typing-inspection==0.4.4
|
| 203 |
typing_extensions==4.16.0
|
| 204 |
urllib3==2.8.0
|
| 205 |
uvicorn==0.54.0
|
| 206 |
+
uvloop==0.23.0
|
| 207 |
+
virtualenv==21.14.5
|
|
|
|
|
|
|
| 208 |
watchfiles==1.3.0
|
| 209 |
websockets==17.1
|
| 210 |
wheel==0.48.0
|
tt_kernel_manifest.json
CHANGED
|
@@ -1,12 +1,12 @@
|
|
| 1 |
{
|
| 2 |
"schema_version": "5.1",
|
| 3 |
"name": "qwen3.8-27b-p150x2",
|
| 4 |
-
"tt_metal_version": "0.65.2.
|
| 5 |
"arch": "blackhole",
|
| 6 |
"device_count": 2,
|
| 7 |
"producer": {
|
| 8 |
"tt_kernel_version": "0.1.0",
|
| 9 |
-
"created_at": "2026-
|
| 10 |
"hostname": "tt-quietbox"
|
| 11 |
},
|
| 12 |
"weights": {
|
|
@@ -27,8 +27,8 @@
|
|
| 27 |
"image": {
|
| 28 |
"registry": "hf",
|
| 29 |
"repository": "qwen3.8-27b-p150x2",
|
| 30 |
-
"tag": "tt-model/qwen3.8-27b-p150x2:
|
| 31 |
-
"digest": "sha256:
|
| 32 |
},
|
| 33 |
"kind": "vllm-plugin",
|
| 34 |
"runtime": {
|
|
@@ -211,26 +211,26 @@
|
|
| 211 |
"from pathlib import Path; assert Path('/opt/tt-metal/models/model_trace_region_sizes.yaml').is_file()"
|
| 212 |
],
|
| 213 |
"built": {
|
| 214 |
-
"image": "tt-model/qwen3.8-27b-p150x2:
|
| 215 |
"repo": "changh95/qwen3.8-27b-p150x2",
|
| 216 |
"tt_model_version": "0.1.0",
|
| 217 |
-
"created_at": "2026-
|
| 218 |
"tt_metal": {
|
| 219 |
-
"sha": "
|
| 220 |
-
"describe": "v0.79.0-dev20260903-
|
| 221 |
"dirty": false,
|
| 222 |
-
"scm_version": "0.65.2.
|
| 223 |
"mode": "local",
|
| 224 |
"remote": "https://github.com/tenstorrent/tt-metal.git",
|
| 225 |
"branch": "qwen36-p150x2",
|
| 226 |
"pushed": true
|
| 227 |
},
|
| 228 |
-
"code_sha256": "
|
| 229 |
"plugin": {
|
| 230 |
"sha": "ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7",
|
| 231 |
"repo": "https://github.com/changh95/vllm-tt-plugin"
|
| 232 |
},
|
| 233 |
-
"image_digest": "sha256:
|
| 234 |
}
|
| 235 |
}
|
| 236 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"schema_version": "5.1",
|
| 3 |
"name": "qwen3.8-27b-p150x2",
|
| 4 |
+
"tt_metal_version": "0.65.2.dev9856",
|
| 5 |
"arch": "blackhole",
|
| 6 |
"device_count": 2,
|
| 7 |
"producer": {
|
| 8 |
"tt_kernel_version": "0.1.0",
|
| 9 |
+
"created_at": "2026-10-02T20:12:52.368580+00:00",
|
| 10 |
"hostname": "tt-quietbox"
|
| 11 |
},
|
| 12 |
"weights": {
|
|
|
|
| 27 |
"image": {
|
| 28 |
"registry": "hf",
|
| 29 |
"repository": "qwen3.8-27b-p150x2",
|
| 30 |
+
"tag": "tt-model/qwen3.8-27b-p150x2:58b3a5c0c045",
|
| 31 |
+
"digest": "sha256:58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2"
|
| 32 |
},
|
| 33 |
"kind": "vllm-plugin",
|
| 34 |
"runtime": {
|
|
|
|
| 211 |
"from pathlib import Path; assert Path('/opt/tt-metal/models/model_trace_region_sizes.yaml').is_file()"
|
| 212 |
],
|
| 213 |
"built": {
|
| 214 |
+
"image": "tt-model/qwen3.8-27b-p150x2:58b3a5c0c045",
|
| 215 |
"repo": "changh95/qwen3.8-27b-p150x2",
|
| 216 |
"tt_model_version": "0.1.0",
|
| 217 |
+
"created_at": "2026-10-02T20:06:49+00:00",
|
| 218 |
"tt_metal": {
|
| 219 |
+
"sha": "fee9e0d35948111be29083c4eb144023d86890d8",
|
| 220 |
+
"describe": "v0.79.0-dev20260903-138-gfee9e0d359",
|
| 221 |
"dirty": false,
|
| 222 |
+
"scm_version": "0.65.2.dev9856",
|
| 223 |
"mode": "local",
|
| 224 |
"remote": "https://github.com/tenstorrent/tt-metal.git",
|
| 225 |
"branch": "qwen36-p150x2",
|
| 226 |
"pushed": true
|
| 227 |
},
|
| 228 |
+
"code_sha256": "b0c865d431db3b437ab3e681b01835c22f2f116164a2fe2e30c635fd43e0d6b9",
|
| 229 |
"plugin": {
|
| 230 |
"sha": "ab5f7f3d795ee519b5d96f41f31cc1e0f635d6a7",
|
| 231 |
"repo": "https://github.com/changh95/vllm-tt-plugin"
|
| 232 |
},
|
| 233 |
+
"image_digest": "sha256:58b3a5c0c045a4853ae1174db6e0e0e77393131f099295b059238f273d38b5c2"
|
| 234 |
}
|
| 235 |
}
|
| 236 |
}
|