Image-Text-to-Text
Transformers
Safetensors
qwen3_5
nvfp4
modelopt
vllm
speculative-decoding
mtp
dgx-spark
conversational
8-bit precision
Instructions to use PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP") model = AutoModelForMultimodalLM.from_pretrained("PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP
- SGLang
How to use PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP with Docker Model Runner:
docker model run hf.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP
Name the checkpoint correctly: Fable-B-F451-NVFP4 (F451 is part of the source merge identity)
Browse files
README.md
CHANGED
|
@@ -1,256 +1,254 @@
|
|
| 1 |
-
---
|
| 2 |
-
base_model:
|
| 3 |
-
- nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
|
| 4 |
-
library_name: transformers
|
| 5 |
-
pipeline_tag: image-text-to-text
|
| 6 |
-
tags:
|
| 7 |
-
- nvfp4
|
| 8 |
-
- modelopt
|
| 9 |
-
- vllm
|
| 10 |
-
- speculative-decoding
|
| 11 |
-
- mtp
|
| 12 |
-
- dgx-spark
|
| 13 |
-
---
|
| 14 |
-
|
| 15 |
-
# Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP
|
| 16 |
-
|
| 17 |
-
NVFP4 quantisation of
|
| 18 |
-
[nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451](https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451),
|
| 19 |
-
built to run on a 128 GB DGX Spark,
|
| 20 |
-
intact**.
|
| 21 |
-
|
| 22 |
-
This is the same quantisation as
|
| 23 |
-
[PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4)
|
| 24 |
-
plus one 105 MB tensor that the other repo is missing. **If you want speculative
|
| 25 |
-
decoding, use this one.** If you do not, either works.
|
| 26 |
-
|
| 27 |
-
| | this repo (`-MTP`) | the plain `-NVFP4` repo |
|
| 28 |
-
|---|---|---|
|
| 29 |
-
| tensors | **2399** | 2398 |
|
| 30 |
-
| `mtp.fc.weight` | **present** | absent |
|
| 31 |
-
| runs `--speculative-config method=mtp` | **yes** | loads, then drafts garbage |
|
| 32 |
-
| size | 20.6 GB | 20.5 GB |
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
`mtp.fc
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
``
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
--
|
| 72 |
-
--
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
```
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
**
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
| |
|
| 120 |
-
|
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
|
| 137 |
-
|
|
| 138 |
-
|
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
-
|
| 165 |
-
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
|
| 182 |
-
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
|
| 187 |
-
|
| 188 |
-
|
| 189 |
-
|
| 190 |
-
|
| 191 |
-
|
| 192 |
-
|
| 193 |
-
|
| 194 |
-
|
| 195 |
-
|
| 196 |
-
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
|
| 200 |
-
|
| 201 |
-
|
| 202 |
-
|
| 203 |
-
|
| 204 |
-
|
| 205 |
-
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
| 210 |
-
```
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
|
| 217 |
-
|
| 218 |
-
|
| 219 |
-
|
| 220 |
-
|
| 221 |
-
|
| 222 |
-
|
| 223 |
-
|
| 224 |
-
|
| 225 |
-
|
| 226 |
-
|
| 227 |
-
|
| 228 |
-
|
| 229 |
-
|
| 230 |
-
|
| 231 |
-
|
| 232 |
-
|
| 233 |
-
|
| 234 |
-
|
| 235 |
-
|
| 236 |
-
|
| 237 |
-
|
| 238 |
-
|
| 239 |
-
|
| 240 |
-
|
| 241 |
-
-
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
|
| 245 |
-
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
|
| 249 |
-
|
| 250 |
-
|
| 251 |
-
|
| 252 |
-
- **
|
| 253 |
-
|
| 254 |
-
|
| 255 |
-
- **vLLM** - the serving engine, including the `Qwen3_5MTP` support that makes
|
| 256 |
-
the MTP head usable at all.
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model:
|
| 3 |
+
- nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
|
| 4 |
+
library_name: transformers
|
| 5 |
+
pipeline_tag: image-text-to-text
|
| 6 |
+
tags:
|
| 7 |
+
- nvfp4
|
| 8 |
+
- modelopt
|
| 9 |
+
- vllm
|
| 10 |
+
- speculative-decoding
|
| 11 |
+
- mtp
|
| 12 |
+
- dgx-spark
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP
|
| 16 |
+
|
| 17 |
+
NVFP4 quantisation of
|
| 18 |
+
[nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451](https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451),
|
| 19 |
+
built to run on a 128 GB DGX Spark, **with the MTP speculative-decoding head
|
| 20 |
+
intact**.
|
| 21 |
+
|
| 22 |
+
This is the same quantisation as
|
| 23 |
+
[PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4)
|
| 24 |
+
plus one 105 MB tensor that the other repo is missing. **If you want speculative
|
| 25 |
+
decoding, use this one.** If you do not, either works.
|
| 26 |
+
|
| 27 |
+
| | this repo (`-MTP`) | the plain `-NVFP4` repo |
|
| 28 |
+
|---|---|---|
|
| 29 |
+
| tensors | **2399** | 2398 |
|
| 30 |
+
| `mtp.fc.weight` | **present** | absent |
|
| 31 |
+
| runs `--speculative-config method=mtp` | **yes** | loads, then drafts garbage |
|
| 32 |
+
| size | 20.6 GB | 20.5 GB |
|
| 33 |
+
|
| 34 |
+
---
|
| 35 |
+
|
| 36 |
+
## Read this first: the missing tensor problem
|
| 37 |
+
|
| 38 |
+
The upstream bf16 repo was **re-uploaded on 2026-07-21** to add
|
| 39 |
+
`mtp.fc.weight` - the fusion projection the MTP head needs to combine its two
|
| 40 |
+
inputs. **A checkpoint pulled before that date has `mtp.layers.0.*` but no
|
| 41 |
+
`mtp.fc`, and this is a silent failure.**
|
| 42 |
+
|
| 43 |
+
vLLM disables strict weight-initialisation checking for quantized configs. The
|
| 44 |
+
missing tensor is allocated, nothing is loaded into it, and it keeps whatever
|
| 45 |
+
was in that memory. The MTP head then drafts from a **random projection**: no
|
| 46 |
+
exception, no warning, nothing in the log. Every drafted token is rejected, you
|
| 47 |
+
pay the drafting cost for nothing, and decode gets *slower*. The only symptom is
|
| 48 |
+
a number you were probably not measuring.
|
| 49 |
+
|
| 50 |
+
Check your own copy:
|
| 51 |
+
|
| 52 |
+
```python
|
| 53 |
+
import json, urllib.request
|
| 54 |
+
idx = json.load(urllib.request.urlopen(
|
| 55 |
+
"https://huggingface.co/<repo>/resolve/main/model.safetensors.index.json"))
|
| 56 |
+
print("mtp.fc.weight:", "mtp.fc.weight" in idx["weight_map"])
|
| 57 |
+
```
|
| 58 |
+
|
| 59 |
+
`False` means MTP will not work for you, however healthy it looks.
|
| 60 |
+
|
| 61 |
+
## Serving it
|
| 62 |
+
|
| 63 |
+
Tested on vLLM `0.25.2.dev0` (arm64/CUDA 13 build), DGX Spark GB10, sm_121a.
|
| 64 |
+
|
| 65 |
+
**Plain, no speculative decoding:**
|
| 66 |
+
|
| 67 |
+
```bash
|
| 68 |
+
vllm serve PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP \
|
| 69 |
+
--kv-cache-dtype fp8 --attention-backend flashinfer \
|
| 70 |
+
--max-model-len 131072 --gpu-memory-utilization 0.72 --max-num-seqs 64 \
|
| 71 |
+
--reasoning-parser qwen3 --trust-remote-code \
|
| 72 |
+
--enable-auto-tool-choice --tool-call-parser qwen3_xml
|
| 73 |
+
```
|
| 74 |
+
|
| 75 |
+
**With MTP (recommended):** add
|
| 76 |
+
|
| 77 |
+
```bash
|
| 78 |
+
--speculative-config '{"method":"mtp","num_speculative_tokens":2}'
|
| 79 |
+
```
|
| 80 |
+
|
| 81 |
+
No separate draft model - vLLM rewrites this checkpoint's own config into a
|
| 82 |
+
`Qwen3_5MTP` draft and reads the head out of these weights. It is a reasoning
|
| 83 |
+
model, so pass `chat_template_kwargs: {"enable_thinking": false}` if you do not
|
| 84 |
+
want the budget spent inside `<think>`.
|
| 85 |
+
|
| 86 |
+
**With DFlash** (much faster on file-editing work, much slower on prose, and it
|
| 87 |
+
needs a patched draft because the published one does not load on vLLM at all):
|
| 88 |
+
see the recipe repo linked below.
|
| 89 |
+
|
| 90 |
+
---
|
| 91 |
+
|
| 92 |
+
## Is speculative decoding worth it? Measured, not guessed
|
| 93 |
+
|
| 94 |
+
540 measurements on this checkpoint: 3 configs x 3 workloads x 10 concurrency
|
| 95 |
+
levels x 6 shuffled rounds, all clean, standard deviation mostly under 1%.
|
| 96 |
+
|
| 97 |
+
### The plain-English version
|
| 98 |
+
|
| 99 |
+
An AI writes one word at a time, and each word means re-reading the whole model
|
| 100 |
+
from memory. That memory trip is the slow part. **Speculative decoding** adds a
|
| 101 |
+
small fast helper that guesses the next few words so the big model can check
|
| 102 |
+
them in one go. Right guesses are free. Wrong guesses are thrown away and cost
|
| 103 |
+
you time.
|
| 104 |
+
|
| 105 |
+
**So it is a bet on whether the next words are predictable.**
|
| 106 |
+
|
| 107 |
+
- **MTP** guesses **2 words** ahead and is usually right. Small bet, reliable.
|
| 108 |
+
- **DFlash** guesses **15 words** ahead at once. Astonishing when the text is
|
| 109 |
+
predictable; wasteful when it is not.
|
| 110 |
+
|
| 111 |
+
Text is predictable when the model is mostly **copying something you just gave
|
| 112 |
+
it** - "here is my file, rename this variable" - which is what coding assistants
|
| 113 |
+
do all day. It is unpredictable when the model is **writing something new**.
|
| 114 |
+
|
| 115 |
+
How often each helper guessed right, on this model:
|
| 116 |
+
|
| 117 |
+
| | writing prose | writing code | editing a file |
|
| 118 |
+
|---|---|---|---|
|
| 119 |
+
| **MTP** | 47% | 76% | **100%** |
|
| 120 |
+
| **DFlash** | **6%** | 19% | 83% |
|
| 121 |
+
|
| 122 |
+
**DFlash guessing right 6% of the time on prose is the whole reason it is not a
|
| 123 |
+
magic speed-up button.**
|
| 124 |
+
|
| 125 |
+

|
| 126 |
+
|
| 127 |
+
Rank the workloads by that hit rate and you have ranked them by outcome, for
|
| 128 |
+
both methods, at every concurrency level. It is not really about concurrency.
|
| 129 |
+
|
| 130 |
+
### MTP is a safe upgrade. Use it by default.
|
| 131 |
+
|
| 132 |
+
Aggregate tokens/sec vs no speculative decoding:
|
| 133 |
+
|
| 134 |
+
| workload | MTP vs no-spec |
|
| 135 |
+
|---|---|
|
| 136 |
+
| **code** | wins at **every** concurrency level tested, up to **+70%** |
|
| 137 |
+
| **edit** | wins at **every** level, up to **+104%** |
|
| 138 |
+
| prose | wins at 1-8, 24, 32 streams (up to +32%); loses at 12 (-11%), 16 (-10%), 64 (-7%) |
|
| 139 |
+
|
| 140 |
+
It costs about **10% of the KV cache pool** and nothing else. On code and editing
|
| 141 |
+
there is no concurrency at which it is the wrong choice. On free-form prose above
|
| 142 |
+
8 concurrent streams it is roughly a coin flip - that is where its acceptance is
|
| 143 |
+
weakest (47%).
|
| 144 |
+
|
| 145 |
+
**Writing code** - MTP (orange) leads base (blue) at every level:
|
| 146 |
+
|
| 147 |
+

|
| 148 |
+
|
| 149 |
+
**Editing files** - DFlash (green) is in another league, MTP still well clear of base:
|
| 150 |
+
|
| 151 |
+

|
| 152 |
+
|
| 153 |
+
**Free-form prose** - DFlash drops below *doing nothing* from 4 streams up:
|
| 154 |
+
|
| 155 |
+

|
| 156 |
+
|
| 157 |
+
### DFlash is a specialist tool with a sharp edge
|
| 158 |
+
|
| 159 |
+
Single stream on editing work it is **+716%** over no speculation and **+311%**
|
| 160 |
+
over MTP. That is real. But:
|
| 161 |
+
|
| 162 |
+
- On **prose** it is beaten by *doing nothing* from 4 concurrent streams up.
|
| 163 |
+
- It costs **40%** of the KV pool (second model + 15 draft slots per stream).
|
| 164 |
+
- It caps at about **22 concurrent streams** on this hardware and queues the
|
| 165 |
+
rest - median time to first token at 64 streams goes to **138 seconds**.
|
| 166 |
+
|
| 167 |
+
Ask for 64 streams and count how many actually run at once. Base and MTP track
|
| 168 |
+
the diagonal; DFlash flattens at 22, so its throughput lines beyond that point
|
| 169 |
+
are a capacity limit rather than a speed result:
|
| 170 |
+
|
| 171 |
+

|
| 172 |
+
|
| 173 |
+
**Rule of thumb: DFlash below its ~22-stream ceiling for edit-heavy agentic
|
| 174 |
+
coding; MTP for everything else and for crowds.**
|
| 175 |
+
|
| 176 |
+
Full charts, tables, the 540-cell dataset and the fix that makes DFlash load on
|
| 177 |
+
vLLM at all:
|
| 178 |
+
|
| 179 |
+
**[-> Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4_Dflash_DGX_recipe](https://github.com/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4_Dflash_DGX_recipe)**
|
| 180 |
+
|
| 181 |
+
### A warning about averages
|
| 182 |
+
|
| 183 |
+
Averaging prose/code/edit assumes editing is exactly a third of your traffic.
|
| 184 |
+
**72% of DFlash's average comes from the edit workload alone.** Drop editing and
|
| 185 |
+
average only prose and code, and MTP wins from 4 streams up. Measure your own
|
| 186 |
+
mix instead - acceptance rate tells you directly:
|
| 187 |
+
|
| 188 |
+
```bash
|
| 189 |
+
curl -s http://127.0.0.1:8000/metrics \
|
| 190 |
+
| grep -E 'spec_decode_num_(accepted|draft)_tokens_total'
|
| 191 |
+
```
|
| 192 |
+
|
| 193 |
+
Near 0.06 is prose-like (do not use DFlash); near 0.83 is edit-like (do).
|
| 194 |
+
|
| 195 |
+
---
|
| 196 |
+
|
| 197 |
+
## How this was made
|
| 198 |
+
|
| 199 |
+
Source: `nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451` - bf16, 1199
|
| 200 |
+
tensors, 55,457,998,304 bytes, a `nuslerp` mergekit merge of two Qwen3.6-27B
|
| 201 |
+
Architect-Polaris Fable variants. Uploaded with nightmedia's permission.
|
| 202 |
+
|
| 203 |
+
Quantised with **NVIDIA ModelOpt 0.45.0**, `examples/llm_ptq/hf_ptq.py`:
|
| 204 |
+
|
| 205 |
+
```
|
| 206 |
+
--qformat nvfp4 --calib_size 512 --dataset cnn_dailymail
|
| 207 |
+
--batch_size 0 --inference_tensor_parallel 1
|
| 208 |
+
```
|
| 209 |
+
|
| 210 |
+
ModelOpt auto-excludes `linear_attn`, the vision tower, `mtp.*`, embeddings and
|
| 211 |
+
`lm_head` from quantisation - see `quant_summary.txt` in this repo for the
|
| 212 |
+
per-layer record, and `hf_quant_config.json` for the exclusion patterns.
|
| 213 |
+
Output: 2398 tensors, 20,487,991,904 bytes.
|
| 214 |
+
|
| 215 |
+
`mtp.fc.weight` ([5120, 10240] BF16, 104,857,600 bytes) was then taken from the
|
| 216 |
+
corrected upstream revision and added as `model-mtp-fc.safetensors`, with the
|
| 217 |
+
index updated - **2399 tensors, 20,592,849,504 bytes**. vLLM's own source
|
| 218 |
+
expects this tensor to be unquantized in NVFP4 checkpoints
|
| 219 |
+
(`qwen3_5_mtp.py`: *"mtp.fc is stored as BF16 in NVFP4 checkpoints ... Force
|
| 220 |
+
unquantized"*), so a spliced BF16 tensor is exactly what a fresh PTQ run would
|
| 221 |
+
have emitted.
|
| 222 |
+
|
| 223 |
+
The three VL processor configs (`preprocessor_config.json`,
|
| 224 |
+
`processor_config.json`, `video_preprocessor_config.json`) are included -
|
| 225 |
+
ModelOpt does not emit them and vLLM hard-fails a VL model at init without them.
|
| 226 |
+
|
| 227 |
+
## Limitations
|
| 228 |
+
|
| 229 |
+
- **Measured on one machine, one model.** DGX Spark GB10, 121.7 GiB unified
|
| 230 |
+
memory, sm_121a, arm64, a community arm64/CUDA-13 vLLM build. No claim about
|
| 231 |
+
other hardware, multi-node, or other quantizations.
|
| 232 |
+
- **Acceptance rates are specific to this checkpoint.** The MTP head was
|
| 233 |
+
inherited through a four-generation mergekit merge and never retrained. A
|
| 234 |
+
different model will have different acceptance, and therefore different
|
| 235 |
+
crossover points.
|
| 236 |
+
- **No quality evaluation was run here.** NVFP4 is a lossy 4-bit quantisation.
|
| 237 |
+
These numbers are throughput only; nothing in this card says the output is as
|
| 238 |
+
good as the bf16 source.
|
| 239 |
+
- **Benchmarks ran with thinking disabled** and 512-token generations. Long
|
| 240 |
+
reasoning traces were not measured.
|
| 241 |
+
- `min_p` and `logit_bias` do not work under speculative decoding (vLLM warns at
|
| 242 |
+
startup).
|
| 243 |
+
|
| 244 |
+
## Credits
|
| 245 |
+
|
| 246 |
+
- **[nightmedia](https://huggingface.co/nightmedia)** - the bf16 merge this is
|
| 247 |
+
quantised from, and its Architect-Polaris merge stages. All the model work is
|
| 248 |
+
theirs; uploaded with their permission. That merge in turn credits
|
| 249 |
+
Qwen/Qwen3.6-27B and its own upstreams.
|
| 250 |
+
- **[z-lab](https://huggingface.co/z-lab)** - the DFlash draft used in the
|
| 251 |
+
comparison.
|
| 252 |
+
- **NVIDIA ModelOpt** - the PTQ pipeline.
|
| 253 |
+
- **vLLM** - the serving engine, including the `Qwen3_5MTP` support that makes
|
| 254 |
+
the MTP head usable at all.
|
|
|
|
|
|