PassingByPixels commited on
Commit
0429235
·
verified ·
1 Parent(s): 06d6d81

Name the checkpoint correctly: Fable-B-F451-NVFP4 (F451 is part of the source merge identity)

Browse files
Files changed (1) hide show
  1. README.md +254 -256
README.md CHANGED
@@ -1,256 +1,254 @@
1
- ---
2
- base_model:
3
- - nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
4
- library_name: transformers
5
- pipeline_tag: image-text-to-text
6
- tags:
7
- - nvfp4
8
- - modelopt
9
- - vllm
10
- - speculative-decoding
11
- - mtp
12
- - dgx-spark
13
- ---
14
-
15
- # Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP
16
-
17
- NVFP4 quantisation of
18
- [nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451](https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451),
19
- built to run on a 128 GB DGX Spark, All of the actual model work here belongs to nightmedia, who built the base model this is quantized from. That model is itself a merge sitting on top of a longer chain of work by other people in the community: **with the MTP speculative-decoding head
20
- intact**.
21
-
22
- This is the same quantisation as
23
- [PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4)
24
- plus one 105 MB tensor that the other repo is missing. **If you want speculative
25
- decoding, use this one.** If you do not, either works.
26
-
27
- | | this repo (`-MTP`) | the plain `-NVFP4` repo |
28
- |---|---|---|
29
- | tensors | **2399** | 2398 |
30
- | `mtp.fc.weight` | **present** | absent |
31
- | runs `--speculative-config method=mtp` | **yes** | loads, then drafts garbage |
32
- | size | 20.6 GB | 20.5 GB |
33
-
34
- DFLASH Recipe can be found here
35
- **[-> Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4_Dflash_DGX_recipe](https://github.com/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4_Dflash_DGX_recipe)**
36
- ---
37
-
38
- ## Read this first: the missing tensor problem
39
-
40
- The upstream bf16 repo was **re-uploaded on 2026-07-21** to add
41
- `mtp.fc.weight` - the fusion projection the MTP head needs to combine its two
42
- inputs. **A checkpoint pulled before that date has `mtp.layers.0.*` but no
43
- `mtp.fc`, and this is a silent failure.**
44
-
45
- vLLM disables strict weight-initialisation checking for quantized configs. The
46
- missing tensor is allocated, nothing is loaded into it, and it keeps whatever
47
- was in that memory. The MTP head then drafts from a **random projection**: no
48
- exception, no warning, nothing in the log. Every drafted token is rejected, you
49
- pay the drafting cost for nothing, and decode gets *slower*. The only symptom is
50
- a number you were probably not measuring.
51
-
52
- Check your own copy:
53
-
54
- ```python
55
- import json, urllib.request
56
- idx = json.load(urllib.request.urlopen(
57
- "https://huggingface.co/<repo>/resolve/main/model.safetensors.index.json"))
58
- print("mtp.fc.weight:", "mtp.fc.weight" in idx["weight_map"])
59
- ```
60
-
61
- `False` means MTP will not work for you, however healthy it looks.
62
-
63
- ## Serving it
64
-
65
- Tested on vLLM `0.25.2.dev0` (arm64/CUDA 13 build), DGX Spark GB10, sm_121a.
66
-
67
- **Plain, no speculative decoding:**
68
-
69
- ```bash
70
- vllm serve PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP \
71
- --kv-cache-dtype fp8 --attention-backend flashinfer \
72
- --max-model-len 131072 --gpu-memory-utilization 0.72 --max-num-seqs 64 \
73
- --reasoning-parser qwen3 --trust-remote-code \
74
- --enable-auto-tool-choice --tool-call-parser qwen3_xml
75
- ```
76
-
77
- **With MTP (recommended):** add
78
-
79
- ```bash
80
- --speculative-config '{"method":"mtp","num_speculative_tokens":2}'
81
- ```
82
-
83
- No separate draft model - vLLM rewrites this checkpoint's own config into a
84
- `Qwen3_5MTP` draft and reads the head out of these weights. It is a reasoning
85
- model, so pass `chat_template_kwargs: {"enable_thinking": false}` if you do not
86
- want the budget spent inside `<think>`.
87
-
88
- **With DFlash** (much faster on file-editing work, much slower on prose, and it
89
- needs a patched draft because the published one does not load on vLLM at all):
90
- see the recipe repo linked below.
91
-
92
- ---
93
-
94
- ## Is speculative decoding worth it? Measured, not guessed
95
-
96
- 540 measurements on this checkpoint: 3 configs x 3 workloads x 10 concurrency
97
- levels x 6 shuffled rounds, all clean, standard deviation mostly under 1%.
98
-
99
- ### The plain-English version
100
-
101
- An AI writes one word at a time, and each word means re-reading the whole model
102
- from memory. That memory trip is the slow part. **Speculative decoding** adds a
103
- small fast helper that guesses the next few words so the big model can check
104
- them in one go. Right guesses are free. Wrong guesses are thrown away and cost
105
- you time.
106
-
107
- **So it is a bet on whether the next words are predictable.**
108
-
109
- - **MTP** guesses **2 words** ahead and is usually right. Small bet, reliable.
110
- - **DFlash** guesses **15 words** ahead at once. Astonishing when the text is
111
- predictable; wasteful when it is not.
112
-
113
- Text is predictable when the model is mostly **copying something you just gave
114
- it** - "here is my file, rename this variable" - which is what coding assistants
115
- do all day. It is unpredictable when the model is **writing something new**.
116
-
117
- How often each helper guessed right, on this model:
118
-
119
- | | writing prose | writing code | editing a file |
120
- |---|---|---|---|
121
- | **MTP** | 47% | 76% | **100%** |
122
- | **DFlash** | **6%** | 19% | 83% |
123
-
124
- **DFlash guessing right 6% of the time on prose is the whole reason it is not a
125
- magic speed-up button.**
126
-
127
- ![draft acceptance rate by workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP/resolve/main/charts/c5.svg)
128
-
129
- Rank the workloads by that hit rate and you have ranked them by outcome, for
130
- both methods, at every concurrency level. It is not really about concurrency.
131
-
132
- ### MTP is a safe upgrade. Use it by default.
133
-
134
- Aggregate tokens/sec vs no speculative decoding:
135
-
136
- | workload | MTP vs no-spec |
137
- |---|---|
138
- | **code** | wins at **every** concurrency level tested, up to **+70%** |
139
- | **edit** | wins at **every** level, up to **+104%** |
140
- | prose | wins at 1-8, 24, 32 streams (up to +32%); loses at 12 (-11%), 16 (-10%), 64 (-7%) |
141
-
142
- It costs about **10% of the KV cache pool** and nothing else. On code and editing
143
- there is no concurrency at which it is the wrong choice. On free-form prose above
144
- 8 concurrent streams it is roughly a coin flip - that is where its acceptance is
145
- weakest (47%).
146
-
147
- **Writing code** - MTP (orange) leads base (blue) at every level:
148
-
149
- ![aggregate throughput, code workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP/resolve/main/charts/c1-code.svg)
150
-
151
- **Editing files** - DFlash (green) is in another league, MTP still well clear of base:
152
-
153
- ![aggregate throughput, edit workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP/resolve/main/charts/c1-edit.svg)
154
-
155
- **Free-form prose** - DFlash drops below *doing nothing* from 4 streams up:
156
-
157
- ![aggregate throughput, prose workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP/resolve/main/charts/c1-prose.svg)
158
-
159
- ### DFlash is a specialist tool with a sharp edge
160
-
161
- Single stream on editing work it is **+716%** over no speculation and **+311%**
162
- over MTP. That is real. But:
163
-
164
- - On **prose** it is beaten by *doing nothing* from 4 concurrent streams up.
165
- - It costs **40%** of the KV pool (second model + 15 draft slots per stream).
166
- - It caps at about **22 concurrent streams** on this hardware and queues the
167
- rest - median time to first token at 64 streams goes to **138 seconds**.
168
-
169
- Ask for 64 streams and count how many actually run at once. Base and MTP track
170
- the diagonal; DFlash flattens at 22, so its throughput lines beyond that point
171
- are a capacity limit rather than a speed result:
172
-
173
- ![concurrency achieved vs requested](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4-MTP/resolve/main/charts/c2c.svg)
174
-
175
- **Rule of thumb: DFlash below its ~22-stream ceiling for edit-heavy agentic
176
- coding; MTP for everything else and for crowds.**
177
-
178
- Full charts, tables, the 540-cell dataset and the fix that makes DFlash load on
179
- vLLM at all:
180
-
181
- **[-> Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4_Dflash_DGX_recipe](https://github.com/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-NVFP4_Dflash_DGX_recipe)**
182
-
183
- ### A warning about averages
184
-
185
- Averaging prose/code/edit assumes editing is exactly a third of your traffic.
186
- **72% of DFlash's average comes from the edit workload alone.** Drop editing and
187
- average only prose and code, and MTP wins from 4 streams up. Measure your own
188
- mix instead - acceptance rate tells you directly:
189
-
190
- ```bash
191
- curl -s http://127.0.0.1:8000/metrics \
192
- | grep -E 'spec_decode_num_(accepted|draft)_tokens_total'
193
- ```
194
-
195
- Near 0.06 is prose-like (do not use DFlash); near 0.83 is edit-like (do).
196
-
197
- ---
198
-
199
- ## How this was made
200
-
201
- Source: `nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451` - bf16, 1199
202
- tensors, 55,457,998,304 bytes, a `nuslerp` mergekit merge of two Qwen3.6-27B
203
- Architect-Polaris Fable variants. Uploaded with nightmedia's permission.
204
-
205
- Quantised with **NVIDIA ModelOpt 0.45.0**, `examples/llm_ptq/hf_ptq.py`:
206
-
207
- ```
208
- --qformat nvfp4 --calib_size 512 --dataset cnn_dailymail
209
- --batch_size 0 --inference_tensor_parallel 1
210
- ```
211
-
212
- ModelOpt auto-excludes `linear_attn`, the vision tower, `mtp.*`, embeddings and
213
- `lm_head` from quantisation - see `quant_summary.txt` in this repo for the
214
- per-layer record, and `hf_quant_config.json` for the exclusion patterns.
215
- Output: 2398 tensors, 20,487,991,904 bytes.
216
-
217
- `mtp.fc.weight` ([5120, 10240] BF16, 104,857,600 bytes) was then taken from the
218
- corrected upstream revision and added as `model-mtp-fc.safetensors`, with the
219
- index updated - **2399 tensors, 20,592,849,504 bytes**. vLLM's own source
220
- expects this tensor to be unquantized in NVFP4 checkpoints
221
- (`qwen3_5_mtp.py`: *"mtp.fc is stored as BF16 in NVFP4 checkpoints ... Force
222
- unquantized"*), so a spliced BF16 tensor is exactly what a fresh PTQ run would
223
- have emitted.
224
-
225
- The three VL processor configs (`preprocessor_config.json`,
226
- `processor_config.json`, `video_preprocessor_config.json`) are included -
227
- ModelOpt does not emit them and vLLM hard-fails a VL model at init without them.
228
-
229
- ## Limitations
230
-
231
- - **Measured on one machine, one model.** DGX Spark GB10, 121.7 GiB unified
232
- memory, sm_121a, arm64, a community arm64/CUDA-13 vLLM build. No claim about
233
- other hardware, multi-node, or other quantizations.
234
- - **Acceptance rates are specific to this checkpoint.** The MTP head was
235
- inherited through a four-generation mergekit merge and never retrained. A
236
- different model will have different acceptance, and therefore different
237
- crossover points.
238
- - **No quality evaluation was run here.** NVFP4 is a lossy 4-bit quantisation.
239
- These numbers are throughput only; nothing in this card says the output is as
240
- good as the bf16 source.
241
- - **Benchmarks ran with thinking disabled** and 512-token generations. Long
242
- reasoning traces were not measured.
243
- - `min_p` and `logit_bias` do not work under speculative decoding (vLLM warns at
244
- startup).
245
-
246
- ## Credits
247
-
248
- - **[nightmedia](https://huggingface.co/nightmedia)** - the bf16 merge this is
249
- quantised from, and its Architect-Polaris merge stages. All the model work is
250
- theirs; uploaded with their permission. That merge in turn credits
251
- Qwen/Qwen3.6-27B and its own upstreams.
252
- - **[z-lab](https://huggingface.co/z-lab)** - the DFlash draft used in the
253
- comparison.
254
- - **NVIDIA ModelOpt** - the PTQ pipeline.
255
- - **vLLM** - the serving engine, including the `Qwen3_5MTP` support that makes
256
- the MTP head usable at all.
 
1
+ ---
2
+ base_model:
3
+ - nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451
4
+ library_name: transformers
5
+ pipeline_tag: image-text-to-text
6
+ tags:
7
+ - nvfp4
8
+ - modelopt
9
+ - vllm
10
+ - speculative-decoding
11
+ - mtp
12
+ - dgx-spark
13
+ ---
14
+
15
+ # Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP
16
+
17
+ NVFP4 quantisation of
18
+ [nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451](https://huggingface.co/nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451),
19
+ built to run on a 128 GB DGX Spark, **with the MTP speculative-decoding head
20
+ intact**.
21
+
22
+ This is the same quantisation as
23
+ [PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4)
24
+ plus one 105 MB tensor that the other repo is missing. **If you want speculative
25
+ decoding, use this one.** If you do not, either works.
26
+
27
+ | | this repo (`-MTP`) | the plain `-NVFP4` repo |
28
+ |---|---|---|
29
+ | tensors | **2399** | 2398 |
30
+ | `mtp.fc.weight` | **present** | absent |
31
+ | runs `--speculative-config method=mtp` | **yes** | loads, then drafts garbage |
32
+ | size | 20.6 GB | 20.5 GB |
33
+
34
+ ---
35
+
36
+ ## Read this first: the missing tensor problem
37
+
38
+ The upstream bf16 repo was **re-uploaded on 2026-07-21** to add
39
+ `mtp.fc.weight` - the fusion projection the MTP head needs to combine its two
40
+ inputs. **A checkpoint pulled before that date has `mtp.layers.0.*` but no
41
+ `mtp.fc`, and this is a silent failure.**
42
+
43
+ vLLM disables strict weight-initialisation checking for quantized configs. The
44
+ missing tensor is allocated, nothing is loaded into it, and it keeps whatever
45
+ was in that memory. The MTP head then drafts from a **random projection**: no
46
+ exception, no warning, nothing in the log. Every drafted token is rejected, you
47
+ pay the drafting cost for nothing, and decode gets *slower*. The only symptom is
48
+ a number you were probably not measuring.
49
+
50
+ Check your own copy:
51
+
52
+ ```python
53
+ import json, urllib.request
54
+ idx = json.load(urllib.request.urlopen(
55
+ "https://huggingface.co/<repo>/resolve/main/model.safetensors.index.json"))
56
+ print("mtp.fc.weight:", "mtp.fc.weight" in idx["weight_map"])
57
+ ```
58
+
59
+ `False` means MTP will not work for you, however healthy it looks.
60
+
61
+ ## Serving it
62
+
63
+ Tested on vLLM `0.25.2.dev0` (arm64/CUDA 13 build), DGX Spark GB10, sm_121a.
64
+
65
+ **Plain, no speculative decoding:**
66
+
67
+ ```bash
68
+ vllm serve PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP \
69
+ --kv-cache-dtype fp8 --attention-backend flashinfer \
70
+ --max-model-len 131072 --gpu-memory-utilization 0.72 --max-num-seqs 64 \
71
+ --reasoning-parser qwen3 --trust-remote-code \
72
+ --enable-auto-tool-choice --tool-call-parser qwen3_xml
73
+ ```
74
+
75
+ **With MTP (recommended):** add
76
+
77
+ ```bash
78
+ --speculative-config '{"method":"mtp","num_speculative_tokens":2}'
79
+ ```
80
+
81
+ No separate draft model - vLLM rewrites this checkpoint's own config into a
82
+ `Qwen3_5MTP` draft and reads the head out of these weights. It is a reasoning
83
+ model, so pass `chat_template_kwargs: {"enable_thinking": false}` if you do not
84
+ want the budget spent inside `<think>`.
85
+
86
+ **With DFlash** (much faster on file-editing work, much slower on prose, and it
87
+ needs a patched draft because the published one does not load on vLLM at all):
88
+ see the recipe repo linked below.
89
+
90
+ ---
91
+
92
+ ## Is speculative decoding worth it? Measured, not guessed
93
+
94
+ 540 measurements on this checkpoint: 3 configs x 3 workloads x 10 concurrency
95
+ levels x 6 shuffled rounds, all clean, standard deviation mostly under 1%.
96
+
97
+ ### The plain-English version
98
+
99
+ An AI writes one word at a time, and each word means re-reading the whole model
100
+ from memory. That memory trip is the slow part. **Speculative decoding** adds a
101
+ small fast helper that guesses the next few words so the big model can check
102
+ them in one go. Right guesses are free. Wrong guesses are thrown away and cost
103
+ you time.
104
+
105
+ **So it is a bet on whether the next words are predictable.**
106
+
107
+ - **MTP** guesses **2 words** ahead and is usually right. Small bet, reliable.
108
+ - **DFlash** guesses **15 words** ahead at once. Astonishing when the text is
109
+ predictable; wasteful when it is not.
110
+
111
+ Text is predictable when the model is mostly **copying something you just gave
112
+ it** - "here is my file, rename this variable" - which is what coding assistants
113
+ do all day. It is unpredictable when the model is **writing something new**.
114
+
115
+ How often each helper guessed right, on this model:
116
+
117
+ | | writing prose | writing code | editing a file |
118
+ |---|---|---|---|
119
+ | **MTP** | 47% | 76% | **100%** |
120
+ | **DFlash** | **6%** | 19% | 83% |
121
+
122
+ **DFlash guessing right 6% of the time on prose is the whole reason it is not a
123
+ magic speed-up button.**
124
+
125
+ ![draft acceptance rate by workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP/resolve/main/charts/c5.svg)
126
+
127
+ Rank the workloads by that hit rate and you have ranked them by outcome, for
128
+ both methods, at every concurrency level. It is not really about concurrency.
129
+
130
+ ### MTP is a safe upgrade. Use it by default.
131
+
132
+ Aggregate tokens/sec vs no speculative decoding:
133
+
134
+ | workload | MTP vs no-spec |
135
+ |---|---|
136
+ | **code** | wins at **every** concurrency level tested, up to **+70%** |
137
+ | **edit** | wins at **every** level, up to **+104%** |
138
+ | prose | wins at 1-8, 24, 32 streams (up to +32%); loses at 12 (-11%), 16 (-10%), 64 (-7%) |
139
+
140
+ It costs about **10% of the KV cache pool** and nothing else. On code and editing
141
+ there is no concurrency at which it is the wrong choice. On free-form prose above
142
+ 8 concurrent streams it is roughly a coin flip - that is where its acceptance is
143
+ weakest (47%).
144
+
145
+ **Writing code** - MTP (orange) leads base (blue) at every level:
146
+
147
+ ![aggregate throughput, code workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP/resolve/main/charts/c1-code.svg)
148
+
149
+ **Editing files** - DFlash (green) is in another league, MTP still well clear of base:
150
+
151
+ ![aggregate throughput, edit workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP/resolve/main/charts/c1-edit.svg)
152
+
153
+ **Free-form prose** - DFlash drops below *doing nothing* from 4 streams up:
154
+
155
+ ![aggregate throughput, prose workload](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP/resolve/main/charts/c1-prose.svg)
156
+
157
+ ### DFlash is a specialist tool with a sharp edge
158
+
159
+ Single stream on editing work it is **+716%** over no speculation and **+311%**
160
+ over MTP. That is real. But:
161
+
162
+ - On **prose** it is beaten by *doing nothing* from 4 concurrent streams up.
163
+ - It costs **40%** of the KV pool (second model + 15 draft slots per stream).
164
+ - It caps at about **22 concurrent streams** on this hardware and queues the
165
+ rest - median time to first token at 64 streams goes to **138 seconds**.
166
+
167
+ Ask for 64 streams and count how many actually run at once. Base and MTP track
168
+ the diagonal; DFlash flattens at 22, so its throughput lines beyond that point
169
+ are a capacity limit rather than a speed result:
170
+
171
+ ![concurrency achieved vs requested](https://huggingface.co/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4-MTP/resolve/main/charts/c2c.svg)
172
+
173
+ **Rule of thumb: DFlash below its ~22-stream ceiling for edit-heavy agentic
174
+ coding; MTP for everything else and for crowds.**
175
+
176
+ Full charts, tables, the 540-cell dataset and the fix that makes DFlash load on
177
+ vLLM at all:
178
+
179
+ **[-> Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4_Dflash_DGX_recipe](https://github.com/PassingByPixels/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451-NVFP4_Dflash_DGX_recipe)**
180
+
181
+ ### A warning about averages
182
+
183
+ Averaging prose/code/edit assumes editing is exactly a third of your traffic.
184
+ **72% of DFlash's average comes from the edit workload alone.** Drop editing and
185
+ average only prose and code, and MTP wins from 4 streams up. Measure your own
186
+ mix instead - acceptance rate tells you directly:
187
+
188
+ ```bash
189
+ curl -s http://127.0.0.1:8000/metrics \
190
+ | grep -E 'spec_decode_num_(accepted|draft)_tokens_total'
191
+ ```
192
+
193
+ Near 0.06 is prose-like (do not use DFlash); near 0.83 is edit-like (do).
194
+
195
+ ---
196
+
197
+ ## How this was made
198
+
199
+ Source: `nightmedia/Qwen3.6-27B-Architect-Polaris2-Fable-B-F451` - bf16, 1199
200
+ tensors, 55,457,998,304 bytes, a `nuslerp` mergekit merge of two Qwen3.6-27B
201
+ Architect-Polaris Fable variants. Uploaded with nightmedia's permission.
202
+
203
+ Quantised with **NVIDIA ModelOpt 0.45.0**, `examples/llm_ptq/hf_ptq.py`:
204
+
205
+ ```
206
+ --qformat nvfp4 --calib_size 512 --dataset cnn_dailymail
207
+ --batch_size 0 --inference_tensor_parallel 1
208
+ ```
209
+
210
+ ModelOpt auto-excludes `linear_attn`, the vision tower, `mtp.*`, embeddings and
211
+ `lm_head` from quantisation - see `quant_summary.txt` in this repo for the
212
+ per-layer record, and `hf_quant_config.json` for the exclusion patterns.
213
+ Output: 2398 tensors, 20,487,991,904 bytes.
214
+
215
+ `mtp.fc.weight` ([5120, 10240] BF16, 104,857,600 bytes) was then taken from the
216
+ corrected upstream revision and added as `model-mtp-fc.safetensors`, with the
217
+ index updated - **2399 tensors, 20,592,849,504 bytes**. vLLM's own source
218
+ expects this tensor to be unquantized in NVFP4 checkpoints
219
+ (`qwen3_5_mtp.py`: *"mtp.fc is stored as BF16 in NVFP4 checkpoints ... Force
220
+ unquantized"*), so a spliced BF16 tensor is exactly what a fresh PTQ run would
221
+ have emitted.
222
+
223
+ The three VL processor configs (`preprocessor_config.json`,
224
+ `processor_config.json`, `video_preprocessor_config.json`) are included -
225
+ ModelOpt does not emit them and vLLM hard-fails a VL model at init without them.
226
+
227
+ ## Limitations
228
+
229
+ - **Measured on one machine, one model.** DGX Spark GB10, 121.7 GiB unified
230
+ memory, sm_121a, arm64, a community arm64/CUDA-13 vLLM build. No claim about
231
+ other hardware, multi-node, or other quantizations.
232
+ - **Acceptance rates are specific to this checkpoint.** The MTP head was
233
+ inherited through a four-generation mergekit merge and never retrained. A
234
+ different model will have different acceptance, and therefore different
235
+ crossover points.
236
+ - **No quality evaluation was run here.** NVFP4 is a lossy 4-bit quantisation.
237
+ These numbers are throughput only; nothing in this card says the output is as
238
+ good as the bf16 source.
239
+ - **Benchmarks ran with thinking disabled** and 512-token generations. Long
240
+ reasoning traces were not measured.
241
+ - `min_p` and `logit_bias` do not work under speculative decoding (vLLM warns at
242
+ startup).
243
+
244
+ ## Credits
245
+
246
+ - **[nightmedia](https://huggingface.co/nightmedia)** - the bf16 merge this is
247
+ quantised from, and its Architect-Polaris merge stages. All the model work is
248
+ theirs; uploaded with their permission. That merge in turn credits
249
+ Qwen/Qwen3.6-27B and its own upstreams.
250
+ - **[z-lab](https://huggingface.co/z-lab)** - the DFlash draft used in the
251
+ comparison.
252
+ - **NVIDIA ModelOpt** - the PTQ pipeline.
253
+ - **vLLM** - the serving engine, including the `Qwen3_5MTP` support that makes
254
+ the MTP head usable at all.