File size: 11,984 Bytes
5b96a97
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f5e378e
 
 
 
5b96a97
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40ecee5
 
 
 
 
 
 
 
 
5b96a97
 
 
 
 
f5e378e
 
 
 
 
 
 
 
 
5b96a97
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f5e378e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5b96a97
f5e378e
5b96a97
f5e378e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5b96a97
 
 
 
 
 
 
 
 
8dc02b8
5b96a97
 
 
 
8dc02b8
facf024
40ecee5
5b96a97
 
 
 
 
 
 
8dc02b8
 
 
 
 
 
40ecee5
facf024
 
 
 
40ecee5
 
facf024
 
5b96a97
8dc02b8
 
 
 
 
 
 
5b96a97
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
  - mxfp4
  - quark
  - amd
  - rocm
  - rdna4
  - gfx1201
  - vllm
  - quantized
---

# Qwen3.8-27B β€” MXFP4 (AMD Quark) for RDNA4

MXFP4 weight quantisation of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B),
built with [AMD Quark](https://quark.docs.amd.com) 0.12.post1 and targeted at **RDNA4**
(gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) β€” GPUs that sit outside the official ROCm
vLLM target list.

**What this buys you on 2Γ—32 GB RDNA4:** the full **262,144-token** context window at
roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get
long β€” see Throughput. Quality is at or above the bf16 reference on every cell measured so far
except one, which is stated below rather than omitted.

## What is and is not quantised

Only **MLP and MoE-expert projections** go to 4-bit. Attention (q/k/v/o and its norms), every
norm, embeddings, `lm_head`, routers/gates and the **entire vision path** stay bf16.

| | count |
|---|---|
| `mlp.{gate,up,down}_proj` | 192 (64 layers Γ— 3) |
| `linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj}` | 240 (48 layers Γ— 5) |
| **total quantised modules** | **432** |
| attention / norms / embeddings / `lm_head` / vision | **0** β€” verified, none |

Verified by tensor inspection: a module counts as quantised only if it carries a real artifact
(`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`).

Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding
attention costs little size and keeps those layers on the fast bf16 path.

For structural comparison, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the decoder's attention as
well β€” 496 quantised modules against 432 here, the difference being exactly the 16 full-attention
layers' q/k/v/o β€” and is **AWQ-calibrated** (`algo_config.name = awq`) where this build is data-free
RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that
build, which is not enough to publish a quality comparison from: strict-match moves by about
Β±0.06 across seeds on this hardware, which is wider than any gap it showed.

- Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized`
- Size: **22.3 GB** across 18 shards (bf16 source β‰ˆ 54 GB)
- Quark `exclude` list: 231 entries

> **The config declares W4A4, not weight-only.** Quark's `mxfp4` scheme enables dynamic fp4
> *activation* quantization by default, so `global_quant_config.input_tensors` reads
> `{dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}`. On the RDNA4 port that
> declaration is **not honoured** β€” the weight-only kernel ignores activation quant, and the
> FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on
> a runtime that *does* honour it, you will get a different numerical path than the one measured
> here. The difference from whole-decoder AMD-style builds is **coverage** (432 quantized modules
> vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width.

## Serving

Speed and window claims here need the RDNA4 port, which has the MXFP4Γ—e4m3 FP8-WMMA kernel:

```bash
docker run --rm -it --device /dev/kfd --device /dev/dri \
  -v /path/to/weights:/model:ro -p 8011:8011 \
  -e VLLM_RDNA_MXFP4_FP8=1 \
  capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \
  serve /model --served-model-name qwen --port 8011 \
  --tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code
```

Source: [`Capicua25x/vllm-rocm-rdna4`](https://github.com/Capicua25x/vllm-rocm-rdna4), branch
`rdna4-port-0.26.1`. `VLLM_RDNA_MXFP4_FP8=0` falls back to the weight-only bf16-unpack kernel.

> **On stock vLLM these weights load and generate correctly, but slower.** Without the
> FP8-WMMA kernel you get the weight-only dequant path β€” roughly 51 tok/s single-stream instead
> of 61 on this hardware β€” and on 32 GB cards you will not reach the 262k window. If you are
> benchmarking this against another quant, check which kernel you are actually on first.

Sampling follows the base model card: thinking `temp 1.0, top_p 0.95, top_k 20, min_p 0`;
non-thinking `temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5`.

## Throughput

Measured on **this exact artifact**, 2026-08-17, on 2Γ— Radeon AI PRO R9700 (TP2, gfx1201) with the
rc6 FP8-WMMA kernel (`VLLM_RDNA_MXFP4_FP8=1`) and native MTP-3 speculative decoding.
`max_tokens: 256`, thinking **on** β€” the shape most deployments actually run.

Compared against stock [`Qwen/Qwen3.8-27B-FP8`](https://huggingface.co/Qwen/Qwen3.8-27B-FP8)
**at matched capacity**: both configurations hold a 262,144-token window on the same two cards,
so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point β€” 131k
window β€” and was only partially swept; it is not compared here.)

**Short prompt (~30 tokens)** β€” per-user tok/s / aggregate tok/s:

| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | 47.7 / 48 | 48.1 / 48 |
| 8 | 28.6 / 213 | **37.1 / 278** |
| 16 | **26.5 / 384** | 21.5 / 322 |
| 32 | 18.2 / **539** | **21.2** / 435 |

**6k-token prompt** β€” closer to a real application's context:

| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | **46.3 / 46** | 36.0 / 36 |
| 8 | **26.2 / 199** | 10.1 / 79 |
| 16 | **16.9 / 260** | 5.5 / 87 |

**The 6k table is the one that matters.** At short prompts the two are close, and FP8 is ahead at
8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply β€” 79 tok/s
aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 β€” while this build
holds its single-stream rate almost unchanged (47.7 β†’ 46.3). If you are serving anything with a
system prompt, retrieved context or conversation history, that is the regime you will be in.

Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross
it by 32 concurrent.

**Thinking-off is not yet measured on these weights.** Figures published elsewhere for the rc6
kernel (61 tok/s single-stream, 649 aggregate) were measured on an *earlier* MXFP4 build of this
model, before this Quark build existed β€” they do not describe this artifact and are omitted
rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be
added here as they are run.

## Quality β€” measured, as of 2026-08-17

Same harness, same seed (1234), same on-spec sampling across all four columns. **bf16 ref** is
the unquantised model on a hosted endpoint; the two FP8 columns are stock `Qwen/Qwen3.8-27B-FP8`
on this same box, differing only in KV cache dtype.

| benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** |
|---|---|---|---|---|---|
| GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.94–0.96 / 0.92–0.94** ᢜ |
| GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** |
| IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** |
| GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** |
| AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
| AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᡃ | 0.800 | **0.780** |
| τ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** |
| τ²-bench airline (Pass^1) | 50 | 0.760 | β€” | β€” | **0.840** |
| HLE | 120 | 0.3083 | β€” | β€” | *running* |
| SWE-bench Verified | 100 | β€” | β€” | β€” | *pending* |
| Terminal-Bench Hard | 44 | β€” | β€” | β€” | *pending* |

ᡃ Scored on the 90 items it served; 10 were refused because the prompt exceeded that
configuration's 131k window. Blended over the full 100 it reads 0.720.

ᢜ **Two runs of this build exist at the same seed and identical settings** β€” 0.94/0.92 and
0.96/0.94 β€” so the honest figure is a range, not a point. The other three columns are single
runs, which is worth knowing before reading small deltas here as real: on this cell one run's
difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items
depending on which run you take.

**τ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against
the bf16 reference's 0.939 β€” eight simulations β€” and sits four behind the FP8 + bf16 KV arm and
three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items
*ahead*. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses.
Retail is still running and will add a third point.

On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling
(`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than
harness noise β€” but it is one domain and a single-digit item count, not a blanket weakness.

Everything else is at or above bf16: GSM8K strict-match **+5 to +6 items** (see note ᢜ β€” two
runs exist), GPQA **+8 items**, and long-context retrieval **identical** to bf16 at ~107k-token
prompts.

All AA-LCR figures are the runner's own judging pass, taken from each arm's `score.json`. A
second judging pass over the same generations moves scores by roughly one item in either
direction; mixing passes between arms would manufacture differences that are not there.

Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and
will be revised as they land; the commit history is the record of what was known when.

## Reproducing the quantisation

Data-free, CPU-only, file-to-file β€” no calibration set, no GPU, ~3 minutes for this model.

```python
from quark.torch.export.api import direct_quantize_checkpoint

EXCLUDE = [
    "lm_head", "*embed_tokens*",
    "*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj",
    "*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*",
    "*.linear_attn.conv1d", "*.linear_attn.norm",
    "*.mlp.gate", "*.mlp.shared_expert_gate",
    "mtp*", "*visual*", "*vision*",
]
```

Two things that are easy to get wrong:

- **`*.mlp.gate` and `*.mlp.gate_proj` are different modules.** The first is the MoE router and
  must stay bf16; the second is the SwiGLU gate projection and *should* be 4-bit. A glob that
  catches both silently quantises the router.
- **When verifying, key on real artifacts**, not on a `_scale` suffix. Several bf16 checkpoints
  in this family ship tensors like `vision_tower.std_scale` or per-layer `layer_scalar` in the
  *original* weights, and a naive check reports leaks on a perfectly correct build.

Check both directions β€” leakage (something quantised that should not be) *and* over-exclusion
(projections that were meant to be 4-bit but stayed bf16) β€” and make a mismatch raise.

Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid
sliding-attention model) are in the port repo; none of the exclude lists transfer between
families.

## Licence and attribution

Apache-2.0, inherited from [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B). The
`LICENSE` file here is byte-identical to upstream's.

**Modification made:** weights of the MLP and linear-attention projections converted from bf16 to
MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to
architecture, tokenizer or chat template. All other tensors are the upstream values.

The gfx1201 enablement this port descends from was first done by **Rob Smith (`tcclaviger`)** on
the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his
`_matmul_fp8_ogs`. See the `NOTICE` in the port repo for the full lineage.