Capicua25x's picture
report GSM8K as a range: two runs exist at the same seed and the card published the higher one; align AA-LCR arm B to its score.json (0.800, not the rejudge pass) so every arm uses the same judging pass
8dc02b8 verified
|
Raw History Blame
12 kB
---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- mxfp4
- quark
- amd
- rocm
- rdna4
- gfx1201
- vllm
- quantized
---
# Qwen3.8-27B β€” MXFP4 (AMD Quark) for RDNA4
MXFP4 weight quantisation of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B),
built with [AMD Quark](https://quark.docs.amd.com) 0.12.post1 and targeted at **RDNA4**
(gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) β€” GPUs that sit outside the official ROCm
vLLM target list.
**What this buys you on 2Γ—32 GB RDNA4:** the full **262,144-token** context window at
roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get
long β€” see Throughput. Quality is at or above the bf16 reference on every cell measured so far
except one, which is stated below rather than omitted.
## What is and is not quantised
Only **MLP and MoE-expert projections** go to 4-bit. Attention (q/k/v/o and its norms), every
norm, embeddings, `lm_head`, routers/gates and the **entire vision path** stay bf16.
| | count |
|---|---|
| `mlp.{gate,up,down}_proj` | 192 (64 layers Γ— 3) |
| `linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj}` | 240 (48 layers Γ— 5) |
| **total quantised modules** | **432** |
| attention / norms / embeddings / `lm_head` / vision | **0** β€” verified, none |
Verified by tensor inspection: a module counts as quantised only if it carries a real artifact
(`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`).
Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding
attention costs little size and keeps those layers on the fast bf16 path.
For structural comparison, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the decoder's attention as
well β€” 496 quantised modules against 432 here, the difference being exactly the 16 full-attention
layers' q/k/v/o β€” and is **AWQ-calibrated** (`algo_config.name = awq`) where this build is data-free
RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that
build, which is not enough to publish a quality comparison from: strict-match moves by about
Β±0.06 across seeds on this hardware, which is wider than any gap it showed.
- Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized`
- Size: **22.3 GB** across 18 shards (bf16 source β‰ˆ 54 GB)
- Quark `exclude` list: 231 entries
> **The config declares W4A4, not weight-only.** Quark's `mxfp4` scheme enables dynamic fp4
> *activation* quantization by default, so `global_quant_config.input_tensors` reads
> `{dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}`. On the RDNA4 port that
> declaration is **not honoured** β€” the weight-only kernel ignores activation quant, and the
> FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on
> a runtime that *does* honour it, you will get a different numerical path than the one measured
> here. The difference from whole-decoder AMD-style builds is **coverage** (432 quantized modules
> vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width.
## Serving
Speed and window claims here need the RDNA4 port, which has the MXFP4Γ—e4m3 FP8-WMMA kernel:
```bash
docker run --rm -it --device /dev/kfd --device /dev/dri \
-v /path/to/weights:/model:ro -p 8011:8011 \
-e VLLM_RDNA_MXFP4_FP8=1 \
capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \
serve /model --served-model-name qwen --port 8011 \
--tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code
```
Source: [`Capicua25x/vllm-rocm-rdna4`](https://github.com/Capicua25x/vllm-rocm-rdna4), branch
`rdna4-port-0.26.1`. `VLLM_RDNA_MXFP4_FP8=0` falls back to the weight-only bf16-unpack kernel.
> **On stock vLLM these weights load and generate correctly, but slower.** Without the
> FP8-WMMA kernel you get the weight-only dequant path β€” roughly 51 tok/s single-stream instead
> of 61 on this hardware β€” and on 32 GB cards you will not reach the 262k window. If you are
> benchmarking this against another quant, check which kernel you are actually on first.
Sampling follows the base model card: thinking `temp 1.0, top_p 0.95, top_k 20, min_p 0`;
non-thinking `temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5`.
## Throughput
Measured on **this exact artifact**, 2026-08-17, on 2Γ— Radeon AI PRO R9700 (TP2, gfx1201) with the
rc6 FP8-WMMA kernel (`VLLM_RDNA_MXFP4_FP8=1`) and native MTP-3 speculative decoding.
`max_tokens: 256`, thinking **on** β€” the shape most deployments actually run.
Compared against stock [`Qwen/Qwen3.8-27B-FP8`](https://huggingface.co/Qwen/Qwen3.8-27B-FP8)
**at matched capacity**: both configurations hold a 262,144-token window on the same two cards,
so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point β€” 131k
window β€” and was only partially swept; it is not compared here.)
**Short prompt (~30 tokens)** β€” per-user tok/s / aggregate tok/s:
| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | 47.7 / 48 | 48.1 / 48 |
| 8 | 28.6 / 213 | **37.1 / 278** |
| 16 | **26.5 / 384** | 21.5 / 322 |
| 32 | 18.2 / **539** | **21.2** / 435 |
**6k-token prompt** β€” closer to a real application's context:
| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | **46.3 / 46** | 36.0 / 36 |
| 8 | **26.2 / 199** | 10.1 / 79 |
| 16 | **16.9 / 260** | 5.5 / 87 |
**The 6k table is the one that matters.** At short prompts the two are close, and FP8 is ahead at
8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply β€” 79 tok/s
aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 β€” while this build
holds its single-stream rate almost unchanged (47.7 β†’ 46.3). If you are serving anything with a
system prompt, retrieved context or conversation history, that is the regime you will be in.
Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross
it by 32 concurrent.
**Thinking-off is not yet measured on these weights.** Figures published elsewhere for the rc6
kernel (61 tok/s single-stream, 649 aggregate) were measured on an *earlier* MXFP4 build of this
model, before this Quark build existed β€” they do not describe this artifact and are omitted
rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be
added here as they are run.
## Quality β€” measured, as of 2026-08-17
Same harness, same seed (1234), same on-spec sampling across all four columns. **bf16 ref** is
the unquantised model on a hosted endpoint; the two FP8 columns are stock `Qwen/Qwen3.8-27B-FP8`
on this same box, differing only in KV cache dtype.
| benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** |
|---|---|---|---|---|---|
| GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.94–0.96 / 0.92–0.94** ᢜ |
| GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** |
| IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** |
| GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** |
| AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
| AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᡃ | 0.800 | **0.780** |
| τ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** |
| τ²-bench airline (Pass^1) | 50 | 0.760 | β€” | β€” | **0.840** |
| HLE | 120 | 0.3083 | β€” | β€” | *running* |
| SWE-bench Verified | 100 | β€” | β€” | β€” | *pending* |
| Terminal-Bench Hard | 44 | β€” | β€” | β€” | *pending* |
ᡃ Scored on the 90 items it served; 10 were refused because the prompt exceeded that
configuration's 131k window. Blended over the full 100 it reads 0.720.
ᢜ **Two runs of this build exist at the same seed and identical settings** β€” 0.94/0.92 and
0.96/0.94 β€” so the honest figure is a range, not a point. The other three columns are single
runs, which is worth knowing before reading small deltas here as real: on this cell one run's
difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items
depending on which run you take.
**τ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against
the bf16 reference's 0.939 β€” eight simulations β€” and sits four behind the FP8 + bf16 KV arm and
three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items
*ahead*. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses.
Retail is still running and will add a third point.
On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling
(`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than
harness noise β€” but it is one domain and a single-digit item count, not a blanket weakness.
Everything else is at or above bf16: GSM8K strict-match **+5 to +6 items** (see note ᢜ β€” two
runs exist), GPQA **+8 items**, and long-context retrieval **identical** to bf16 at ~107k-token
prompts.
All AA-LCR figures are the runner's own judging pass, taken from each arm's `score.json`. A
second judging pass over the same generations moves scores by roughly one item in either
direction; mixing passes between arms would manufacture differences that are not there.
Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and
will be revised as they land; the commit history is the record of what was known when.
## Reproducing the quantisation
Data-free, CPU-only, file-to-file β€” no calibration set, no GPU, ~3 minutes for this model.
```python
from quark.torch.export.api import direct_quantize_checkpoint
EXCLUDE = [
"lm_head", "*embed_tokens*",
"*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj",
"*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*",
"*.linear_attn.conv1d", "*.linear_attn.norm",
"*.mlp.gate", "*.mlp.shared_expert_gate",
"mtp*", "*visual*", "*vision*",
]
```
Two things that are easy to get wrong:
- **`*.mlp.gate` and `*.mlp.gate_proj` are different modules.** The first is the MoE router and
must stay bf16; the second is the SwiGLU gate projection and *should* be 4-bit. A glob that
catches both silently quantises the router.
- **When verifying, key on real artifacts**, not on a `_scale` suffix. Several bf16 checkpoints
in this family ship tensors like `vision_tower.std_scale` or per-layer `layer_scalar` in the
*original* weights, and a naive check reports leaks on a perfectly correct build.
Check both directions β€” leakage (something quantised that should not be) *and* over-exclusion
(projections that were meant to be 4-bit but stayed bf16) β€” and make a mismatch raise.
Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid
sliding-attention model) are in the port repo; none of the exclude lists transfer between
families.
## Licence and attribution
Apache-2.0, inherited from [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B). The
`LICENSE` file here is byte-identical to upstream's.
**Modification made:** weights of the MLP and linear-attention projections converted from bf16 to
MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to
architecture, tokenizer or chat template. All other tensors are the upstream values.
The gfx1201 enablement this port descends from was first done by **Rob Smith (`tcclaviger`)** on
the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his
`_matmul_fp8_ogs`. See the `NOTICE` in the port repo for the full lineage.