Capicua25x's picture
report GSM8K as a range: two runs exist at the same seed and the card published the higher one; align AA-LCR arm B to its score.json (0.800, not the rejudge pass) so every arm uses the same judging pass
8dc02b8 verified
|
Raw History Blame
12 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
  - mxfp4
  - quark
  - amd
  - rocm
  - rdna4
  - gfx1201
  - vllm
  - quantized

Qwen3.8-27B β€” MXFP4 (AMD Quark) for RDNA4

MXFP4 weight quantisation of Qwen/Qwen3.8-27B, built with AMD Quark 0.12.post1 and targeted at RDNA4 (gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) β€” GPUs that sit outside the official ROCm vLLM target list.

What this buys you on 2Γ—32 GB RDNA4: the full 262,144-token context window at roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get long β€” see Throughput. Quality is at or above the bf16 reference on every cell measured so far except one, which is stated below rather than omitted.

What is and is not quantised

Only MLP and MoE-expert projections go to 4-bit. Attention (q/k/v/o and its norms), every norm, embeddings, lm_head, routers/gates and the entire vision path stay bf16.

count
mlp.{gate,up,down}_proj 192 (64 layers Γ— 3)
linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj} 240 (48 layers Γ— 5)
total quantised modules 432
attention / norms / embeddings / lm_head / vision 0 β€” verified, none

Verified by tensor inspection: a module counts as quantised only if it carries a real artifact (weight_scale, weight_packed, qweight, weight_zero_point).

Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding attention costs little size and keeps those layers on the fast bf16 path.

For structural comparison, amd/Qwen3.8-27B-Quark-AWQ-MXFP4 quantises the decoder's attention as well β€” 496 quantised modules against 432 here, the difference being exactly the 16 full-attention layers' q/k/v/o β€” and is AWQ-calibrated (algo_config.name = awq) where this build is data-free RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that build, which is not enough to publish a quality comparison from: strict-match moves by about Β±0.06 across seeds on this hardware, which is wider than any gap it showed.

  • Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), pack_method: reorder, weight_format: real_quantized
  • Size: 22.3 GB across 18 shards (bf16 source β‰ˆ 54 GB)
  • Quark exclude list: 231 entries

The config declares W4A4, not weight-only. Quark's mxfp4 scheme enables dynamic fp4 activation quantization by default, so global_quant_config.input_tensors reads {dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}. On the RDNA4 port that declaration is not honoured β€” the weight-only kernel ignores activation quant, and the FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on a runtime that does honour it, you will get a different numerical path than the one measured here. The difference from whole-decoder AMD-style builds is coverage (432 quantized modules vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width.

Serving

Speed and window claims here need the RDNA4 port, which has the MXFP4Γ—e4m3 FP8-WMMA kernel:

docker run --rm -it --device /dev/kfd --device /dev/dri \
  -v /path/to/weights:/model:ro -p 8011:8011 \
  -e VLLM_RDNA_MXFP4_FP8=1 \
  capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \
  serve /model --served-model-name qwen --port 8011 \
  --tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code

Source: Capicua25x/vllm-rocm-rdna4, branch rdna4-port-0.26.1. VLLM_RDNA_MXFP4_FP8=0 falls back to the weight-only bf16-unpack kernel.

On stock vLLM these weights load and generate correctly, but slower. Without the FP8-WMMA kernel you get the weight-only dequant path β€” roughly 51 tok/s single-stream instead of 61 on this hardware β€” and on 32 GB cards you will not reach the 262k window. If you are benchmarking this against another quant, check which kernel you are actually on first.

Sampling follows the base model card: thinking temp 1.0, top_p 0.95, top_k 20, min_p 0; non-thinking temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5.

Throughput

Measured on this exact artifact, 2026-08-17, on 2Γ— Radeon AI PRO R9700 (TP2, gfx1201) with the rc6 FP8-WMMA kernel (VLLM_RDNA_MXFP4_FP8=1) and native MTP-3 speculative decoding. max_tokens: 256, thinking on β€” the shape most deployments actually run.

Compared against stock Qwen/Qwen3.8-27B-FP8 at matched capacity: both configurations hold a 262,144-token window on the same two cards, so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point β€” 131k window β€” and was only partially swept; it is not compared here.)

Short prompt (~30 tokens) β€” per-user tok/s / aggregate tok/s:

concurrent MXFP4 (this build) FP8 + fp8 KV
1 47.7 / 48 48.1 / 48
8 28.6 / 213 37.1 / 278
16 26.5 / 384 21.5 / 322
32 18.2 / 539 21.2 / 435

6k-token prompt β€” closer to a real application's context:

concurrent MXFP4 (this build) FP8 + fp8 KV
1 46.3 / 46 36.0 / 36
8 26.2 / 199 10.1 / 79
16 16.9 / 260 5.5 / 87

The 6k table is the one that matters. At short prompts the two are close, and FP8 is ahead at 8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply β€” 79 tok/s aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 β€” while this build holds its single-stream rate almost unchanged (47.7 β†’ 46.3). If you are serving anything with a system prompt, retrieved context or conversation history, that is the regime you will be in.

Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross it by 32 concurrent.

Thinking-off is not yet measured on these weights. Figures published elsewhere for the rc6 kernel (61 tok/s single-stream, 649 aggregate) were measured on an earlier MXFP4 build of this model, before this Quark build existed β€” they do not describe this artifact and are omitted rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be added here as they are run.

Quality β€” measured, as of 2026-08-17

Same harness, same seed (1234), same on-spec sampling across all four columns. bf16 ref is the unquantised model on a hosted endpoint; the two FP8 columns are stock Qwen/Qwen3.8-27B-FP8 on this same box, differing only in KV cache dtype.

benchmark n bf16 ref FP8 + bf16 KV FP8 + fp8 KV MXFP4 (this)
GSM8K, thinking (flex / strict) 50 0.96 / 0.82 0.96 / 0.90 0.94 / 0.70 0.94–0.96 / 0.92–0.94 ᢜ
GSM8K, no thinking (flex / strict) 50 0.98 / 0.98 0.98 / 0.98 0.98 / 0.98 0.98 / 0.98
IFEval (inst / prompt, strict) 80 .9688 / .9500 .9688 / .9500 .9688 / .9500 .9688 / .9500
GPQA-Diamond (flexible) 60 0.7833 0.8333 0.8333 0.9167
AIME 2025 30 0.9333 1.0000 0.9667 0.9333
AA-LCR (~107k-token prompts, judge-scored) 100 0.780 0.800 ᡃ 0.800 0.780
τ²-bench telecom (Pass^1) 114 0.939 0.904 0.895 0.868
τ²-bench airline (Pass^1) 50 0.760 β€” β€” 0.840
HLE 120 0.3083 β€” β€” running
SWE-bench Verified 100 β€” β€” β€” pending
Terminal-Bench Hard 44 β€” β€” β€” pending

ᡃ Scored on the 90 items it served; 10 were refused because the prompt exceeded that configuration's 131k window. Blended over the full 100 it reads 0.720.

ᢜ Two runs of this build exist at the same seed and identical settings β€” 0.94/0.92 and 0.96/0.94 β€” so the honest figure is a range, not a point. The other three columns are single runs, which is worth knowing before reading small deltas here as real: on this cell one run's difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items depending on which run you take.

τ² is domain-split, and the split is the finding. On telecom this build scores 0.868 against the bf16 reference's 0.939 β€” eight simulations β€” and sits four behind the FP8 + bf16 KV arm and three behind FP8 + fp8 KV. On airline it scores 0.840 against the reference's 0.760, four items ahead. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses. Retail is still running and will add a third point.

On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling (too_many_errors, scored 0), so the shortfall is a genuine capability difference rather than harness noise β€” but it is one domain and a single-digit item count, not a blanket weakness.

Everything else is at or above bf16: GSM8K strict-match +5 to +6 items (see note ᢜ β€” two runs exist), GPQA +8 items, and long-context retrieval identical to bf16 at ~107k-token prompts.

All AA-LCR figures are the runner's own judging pass, taken from each arm's score.json. A second judging pass over the same generations moves scores by roughly one item in either direction; mixing passes between arms would manufacture differences that are not there.

Cells marked running / pending are genuinely unfinished, not withheld. This card is dated and will be revised as they land; the commit history is the record of what was known when.

Reproducing the quantisation

Data-free, CPU-only, file-to-file β€” no calibration set, no GPU, ~3 minutes for this model.

from quark.torch.export.api import direct_quantize_checkpoint

EXCLUDE = [
    "lm_head", "*embed_tokens*",
    "*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj",
    "*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*",
    "*.linear_attn.conv1d", "*.linear_attn.norm",
    "*.mlp.gate", "*.mlp.shared_expert_gate",
    "mtp*", "*visual*", "*vision*",
]

Two things that are easy to get wrong:

  • *.mlp.gate and *.mlp.gate_proj are different modules. The first is the MoE router and must stay bf16; the second is the SwiGLU gate projection and should be 4-bit. A glob that catches both silently quantises the router.
  • When verifying, key on real artifacts, not on a _scale suffix. Several bf16 checkpoints in this family ship tensors like vision_tower.std_scale or per-layer layer_scalar in the original weights, and a naive check reports leaks on a perfectly correct build.

Check both directions β€” leakage (something quantised that should not be) and over-exclusion (projections that were meant to be 4-bit but stayed bf16) β€” and make a mismatch raise.

Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid sliding-attention model) are in the port repo; none of the exclude lists transfer between families.

Licence and attribution

Apache-2.0, inherited from Qwen/Qwen3.8-27B. The LICENSE file here is byte-identical to upstream's.

Modification made: weights of the MLP and linear-attention projections converted from bf16 to MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to architecture, tokenizer or chat template. All other tensors are the upstream values.

The gfx1201 enablement this port descends from was first done by Rob Smith (tcclaviger) on the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his _matmul_fp8_ogs. See the NOTICE in the port repo for the full lineage.