--- license: apache-2.0 base_model: Qwen/Qwen3.8-27B base_model_relation: quantized pipeline_tag: image-text-to-text library_name: transformers tags: - mxfp4 - quark - amd - rocm - rdna4 - gfx1201 - vllm - quantized --- # Qwen3.8-27B — MXFP4 (AMD Quark) for RDNA4 MXFP4 weight quantisation of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B), built with [AMD Quark](https://quark.docs.amd.com) 0.12.post1 and targeted at **RDNA4** (gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) — GPUs that sit outside the official ROCm vLLM target list. **What this buys you on 2×32 GB RDNA4:** the full **262,144-token** context window at roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get long — see Throughput. Quality is at or above the bf16 reference on every cell measured so far except one, which is stated below rather than omitted. ## What is and is not quantised Only **MLP and MoE-expert projections** go to 4-bit. Attention (q/k/v/o and its norms), every norm, embeddings, `lm_head`, routers/gates and the **entire vision path** stay bf16. | | count | |---|---| | `mlp.{gate,up,down}_proj` | 192 (64 layers × 3) | | `linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj}` | 240 (48 layers × 5) | | **total quantised modules** | **432** | | attention / norms / embeddings / `lm_head` / vision | **0** — verified, none | Verified by tensor inspection: a module counts as quantised only if it carries a real artifact (`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`). Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding attention costs little size and keeps those layers on the fast bf16 path. For structural comparison, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the decoder's attention as well — 496 quantised modules against 432 here, the difference being exactly the 16 full-attention layers' q/k/v/o — and is **AWQ-calibrated** (`algo_config.name = awq`) where this build is data-free RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that build, which is not enough to publish a quality comparison from: strict-match moves by about ±0.06 across seeds on this hardware, which is wider than any gap it showed. - Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized` - Size: **22.3 GB** across 18 shards (bf16 source ≈ 54 GB) - Quark `exclude` list: 231 entries > **The config declares W4A4, not weight-only.** Quark's `mxfp4` scheme enables dynamic fp4 > *activation* quantization by default, so `global_quant_config.input_tensors` reads > `{dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}`. On the RDNA4 port that > declaration is **not honoured** — the weight-only kernel ignores activation quant, and the > FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on > a runtime that *does* honour it, you will get a different numerical path than the one measured > here. The difference from whole-decoder AMD-style builds is **coverage** (432 quantized modules > vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width. ## Serving Speed and window claims here need the RDNA4 port, which has the MXFP4×e4m3 FP8-WMMA kernel: ```bash docker run --rm -it --device /dev/kfd --device /dev/dri \ -v /path/to/weights:/model:ro -p 8011:8011 \ -e VLLM_RDNA_MXFP4_FP8=1 \ capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \ serve /model --served-model-name qwen --port 8011 \ --tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code ``` Source: [`Capicua25x/vllm-rocm-rdna4`](https://github.com/Capicua25x/vllm-rocm-rdna4), branch `rdna4-port-0.26.1`. `VLLM_RDNA_MXFP4_FP8=0` falls back to the weight-only bf16-unpack kernel. > **On stock vLLM these weights load and generate correctly, but slower.** Without the > FP8-WMMA kernel you get the weight-only dequant path — roughly 51 tok/s single-stream instead > of 61 on this hardware — and on 32 GB cards you will not reach the 262k window. If you are > benchmarking this against another quant, check which kernel you are actually on first. Sampling follows the base model card: thinking `temp 1.0, top_p 0.95, top_k 20, min_p 0`; non-thinking `temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5`. ## Throughput Measured on **this exact artifact**, 2026-08-17, on 2× Radeon AI PRO R9700 (TP2, gfx1201) with the rc6 FP8-WMMA kernel (`VLLM_RDNA_MXFP4_FP8=1`) and native MTP-3 speculative decoding. `max_tokens: 256`, thinking **on** — the shape most deployments actually run. Compared against stock [`Qwen/Qwen3.8-27B-FP8`](https://huggingface.co/Qwen/Qwen3.8-27B-FP8) **at matched capacity**: both configurations hold a 262,144-token window on the same two cards, so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point — 131k window — and was only partially swept; it is not compared here.) **Short prompt (~30 tokens)** — per-user tok/s / aggregate tok/s: | concurrent | MXFP4 (this build) | FP8 + fp8 KV | |---|---|---| | 1 | 47.7 / 48 | 48.1 / 48 | | 8 | 28.6 / 213 | **37.1 / 278** | | 16 | **26.5 / 384** | 21.5 / 322 | | 32 | 18.2 / **539** | **21.2** / 435 | **6k-token prompt** — closer to a real application's context: | concurrent | MXFP4 (this build) | FP8 + fp8 KV | |---|---|---| | 1 | **46.3 / 46** | 36.0 / 36 | | 8 | **26.2 / 199** | 10.1 / 79 | | 16 | **16.9 / 260** | 5.5 / 87 | **The 6k table is the one that matters.** At short prompts the two are close, and FP8 is ahead at 8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply — 79 tok/s aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 — while this build holds its single-stream rate almost unchanged (47.7 → 46.3). If you are serving anything with a system prompt, retrieved context or conversation history, that is the regime you will be in. Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross it by 32 concurrent. **Thinking-off is not yet measured on these weights.** Figures published elsewhere for the rc6 kernel (61 tok/s single-stream, 649 aggregate) were measured on an *earlier* MXFP4 build of this model, before this Quark build existed — they do not describe this artifact and are omitted rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be added here as they are run. ## Quality — measured, as of 2026-08-17 Same harness, same seed (1234), same on-spec sampling across all four columns. **bf16 ref** is the unquantised model on a hosted endpoint; the two FP8 columns are stock `Qwen/Qwen3.8-27B-FP8` on this same box, differing only in KV cache dtype. | benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** | |---|---|---|---|---|---| | GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.94–0.96 / 0.92–0.94** ᶜ | | GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** | | IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** | | GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** | | AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** | | AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 ᵃ | 0.800 | **0.780** | | τ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** | | τ²-bench airline (Pass^1) | 50 | 0.760 | — | — | **0.840** | | HLE | 120 | 0.3083 | — | — | *running* | | SWE-bench Verified | 100 | — | — | — | *pending* | | Terminal-Bench Hard | 44 | — | — | — | *pending* | ᵃ Scored on the 90 items it served; 10 were refused because the prompt exceeded that configuration's 131k window. Blended over the full 100 it reads 0.720. ᶜ **Two runs of this build exist at the same seed and identical settings** — 0.94/0.92 and 0.96/0.94 — so the honest figure is a range, not a point. The other three columns are single runs, which is worth knowing before reading small deltas here as real: on this cell one run's difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items depending on which run you take. **τ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against the bf16 reference's 0.939 — eight simulations — and sits four behind the FP8 + bf16 KV arm and three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items *ahead*. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses. Retail is still running and will add a third point. On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling (`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than harness noise — but it is one domain and a single-digit item count, not a blanket weakness. Everything else is at or above bf16: GSM8K strict-match **+5 to +6 items** (see note ᶜ — two runs exist), GPQA **+8 items**, and long-context retrieval **identical** to bf16 at ~107k-token prompts. All AA-LCR figures are the runner's own judging pass, taken from each arm's `score.json`. A second judging pass over the same generations moves scores by roughly one item in either direction; mixing passes between arms would manufacture differences that are not there. Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and will be revised as they land; the commit history is the record of what was known when. ## Reproducing the quantisation Data-free, CPU-only, file-to-file — no calibration set, no GPU, ~3 minutes for this model. ```python from quark.torch.export.api import direct_quantize_checkpoint EXCLUDE = [ "lm_head", "*embed_tokens*", "*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj", "*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*", "*.linear_attn.conv1d", "*.linear_attn.norm", "*.mlp.gate", "*.mlp.shared_expert_gate", "mtp*", "*visual*", "*vision*", ] ``` Two things that are easy to get wrong: - **`*.mlp.gate` and `*.mlp.gate_proj` are different modules.** The first is the MoE router and must stay bf16; the second is the SwiGLU gate projection and *should* be 4-bit. A glob that catches both silently quantises the router. - **When verifying, key on real artifacts**, not on a `_scale` suffix. Several bf16 checkpoints in this family ship tensors like `vision_tower.std_scale` or per-layer `layer_scalar` in the *original* weights, and a naive check reports leaks on a perfectly correct build. Check both directions — leakage (something quantised that should not be) *and* over-exclusion (projections that were meant to be 4-bit but stayed bf16) — and make a mismatch raise. Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid sliding-attention model) are in the port repo; none of the exclude lists transfer between families. ## Licence and attribution Apache-2.0, inherited from [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B). The `LICENSE` file here is byte-identical to upstream's. **Modification made:** weights of the MLP and linear-attention projections converted from bf16 to MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to architecture, tokenizer or chat template. All other tensors are the upstream values. The gfx1201 enablement this port descends from was first done by **Rob Smith (`tcclaviger`)** on the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his `_matmul_fp8_ogs`. See the `NOTICE` in the port repo for the full lineage.