Instructions to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4") model = AutoModelForMultimodalLM.from_pretrained("Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4
- SGLang
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with Docker Model Runner:
docker model run hf.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4
Download README.md from Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4: direct link, hf CLI and curl.
- Browser
- Download file 12 kB
-
https://huggingface.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4/resolve/7a364c7f821c7197fbb4253a822236ceafaf4a05/README.md
- Command line
-
hf download hf://Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4@7a364c7f821c7197fbb4253a822236ceafaf4a05/README.md
-
curl -L -o README.md https://huggingface.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4/resolve/7a364c7f821c7197fbb4253a822236ceafaf4a05/README.md
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- mxfp4
- quark
- amd
- rocm
- rdna4
- gfx1201
- vllm
- quantized
Qwen3.8-27B β MXFP4 (AMD Quark) for RDNA4
MXFP4 weight quantisation of Qwen/Qwen3.8-27B,
built with AMD Quark 0.12.post1 and targeted at RDNA4
(gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) β GPUs that sit outside the official ROCm
vLLM target list.
What this buys you on 2Γ32 GB RDNA4: the full 262,144-token context window at roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get long β see Throughput. Quality is at or above the bf16 reference on every cell measured so far except one, which is stated below rather than omitted.
What is and is not quantised
Only MLP and MoE-expert projections go to 4-bit. Attention (q/k/v/o and its norms), every
norm, embeddings, lm_head, routers/gates and the entire vision path stay bf16.
| count | |
|---|---|
mlp.{gate,up,down}_proj |
192 (64 layers Γ 3) |
linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj} |
240 (48 layers Γ 5) |
| total quantised modules | 432 |
attention / norms / embeddings / lm_head / vision |
0 β verified, none |
Verified by tensor inspection: a module counts as quantised only if it carries a real artifact
(weight_scale, weight_packed, qweight, weight_zero_point).
Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding attention costs little size and keeps those layers on the fast bf16 path.
For structural comparison, amd/Qwen3.8-27B-Quark-AWQ-MXFP4 quantises the decoder's attention as
well β 496 quantised modules against 432 here, the difference being exactly the 16 full-attention
layers' q/k/v/o β and is AWQ-calibrated (algo_config.name = awq) where this build is data-free
RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that
build, which is not enough to publish a quality comparison from: strict-match moves by about
Β±0.06 across seeds on this hardware, which is wider than any gap it showed.
- Format: MXFP4 (E2M1 + E8M0 scale per 32 weights),
pack_method: reorder,weight_format: real_quantized - Size: 22.3 GB across 18 shards (bf16 source β 54 GB)
- Quark
excludelist: 231 entries
The config declares W4A4, not weight-only. Quark's
mxfp4scheme enables dynamic fp4 activation quantization by default, soglobal_quant_config.input_tensorsreads{dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}. On the RDNA4 port that declaration is not honoured β the weight-only kernel ignores activation quant, and the FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on a runtime that does honour it, you will get a different numerical path than the one measured here. The difference from whole-decoder AMD-style builds is coverage (432 quantized modules vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width.
Serving
Speed and window claims here need the RDNA4 port, which has the MXFP4Γe4m3 FP8-WMMA kernel:
docker run --rm -it --device /dev/kfd --device /dev/dri \
-v /path/to/weights:/model:ro -p 8011:8011 \
-e VLLM_RDNA_MXFP4_FP8=1 \
capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \
serve /model --served-model-name qwen --port 8011 \
--tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code
Source: Capicua25x/vllm-rocm-rdna4, branch
rdna4-port-0.26.1. VLLM_RDNA_MXFP4_FP8=0 falls back to the weight-only bf16-unpack kernel.
On stock vLLM these weights load and generate correctly, but slower. Without the FP8-WMMA kernel you get the weight-only dequant path β roughly 51 tok/s single-stream instead of 61 on this hardware β and on 32 GB cards you will not reach the 262k window. If you are benchmarking this against another quant, check which kernel you are actually on first.
Sampling follows the base model card: thinking temp 1.0, top_p 0.95, top_k 20, min_p 0;
non-thinking temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5.
Throughput
Measured on this exact artifact, 2026-08-17, on 2Γ Radeon AI PRO R9700 (TP2, gfx1201) with the
rc6 FP8-WMMA kernel (VLLM_RDNA_MXFP4_FP8=1) and native MTP-3 speculative decoding.
max_tokens: 256, thinking on β the shape most deployments actually run.
Compared against stock Qwen/Qwen3.8-27B-FP8
at matched capacity: both configurations hold a 262,144-token window on the same two cards,
so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point β 131k
window β and was only partially swept; it is not compared here.)
Short prompt (~30 tokens) β per-user tok/s / aggregate tok/s:
| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | 47.7 / 48 | 48.1 / 48 |
| 8 | 28.6 / 213 | 37.1 / 278 |
| 16 | 26.5 / 384 | 21.5 / 322 |
| 32 | 18.2 / 539 | 21.2 / 435 |
6k-token prompt β closer to a real application's context:
| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | 46.3 / 46 | 36.0 / 36 |
| 8 | 26.2 / 199 | 10.1 / 79 |
| 16 | 16.9 / 260 | 5.5 / 87 |
The 6k table is the one that matters. At short prompts the two are close, and FP8 is ahead at 8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply β 79 tok/s aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 β while this build holds its single-stream rate almost unchanged (47.7 β 46.3). If you are serving anything with a system prompt, retrieved context or conversation history, that is the regime you will be in.
Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross it by 32 concurrent.
Thinking-off is not yet measured on these weights. Figures published elsewhere for the rc6 kernel (61 tok/s single-stream, 649 aggregate) were measured on an earlier MXFP4 build of this model, before this Quark build existed β they do not describe this artifact and are omitted rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be added here as they are run.
Quality β measured, as of 2026-08-17
Same harness, same seed (1234), same on-spec sampling across all four columns. bf16 ref is
the unquantised model on a hosted endpoint; the two FP8 columns are stock Qwen/Qwen3.8-27B-FP8
on this same box, differing only in KV cache dtype.
| benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | MXFP4 (this) |
|---|---|---|---|---|---|
| GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | 0.94β0.96 / 0.92β0.94 αΆ |
| GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 |
| IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 |
| GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | 0.9167 |
| AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | 0.9333 |
| AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 α΅ | 0.800 | 0.780 |
| ΟΒ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | 0.868 |
| ΟΒ²-bench airline (Pass^1) | 50 | 0.760 | β | β | 0.840 |
| HLE | 120 | 0.3083 | β | β | running |
| SWE-bench Verified | 100 | β | β | β | pending |
| Terminal-Bench Hard | 44 | β | β | β | pending |
α΅ Scored on the 90 items it served; 10 were refused because the prompt exceeded that configuration's 131k window. Blended over the full 100 it reads 0.720.
αΆ Two runs of this build exist at the same seed and identical settings β 0.94/0.92 and 0.96/0.94 β so the honest figure is a range, not a point. The other three columns are single runs, which is worth knowing before reading small deltas here as real: on this cell one run's difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items depending on which run you take.
ΟΒ² is domain-split, and the split is the finding. On telecom this build scores 0.868 against the bf16 reference's 0.939 β eight simulations β and sits four behind the FP8 + bf16 KV arm and three behind FP8 + fp8 KV. On airline it scores 0.840 against the reference's 0.760, four items ahead. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses. Retail is still running and will add a third point.
On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling
(too_many_errors, scored 0), so the shortfall is a genuine capability difference rather than
harness noise β but it is one domain and a single-digit item count, not a blanket weakness.
Everything else is at or above bf16: GSM8K strict-match +5 to +6 items (see note αΆ β two runs exist), GPQA +8 items, and long-context retrieval identical to bf16 at ~107k-token prompts.
All AA-LCR figures are the runner's own judging pass, taken from each arm's score.json. A
second judging pass over the same generations moves scores by roughly one item in either
direction; mixing passes between arms would manufacture differences that are not there.
Cells marked running / pending are genuinely unfinished, not withheld. This card is dated and will be revised as they land; the commit history is the record of what was known when.
Reproducing the quantisation
Data-free, CPU-only, file-to-file β no calibration set, no GPU, ~3 minutes for this model.
from quark.torch.export.api import direct_quantize_checkpoint
EXCLUDE = [
"lm_head", "*embed_tokens*",
"*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj",
"*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*",
"*.linear_attn.conv1d", "*.linear_attn.norm",
"*.mlp.gate", "*.mlp.shared_expert_gate",
"mtp*", "*visual*", "*vision*",
]
Two things that are easy to get wrong:
*.mlp.gateand*.mlp.gate_projare different modules. The first is the MoE router and must stay bf16; the second is the SwiGLU gate projection and should be 4-bit. A glob that catches both silently quantises the router.- When verifying, key on real artifacts, not on a
_scalesuffix. Several bf16 checkpoints in this family ship tensors likevision_tower.std_scaleor per-layerlayer_scalarin the original weights, and a naive check reports leaks on a perfectly correct build.
Check both directions β leakage (something quantised that should not be) and over-exclusion (projections that were meant to be 4-bit but stayed bf16) β and make a mismatch raise.
Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid sliding-attention model) are in the port repo; none of the exclude lists transfer between families.
Licence and attribution
Apache-2.0, inherited from Qwen/Qwen3.8-27B. The
LICENSE file here is byte-identical to upstream's.
Modification made: weights of the MLP and linear-attention projections converted from bf16 to MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to architecture, tokenizer or chat template. All other tensors are the upstream values.
The gfx1201 enablement this port descends from was first done by Rob Smith (tcclaviger) on
the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his
_matmul_fp8_ogs. See the NOTICE in the port repo for the full lineage.