Image-Text-to-Text
Transformers
Safetensors
qwen3_5
4-bit precision
mxfp4
quark
amd
rocm
rdna4
gfx1201
vllm
quantized
conversational
8-bit precision
Instructions to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4") model = AutoModelForMultimodalLM.from_pretrained("Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4
- SGLang
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with Docker Model Runner:
docker model run hf.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4
File size: 11,984 Bytes
5b96a97 f5e378e 5b96a97 40ecee5 5b96a97 f5e378e 5b96a97 f5e378e 5b96a97 f5e378e 5b96a97 f5e378e 5b96a97 8dc02b8 5b96a97 8dc02b8 facf024 40ecee5 5b96a97 8dc02b8 40ecee5 facf024 40ecee5 facf024 5b96a97 8dc02b8 5b96a97 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 | ---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- mxfp4
- quark
- amd
- rocm
- rdna4
- gfx1201
- vllm
- quantized
---
# Qwen3.8-27B β MXFP4 (AMD Quark) for RDNA4
MXFP4 weight quantisation of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B),
built with [AMD Quark](https://quark.docs.amd.com) 0.12.post1 and targeted at **RDNA4**
(gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) β GPUs that sit outside the official ROCm
vLLM target list.
**What this buys you on 2Γ32 GB RDNA4:** the full **262,144-token** context window at
roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get
long β see Throughput. Quality is at or above the bf16 reference on every cell measured so far
except one, which is stated below rather than omitted.
## What is and is not quantised
Only **MLP and MoE-expert projections** go to 4-bit. Attention (q/k/v/o and its norms), every
norm, embeddings, `lm_head`, routers/gates and the **entire vision path** stay bf16.
| | count |
|---|---|
| `mlp.{gate,up,down}_proj` | 192 (64 layers Γ 3) |
| `linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj}` | 240 (48 layers Γ 5) |
| **total quantised modules** | **432** |
| attention / norms / embeddings / `lm_head` / vision | **0** β verified, none |
Verified by tensor inspection: a module counts as quantised only if it carries a real artifact
(`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`).
Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding
attention costs little size and keeps those layers on the fast bf16 path.
For structural comparison, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the decoder's attention as
well β 496 quantised modules against 432 here, the difference being exactly the 16 full-attention
layers' q/k/v/o β and is **AWQ-calibrated** (`algo_config.name = awq`) where this build is data-free
RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that
build, which is not enough to publish a quality comparison from: strict-match moves by about
Β±0.06 across seeds on this hardware, which is wider than any gap it showed.
- Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized`
- Size: **22.3 GB** across 18 shards (bf16 source β 54 GB)
- Quark `exclude` list: 231 entries
> **The config declares W4A4, not weight-only.** Quark's `mxfp4` scheme enables dynamic fp4
> *activation* quantization by default, so `global_quant_config.input_tensors` reads
> `{dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}`. On the RDNA4 port that
> declaration is **not honoured** β the weight-only kernel ignores activation quant, and the
> FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on
> a runtime that *does* honour it, you will get a different numerical path than the one measured
> here. The difference from whole-decoder AMD-style builds is **coverage** (432 quantized modules
> vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width.
## Serving
Speed and window claims here need the RDNA4 port, which has the MXFP4Γe4m3 FP8-WMMA kernel:
```bash
docker run --rm -it --device /dev/kfd --device /dev/dri \
-v /path/to/weights:/model:ro -p 8011:8011 \
-e VLLM_RDNA_MXFP4_FP8=1 \
capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \
serve /model --served-model-name qwen --port 8011 \
--tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code
```
Source: [`Capicua25x/vllm-rocm-rdna4`](https://github.com/Capicua25x/vllm-rocm-rdna4), branch
`rdna4-port-0.26.1`. `VLLM_RDNA_MXFP4_FP8=0` falls back to the weight-only bf16-unpack kernel.
> **On stock vLLM these weights load and generate correctly, but slower.** Without the
> FP8-WMMA kernel you get the weight-only dequant path β roughly 51 tok/s single-stream instead
> of 61 on this hardware β and on 32 GB cards you will not reach the 262k window. If you are
> benchmarking this against another quant, check which kernel you are actually on first.
Sampling follows the base model card: thinking `temp 1.0, top_p 0.95, top_k 20, min_p 0`;
non-thinking `temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5`.
## Throughput
Measured on **this exact artifact**, 2026-08-17, on 2Γ Radeon AI PRO R9700 (TP2, gfx1201) with the
rc6 FP8-WMMA kernel (`VLLM_RDNA_MXFP4_FP8=1`) and native MTP-3 speculative decoding.
`max_tokens: 256`, thinking **on** β the shape most deployments actually run.
Compared against stock [`Qwen/Qwen3.8-27B-FP8`](https://huggingface.co/Qwen/Qwen3.8-27B-FP8)
**at matched capacity**: both configurations hold a 262,144-token window on the same two cards,
so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point β 131k
window β and was only partially swept; it is not compared here.)
**Short prompt (~30 tokens)** β per-user tok/s / aggregate tok/s:
| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | 47.7 / 48 | 48.1 / 48 |
| 8 | 28.6 / 213 | **37.1 / 278** |
| 16 | **26.5 / 384** | 21.5 / 322 |
| 32 | 18.2 / **539** | **21.2** / 435 |
**6k-token prompt** β closer to a real application's context:
| concurrent | MXFP4 (this build) | FP8 + fp8 KV |
|---|---|---|
| 1 | **46.3 / 46** | 36.0 / 36 |
| 8 | **26.2 / 199** | 10.1 / 79 |
| 16 | **16.9 / 260** | 5.5 / 87 |
**The 6k table is the one that matters.** At short prompts the two are close, and FP8 is ahead at
8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply β 79 tok/s
aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 β while this build
holds its single-stream rate almost unchanged (47.7 β 46.3). If you are serving anything with a
system prompt, retrieved context or conversation history, that is the regime you will be in.
Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross
it by 32 concurrent.
**Thinking-off is not yet measured on these weights.** Figures published elsewhere for the rc6
kernel (61 tok/s single-stream, 649 aggregate) were measured on an *earlier* MXFP4 build of this
model, before this Quark build existed β they do not describe this artifact and are omitted
rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be
added here as they are run.
## Quality β measured, as of 2026-08-17
Same harness, same seed (1234), same on-spec sampling across all four columns. **bf16 ref** is
the unquantised model on a hosted endpoint; the two FP8 columns are stock `Qwen/Qwen3.8-27B-FP8`
on this same box, differing only in KV cache dtype.
| benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** |
|---|---|---|---|---|---|
| GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.94β0.96 / 0.92β0.94** αΆ |
| GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** |
| IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** |
| GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** |
| AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** |
| AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 α΅ | 0.800 | **0.780** |
| ΟΒ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** |
| ΟΒ²-bench airline (Pass^1) | 50 | 0.760 | β | β | **0.840** |
| HLE | 120 | 0.3083 | β | β | *running* |
| SWE-bench Verified | 100 | β | β | β | *pending* |
| Terminal-Bench Hard | 44 | β | β | β | *pending* |
α΅ Scored on the 90 items it served; 10 were refused because the prompt exceeded that
configuration's 131k window. Blended over the full 100 it reads 0.720.
αΆ **Two runs of this build exist at the same seed and identical settings** β 0.94/0.92 and
0.96/0.94 β so the honest figure is a range, not a point. The other three columns are single
runs, which is worth knowing before reading small deltas here as real: on this cell one run's
difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items
depending on which run you take.
**ΟΒ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against
the bf16 reference's 0.939 β eight simulations β and sits four behind the FP8 + bf16 KV arm and
three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items
*ahead*. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses.
Retail is still running and will add a third point.
On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling
(`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than
harness noise β but it is one domain and a single-digit item count, not a blanket weakness.
Everything else is at or above bf16: GSM8K strict-match **+5 to +6 items** (see note αΆ β two
runs exist), GPQA **+8 items**, and long-context retrieval **identical** to bf16 at ~107k-token
prompts.
All AA-LCR figures are the runner's own judging pass, taken from each arm's `score.json`. A
second judging pass over the same generations moves scores by roughly one item in either
direction; mixing passes between arms would manufacture differences that are not there.
Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and
will be revised as they land; the commit history is the record of what was known when.
## Reproducing the quantisation
Data-free, CPU-only, file-to-file β no calibration set, no GPU, ~3 minutes for this model.
```python
from quark.torch.export.api import direct_quantize_checkpoint
EXCLUDE = [
"lm_head", "*embed_tokens*",
"*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj",
"*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*",
"*.linear_attn.conv1d", "*.linear_attn.norm",
"*.mlp.gate", "*.mlp.shared_expert_gate",
"mtp*", "*visual*", "*vision*",
]
```
Two things that are easy to get wrong:
- **`*.mlp.gate` and `*.mlp.gate_proj` are different modules.** The first is the MoE router and
must stay bf16; the second is the SwiGLU gate projection and *should* be 4-bit. A glob that
catches both silently quantises the router.
- **When verifying, key on real artifacts**, not on a `_scale` suffix. Several bf16 checkpoints
in this family ship tensors like `vision_tower.std_scale` or per-layer `layer_scalar` in the
*original* weights, and a naive check reports leaks on a perfectly correct build.
Check both directions β leakage (something quantised that should not be) *and* over-exclusion
(projections that were meant to be 4-bit but stayed bf16) β and make a mismatch raise.
Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid
sliding-attention model) are in the port repo; none of the exclude lists transfer between
families.
## Licence and attribution
Apache-2.0, inherited from [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B). The
`LICENSE` file here is byte-identical to upstream's.
**Modification made:** weights of the MLP and linear-attention projections converted from bf16 to
MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to
architecture, tokenizer or chat template. All other tensors are the upstream values.
The gfx1201 enablement this port descends from was first done by **Rob Smith (`tcclaviger`)** on
the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his
`_matmul_fp8_ogs`. See the `NOTICE` in the port repo for the full lineage.
|