Image-Text-to-Text
Transformers
Safetensors
qwen3_5
4-bit precision
mxfp4
quark
amd
rocm
rdna4
gfx1201
vllm
quantized
conversational
8-bit precision
Instructions to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4") model = AutoModelForMultimodalLM.from_pretrained("Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4
- SGLang
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4 with Docker Model Runner:
docker model run hf.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4
report GSM8K as a range: two runs exist at the same seed and the card published the higher one; align AA-LCR arm B to its score.json (0.800, not the rejudge pass) so every arm uses the same judging pass
8dc02b8 verified |
Download README.md from Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4: direct link, hf CLI and curl.
- Browser
- Download file 12 kB
-
https://huggingface.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4/resolve/7a364c7f821c7197fbb4253a822236ceafaf4a05/README.md
- Command line
-
hf download hf://Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4@7a364c7f821c7197fbb4253a822236ceafaf4a05/README.md
-
curl -L -o README.md https://huggingface.co/Capicua25x/Qwen3.8-27B-MXFP4-Quark-RDNA4/resolve/7a364c7f821c7197fbb4253a822236ceafaf4a05/README.md
12 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen3.8-27B | |
| base_model_relation: quantized | |
| pipeline_tag: image-text-to-text | |
| library_name: transformers | |
| tags: | |
| - mxfp4 | |
| - quark | |
| - amd | |
| - rocm | |
| - rdna4 | |
| - gfx1201 | |
| - vllm | |
| - quantized | |
| # Qwen3.8-27B β MXFP4 (AMD Quark) for RDNA4 | |
| MXFP4 weight quantisation of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B), | |
| built with [AMD Quark](https://quark.docs.amd.com) 0.12.post1 and targeted at **RDNA4** | |
| (gfx1200/gfx1201: Radeon AI PRO R9700, RX 9070 XT) β GPUs that sit outside the official ROCm | |
| vLLM target list. | |
| **What this buys you on 2Γ32 GB RDNA4:** the full **262,144-token** context window at | |
| roughly the single-stream speed of stock FP8, and markedly better throughput once prompts get | |
| long β see Throughput. Quality is at or above the bf16 reference on every cell measured so far | |
| except one, which is stated below rather than omitted. | |
| ## What is and is not quantised | |
| Only **MLP and MoE-expert projections** go to 4-bit. Attention (q/k/v/o and its norms), every | |
| norm, embeddings, `lm_head`, routers/gates and the **entire vision path** stay bf16. | |
| | | count | | |
| |---|---| | |
| | `mlp.{gate,up,down}_proj` | 192 (64 layers Γ 3) | | |
| | `linear_attn.{in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj}` | 240 (48 layers Γ 5) | | |
| | **total quantised modules** | **432** | | |
| | attention / norms / embeddings / `lm_head` / vision | **0** β verified, none | | |
| Verified by tensor inspection: a module counts as quantised only if it carries a real artifact | |
| (`weight_scale`, `weight_packed`, `qweight`, `weight_zero_point`). | |
| Keeping attention in bf16 is deliberate: the MLP stack is where the parameters are, so excluding | |
| attention costs little size and keeps those layers on the fast bf16 path. | |
| For structural comparison, `amd/Qwen3.8-27B-Quark-AWQ-MXFP4` quantises the decoder's attention as | |
| well β 496 quantised modules against 432 here, the difference being exactly the 16 full-attention | |
| layers' q/k/v/o β and is **AWQ-calibrated** (`algo_config.name = awq`) where this build is data-free | |
| RTN. Those are the two real differences. We have a single unrepeated n=50 GSM8K run against that | |
| build, which is not enough to publish a quality comparison from: strict-match moves by about | |
| Β±0.06 across seeds on this hardware, which is wider than any gap it showed. | |
| - Format: MXFP4 (E2M1 + E8M0 scale per 32 weights), `pack_method: reorder`, `weight_format: real_quantized` | |
| - Size: **22.3 GB** across 18 shards (bf16 source β 54 GB) | |
| - Quark `exclude` list: 231 entries | |
| > **The config declares W4A4, not weight-only.** Quark's `mxfp4` scheme enables dynamic fp4 | |
| > *activation* quantization by default, so `global_quant_config.input_tensors` reads | |
| > `{dtype: fp4, is_dynamic: true, per_group, group_size 32, e8m0}`. On the RDNA4 port that | |
| > declaration is **not honoured** β the weight-only kernel ignores activation quant, and the | |
| > FP8-WMMA kernel uses its own per-(token, 32-K-group) dynamic e4m3. If you load these weights on | |
| > a runtime that *does* honour it, you will get a different numerical path than the one measured | |
| > here. The difference from whole-decoder AMD-style builds is **coverage** (432 quantized modules | |
| > vs 496, the delta being the 16 full-attention layers' q/k/v/o), not activation width. | |
| ## Serving | |
| Speed and window claims here need the RDNA4 port, which has the MXFP4Γe4m3 FP8-WMMA kernel: | |
| ```bash | |
| docker run --rm -it --device /dev/kfd --device /dev/dri \ | |
| -v /path/to/weights:/model:ro -p 8011:8011 \ | |
| -e VLLM_RDNA_MXFP4_FP8=1 \ | |
| capicua25x/vllm-rocm-rdna4:0.26.1-rdna4-rc6 \ | |
| serve /model --served-model-name qwen --port 8011 \ | |
| --tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code | |
| ``` | |
| Source: [`Capicua25x/vllm-rocm-rdna4`](https://github.com/Capicua25x/vllm-rocm-rdna4), branch | |
| `rdna4-port-0.26.1`. `VLLM_RDNA_MXFP4_FP8=0` falls back to the weight-only bf16-unpack kernel. | |
| > **On stock vLLM these weights load and generate correctly, but slower.** Without the | |
| > FP8-WMMA kernel you get the weight-only dequant path β roughly 51 tok/s single-stream instead | |
| > of 61 on this hardware β and on 32 GB cards you will not reach the 262k window. If you are | |
| > benchmarking this against another quant, check which kernel you are actually on first. | |
| Sampling follows the base model card: thinking `temp 1.0, top_p 0.95, top_k 20, min_p 0`; | |
| non-thinking `temp 0.7, top_p 0.8, top_k 20, presence_penalty 1.5`. | |
| ## Throughput | |
| Measured on **this exact artifact**, 2026-08-17, on 2Γ Radeon AI PRO R9700 (TP2, gfx1201) with the | |
| rc6 FP8-WMMA kernel (`VLLM_RDNA_MXFP4_FP8=1`) and native MTP-3 speculative decoding. | |
| `max_tokens: 256`, thinking **on** β the shape most deployments actually run. | |
| Compared against stock [`Qwen/Qwen3.8-27B-FP8`](https://huggingface.co/Qwen/Qwen3.8-27B-FP8) | |
| **at matched capacity**: both configurations hold a 262,144-token window on the same two cards, | |
| so this is like-for-like. (Stock FP8 with a bf16 KV cache is a different operating point β 131k | |
| window β and was only partially swept; it is not compared here.) | |
| **Short prompt (~30 tokens)** β per-user tok/s / aggregate tok/s: | |
| | concurrent | MXFP4 (this build) | FP8 + fp8 KV | | |
| |---|---|---| | |
| | 1 | 47.7 / 48 | 48.1 / 48 | | |
| | 8 | 28.6 / 213 | **37.1 / 278** | | |
| | 16 | **26.5 / 384** | 21.5 / 322 | | |
| | 32 | 18.2 / **539** | **21.2** / 435 | | |
| **6k-token prompt** β closer to a real application's context: | |
| | concurrent | MXFP4 (this build) | FP8 + fp8 KV | | |
| |---|---|---| | |
| | 1 | **46.3 / 46** | 36.0 / 36 | | |
| | 8 | **26.2 / 199** | 10.1 / 79 | | |
| | 16 | **16.9 / 260** | 5.5 / 87 | | |
| **The 6k table is the one that matters.** At short prompts the two are close, and FP8 is ahead at | |
| 8 concurrent. But at realistic prompt lengths the FP8 fp8-KV path degrades sharply β 79 tok/s | |
| aggregate at 8 concurrent against this build's 199, and 87 against 260 at 16 β while this build | |
| holds its single-stream rate almost unchanged (47.7 β 46.3). If you are serving anything with a | |
| system prompt, retrieved context or conversation history, that is the regime you will be in. | |
| Per-user rates below ~20 tok/s fall under a usable interactive floor; both configurations cross | |
| it by 32 concurrent. | |
| **Thinking-off is not yet measured on these weights.** Figures published elsewhere for the rc6 | |
| kernel (61 tok/s single-stream, 649 aggregate) were measured on an *earlier* MXFP4 build of this | |
| model, before this Quark build existed β they do not describe this artifact and are omitted | |
| rather than borrowed. Think-off sweeps, and a full sweep of the 131k FP8 configuration, will be | |
| added here as they are run. | |
| ## Quality β measured, as of 2026-08-17 | |
| Same harness, same seed (1234), same on-spec sampling across all four columns. **bf16 ref** is | |
| the unquantised model on a hosted endpoint; the two FP8 columns are stock `Qwen/Qwen3.8-27B-FP8` | |
| on this same box, differing only in KV cache dtype. | |
| | benchmark | n | bf16 ref | FP8 + bf16 KV | FP8 + fp8 KV | **MXFP4 (this)** | | |
| |---|---|---|---|---|---| | |
| | GSM8K, thinking (flex / strict) | 50 | 0.96 / 0.82 | 0.96 / 0.90 | 0.94 / 0.70 | **0.94β0.96 / 0.92β0.94** αΆ | | |
| | GSM8K, no thinking (flex / strict) | 50 | 0.98 / 0.98 | 0.98 / 0.98 | 0.98 / 0.98 | **0.98 / 0.98** | | |
| | IFEval (inst / prompt, strict) | 80 | .9688 / .9500 | .9688 / .9500 | .9688 / .9500 | **.9688 / .9500** | | |
| | GPQA-Diamond (flexible) | 60 | 0.7833 | 0.8333 | 0.8333 | **0.9167** | | |
| | AIME 2025 | 30 | 0.9333 | 1.0000 | 0.9667 | **0.9333** | | |
| | AA-LCR (~107k-token prompts, judge-scored) | 100 | 0.780 | 0.800 α΅ | 0.800 | **0.780** | | |
| | ΟΒ²-bench telecom (Pass^1) | 114 | 0.939 | 0.904 | 0.895 | **0.868** | | |
| | ΟΒ²-bench airline (Pass^1) | 50 | 0.760 | β | β | **0.840** | | |
| | HLE | 120 | 0.3083 | β | β | *running* | | |
| | SWE-bench Verified | 100 | β | β | β | *pending* | | |
| | Terminal-Bench Hard | 44 | β | β | β | *pending* | | |
| α΅ Scored on the 90 items it served; 10 were refused because the prompt exceeded that | |
| configuration's 131k window. Blended over the full 100 it reads 0.720. | |
| αΆ **Two runs of this build exist at the same seed and identical settings** β 0.94/0.92 and | |
| 0.96/0.94 β so the honest figure is a range, not a point. The other three columns are single | |
| runs, which is worth knowing before reading small deltas here as real: on this cell one run's | |
| difference is one item. Against the reference's 0.82 strict, this build is +5 or +6 items | |
| depending on which run you take. | |
| **ΟΒ² is domain-split, and the split is the finding.** On telecom this build scores 0.868 against | |
| the bf16 reference's 0.939 β eight simulations β and sits four behind the FP8 + bf16 KV arm and | |
| three behind FP8 + fp8 KV. On airline it scores **0.840 against the reference's 0.760**, four items | |
| *ahead*. Multi-turn tool use is therefore not uniformly degraded; telecom is where it loses. | |
| Retail is still running and will add a third point. | |
| On telecom, 113 of 114 simulations ended normally and one hit the harness's error ceiling | |
| (`too_many_errors`, scored 0), so the shortfall is a genuine capability difference rather than | |
| harness noise β but it is one domain and a single-digit item count, not a blanket weakness. | |
| Everything else is at or above bf16: GSM8K strict-match **+5 to +6 items** (see note αΆ β two | |
| runs exist), GPQA **+8 items**, and long-context retrieval **identical** to bf16 at ~107k-token | |
| prompts. | |
| All AA-LCR figures are the runner's own judging pass, taken from each arm's `score.json`. A | |
| second judging pass over the same generations moves scores by roughly one item in either | |
| direction; mixing passes between arms would manufacture differences that are not there. | |
| Cells marked *running* / *pending* are genuinely unfinished, not withheld. This card is dated and | |
| will be revised as they land; the commit history is the record of what was known when. | |
| ## Reproducing the quantisation | |
| Data-free, CPU-only, file-to-file β no calibration set, no GPU, ~3 minutes for this model. | |
| ```python | |
| from quark.torch.export.api import direct_quantize_checkpoint | |
| EXCLUDE = [ | |
| "lm_head", "*embed_tokens*", | |
| "*.self_attn.q_proj", "*.self_attn.k_proj", "*.self_attn.v_proj", "*.self_attn.o_proj", | |
| "*.self_attn.q_norm", "*.self_attn.k_norm", "*norm*", | |
| "*.linear_attn.conv1d", "*.linear_attn.norm", | |
| "*.mlp.gate", "*.mlp.shared_expert_gate", | |
| "mtp*", "*visual*", "*vision*", | |
| ] | |
| ``` | |
| Two things that are easy to get wrong: | |
| - **`*.mlp.gate` and `*.mlp.gate_proj` are different modules.** The first is the MoE router and | |
| must stay bf16; the second is the SwiGLU gate projection and *should* be 4-bit. A glob that | |
| catches both silently quantises the router. | |
| - **When verifying, key on real artifacts**, not on a `_scale` suffix. Several bf16 checkpoints | |
| in this family ship tensors like `vision_tower.std_scale` or per-layer `layer_scalar` in the | |
| *original* weights, and a naive check reports leaks on a perfectly correct build. | |
| Check both directions β leakage (something quantised that should not be) *and* over-exclusion | |
| (projections that were meant to be 4-bit but stayed bf16) β and make a mismatch raise. | |
| Per-family recipes for other architectures (Gemma-4 dense and MoE, Mistral-Small, and a hybrid | |
| sliding-attention model) are in the port repo; none of the exclude lists transfer between | |
| families. | |
| ## Licence and attribution | |
| Apache-2.0, inherited from [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B). The | |
| `LICENSE` file here is byte-identical to upstream's. | |
| **Modification made:** weights of the MLP and linear-attention projections converted from bf16 to | |
| MXFP4 via AMD Quark 0.12.post1, as described above. No fine-tuning, no distillation, no change to | |
| architecture, tokenizer or chat template. All other tensors are the upstream values. | |
| The gfx1201 enablement this port descends from was first done by **Rob Smith (`tcclaviger`)** on | |
| the vLLM 0.18.1 line; the FP8-WMMA kernel takes its per-K-group scale-fold design from his | |
| `_matmul_fp8_ogs`. See the `NOTICE` in the port repo for the full lineage. | |