Qwen3.8 27B rotated INT3 weights for Paiton on RDNA4
These are 3-bit replacement weights for the decoder projections of Qwen/Qwen3.8-27B, calibrated with our own GPTQ run on permissively licensed data. They are an add-on to the current Paiton Qwen3.8 release on one AMD Radeon AI PRO R9700, not a standalone checkpoint. The same target checkpoint supplies all other tensors, and the same DFlash2 drafter is used. Against the previous MXFP4 image, weighted decode rises 19.9 %, prefill rises 5–13 %, and model memory falls from 19.18 to 15.89 GiB. The cost is about 3 points of MMLU-Pro knowledge recall.
GPU: Radeon AI PRO R9700 (gfx1201, 32 GB) · Size: 9.55 GB · Context: 65K profile · Text only
Requirements
| Component | Requirement |
|---|---|
| GPU | One AMD Radeon AI PRO R9700 (gfx1201, RDNA4, 32 GB) |
| Host | Linux x86-64, Python 3, Docker, AMD GPU device access (/dev/kfd, /dev/dri) |
| Runtime image | ghcr.io/eliovp/paiton-vllm-plugin:qwen38-rocm10-vllm029-20260928-r1@sha256:487c97d51e5b4a3fcd0a206e53d842a52dd56a199d8ee3e884f48815093a80d4 |
| Launcher | models/Qwen3.8-MXFP4-DFlash2/run-rocm10-65k.sh in the public plugin repository |
| Target checkpoint | unsloth/Qwen3.8-27B-NVFP4 at f0b7c9e722f5565102fff8481c99e4d86ae099c7 |
| Draft checkpoint | tcclaviger/Qwen3.8-27B-DFlash2-FP8 at ee0cb26a8279b7910cc28d82a8a3e15e4728d56f |
| These weights | This repository at d74ae7d5f5f1b4b45dd12fb1271e3664283a2ec1 |
Use exactly these revisions. The rotated Gated DeltaNet b/a rows were derived from the target revision's bytes.
Quick start
git clone https://github.com/Eliovp-BV/paiton-vllm-plugin.git && cd paiton-vllm-plugin
export PAITON_TARGET_DIR="$PWD/model-cache/qwen38-nvfp4"
export PAITON_DRAFT_DIR="$PWD/model-cache/qwen38-dflash2"
export PAITON_W3ROT_DIR="$PWD/model-cache/qwen38-w3rot-int3"
export PAITON_CACHE_DIR="$PWD/runtime-cache/qwen38-rocm10-65k-w3a4"
mkdir -p "$PAITON_TARGET_DIR" "$PAITON_DRAFT_DIR" "$PAITON_W3ROT_DIR" "$PAITON_CACHE_DIR"
hf download unsloth/Qwen3.8-27B-NVFP4 \
--revision f0b7c9e722f5565102fff8481c99e4d86ae099c7 --local-dir "$PAITON_TARGET_DIR"
hf download tcclaviger/Qwen3.8-27B-DFlash2-FP8 \
--revision ee0cb26a8279b7910cc28d82a8a3e15e4728d56f --local-dir "$PAITON_DRAFT_DIR"
hf download EliovpAI/Qwen3.8-27B-W3Rot-INT3-Paiton-RDNA4 \
--revision d74ae7d5f5f1b4b45dd12fb1271e3664283a2ec1 --local-dir "$PAITON_W3ROT_DIR"
(cd "$PAITON_W3ROT_DIR" && sha256sum -c SHA256SUMS)
bash models/Qwen3.8-MXFP4-DFlash2/run-rocm10.sh
When PAITON_W3ROT_DIR is set, the launcher serves these 3-bit weights automatically. --weights mxfp4
forces the MXFP4 path on the same image. This starts the 65K mode (up to eight requests at once), where these
weights also enable the image's 4-bit KV cache: 1.7× the attention tokens in about the same peak VRAM
(--kv-cache fp8 keeps the FP8 cache). --context 200000 starts the long-context mode on the same image, and
--vision adds image input. Once the server is ready, it accepts OpenAI-compatible requests at
http://127.0.0.1:18982/v1 with API model name Qwen3.8.
Method
- Weights. The MLP, attention and Gated DeltaNet projections of all 64 layers (24.3 billion weights) are symmetric INT3 with one BF16 scale per 128 input elements (3.125 bits per weight).
- Rotation. MLP gate/up, attention q/k/v and Gated DeltaNet
in_projare stored in a block-diagonal Hadamard-128 basis, and the runtime rotates their inputs to match. - Calibration. We ran sequential GPTQ in the rotated basis with an MSE scale search on 292,864 tokens of permissively licensed data. CALIBRATION.md lists the sources.
- Prefill. The rotated projections take 4-bit activations, scaled per token and per 128-element group.
- Everywhere else. The other projections, and all projections during decode, take 8-bit (FP8) activations.
- Execution. Native RDNA4 kernels in the Paiton image run these weights. All other tensors come from the pinned target checkpoint.
Accuracy
These results were measured in-model on 2026-09-26 with the Paiton image, using greedy decoding with thinking off. Each benchmark is paired against the MXFP4 image on identical items. Δ is in points with a normal-approximation 95 % CI, and p is the exact McNemar test.
| Benchmark | MXFP4 | This model (W3A4) | Δ [95 % CI] | McNemar p |
|---|---|---|---|---|
| GSM8K 5-shot (1,319) | 95.68 % | 95.30 % | −0.38 [−1.44, +0.68] | 0.58 |
| HumanEval pass@1 (164) | 95.12 % | 93.90 % | −1.22 [−4.99, +2.56] | 0.75 |
| MMLU-Pro subset, 0-shot direct answer (14 × 100) | 62.57 % | 59.71 % | −2.86 [−4.81, −0.90] | 0.005 |
| Needle at 61,440 tokens (80) | 100 % | 100 % | 0 | 1 |
DFlash2 mean acceptance length relative to MXFP4 changes by +0.1 % on 38 chat prompts, −0.4 % on GSM8K, +2.8 % on HumanEval, +1.8 % on MMLU-Pro and +2.0 % on the needle test.
Math, code and long-context retrieval stay within noise. Knowledge recall drops by about 3 MMLU-Pro points,
which is the trade-off of 3-bit weights. The MXFP4 path (--weights mxfp4) remains the choice for maximum
knowledge accuracy.
Performance
The benchmark is BetterBench 0.6.0 (quick preset) with the 65K release profile on one R9700, using two fresh
server processes per image. The reference is the previous MXFP4 image, qwen38-rocm10-vllm029-65k-20260924-r3.
These runs used the earlier v3 calibration of these weights. Speed does not depend on the calibration data,
because the tensor format and runtime are identical. Changes are computed from unrounded values.
| Metric | MXFP4 | These weights | Change |
|---|---|---|---|
| Weighted decode, tok/s | 153.8 | 184.4 | +19.9 % |
| chat / code / file edit | 121.2 / 179.9 / 179.9 | 135.8 / 226.0 / 195.0 | +12.0 / +25.6 / +8.5 % |
| json / math / prose | 217.5 / 183.8 / 78.5 | 269.0 / 228.9 / 94.7 | +23.7 / +24.5 / +20.7 % |
| reasoning / summarization | 117.8 / 138.4 | 133.5 / 158.4 | +13.3 / +14.4 % |
| Aggregate at 1 / 2 / 4 / 8 requests, tok/s | 122.0 / 204.2 / 308.2 / 425.3 | 148.8 / 249.3 / 368.3 / 492.1 | +22.0 / +22.1 / +19.5 / +15.7 % |
| Prefill at 2K / 8K / 16K, input tok/s | 3,689 / 3,834 / 3,871 | 4,156 / 4,165 / 4,103 | +12.7 / +8.6 / +6.0 % |
| Prefill at 32K / 64K, input tok/s | 3,751 / 3,455 | 3,958 / 3,629 | +5.5 / +5.0 % |
Time to first token shortens by the same factors as prefill throughput. Model memory falls from 19.18 GiB to 15.89 GiB, and the launcher gives the freed VRAM to the KV cache: 250,578 cached tokens instead of 174,634, or 3.8 concurrent 65K-token requests instead of 2.7. Since 28 September, the 65K mode stores that cache in 4 bits by default with these weights: vLLM reports 393,216 tokens, and six 61K-token requests run together. Its accuracy and long-session checks are in the plugin README.
Files and verification
The repository holds 64 layer shards (w3rot-L00.safetensors … w3rot-L63.safetensors, 1,296 tensors),
w3rot-index.json, w3rot-manifest.json, CALIBRATION.md, LICENSE, THIRD_PARTY_NOTICES.md and SHA256SUMS.
The manifest records the format, per-shard SHA-256, b/a source hashes and calibration provenance, including 33
dataset entries with revisions and licenses. The accuracy results above used this manifest, which is published
unchanged. The paths in its sources section refer to the calibration environment.
sha256sum -c SHA256SUMS
sha256sum w3rot-manifest.json # 335c5c834fb3f455203908d2b5aff8131aca2fe64ad73d59feddf32247d41a70
Limitations
- The weights are lossy. MMLU-Pro scores about 3 points below MXFP4, outputs differ from the MXFP4 path, and greedy generations can diverge.
- The evaluation covers GSM8K, HumanEval, a 1,400-question MMLU-Pro subset, 38 chat prompts and needle retrieval at 61,440 tokens. Multilingual, tool-calling and longer-context quality are not measured.
- The 28 September image serves these weights in its 65K and long-context modes. They cover the language model;
with
--vision, image input uses the target checkpoint's vision encoder. One R9700 only; Transformers, stock vLLM and llama.cpp cannot load them. - The launcher selects the 4-bit KV cache only in the 65K mode, where it was measured; the long-context mode and other settings use the FP8 cache.
- The weights are tied to the pinned target revision and to Paiton's rotation and activation formats.
License and attribution
These weights are released under Apache-2.0. They are a quantized derivative of Qwen/Qwen3.8-27B (Apache-2.0), and the rotated b/a rows derive from unsloth/Qwen3.8-27B-NVFP4 (Apache-2.0). The method builds on GPTQ and GSQ (Apache-2.0). The target checkpoint and the DFlash2 drafter (tcclaviger/Qwen3.8-27B-DFlash2-FP8, Apache-2.0) are not redistributed here. The calibration data requires these attributions:
- QASC. Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen and Ashish Sabharwal (Allen Institute for AI), allenai/qasc. It is licensed under CC BY 4.0 and provided as is. We reformatted it into calibration prompts, and it is not redistributed.
- ODC-By datasets. The calibration contains information from FineWeb-Edu, FineWeb-2 and FinePDFs-Edu (Hugging Face). They are made available under the ODC Attribution License 1.0, and their use is also subject to Common Crawl's Terms of Use.
THIRD_PARTY_NOTICES.md has the full notices · Paiton · Public Paiton plugin
Model tree for EliovpAI/Qwen3.8-27B-W3Rot-INT3-Paiton-RDNA4
Base model
Qwen/Qwen3.8-27B