Qwen3.8 27B rotated INT3 weights for Paiton on RDNA4

These are 3-bit replacement weights for the decoder projections of Qwen/Qwen3.8-27B, calibrated with our own GPTQ run on permissively licensed data. They are an add-on to the current Paiton Qwen3.8 release on one AMD Radeon AI PRO R9700, not a standalone checkpoint. The same target checkpoint supplies all other tensors, and the same DFlash2 drafter is used. Against the previous MXFP4 image, weighted decode rises 19.9 %, prefill rises 5–13 %, and model memory falls from 19.18 to 15.89 GiB. The cost is about 3 points of MMLU-Pro knowledge recall.

GPU: Radeon AI PRO R9700 (gfx1201, 32 GB) · Size: 9.55 GB · Context: 65K profile · Text only

Requirements

Component Requirement
GPU One AMD Radeon AI PRO R9700 (gfx1201, RDNA4, 32 GB)
Host Linux x86-64, Python 3, Docker, AMD GPU device access (/dev/kfd, /dev/dri)
Runtime image ghcr.io/eliovp/paiton-vllm-plugin:qwen38-rocm10-vllm029-20260928-r1@sha256:487c97d51e5b4a3fcd0a206e53d842a52dd56a199d8ee3e884f48815093a80d4
Launcher models/Qwen3.8-MXFP4-DFlash2/run-rocm10-65k.sh in the public plugin repository
Target checkpoint unsloth/Qwen3.8-27B-NVFP4 at f0b7c9e722f5565102fff8481c99e4d86ae099c7
Draft checkpoint tcclaviger/Qwen3.8-27B-DFlash2-FP8 at ee0cb26a8279b7910cc28d82a8a3e15e4728d56f
These weights This repository at d74ae7d5f5f1b4b45dd12fb1271e3664283a2ec1

Use exactly these revisions. The rotated Gated DeltaNet b/a rows were derived from the target revision's bytes.

Quick start

git clone https://github.com/Eliovp-BV/paiton-vllm-plugin.git && cd paiton-vllm-plugin

export PAITON_TARGET_DIR="$PWD/model-cache/qwen38-nvfp4"
export PAITON_DRAFT_DIR="$PWD/model-cache/qwen38-dflash2"
export PAITON_W3ROT_DIR="$PWD/model-cache/qwen38-w3rot-int3"
export PAITON_CACHE_DIR="$PWD/runtime-cache/qwen38-rocm10-65k-w3a4"
mkdir -p "$PAITON_TARGET_DIR" "$PAITON_DRAFT_DIR" "$PAITON_W3ROT_DIR" "$PAITON_CACHE_DIR"

hf download unsloth/Qwen3.8-27B-NVFP4 \
  --revision f0b7c9e722f5565102fff8481c99e4d86ae099c7 --local-dir "$PAITON_TARGET_DIR"
hf download tcclaviger/Qwen3.8-27B-DFlash2-FP8 \
  --revision ee0cb26a8279b7910cc28d82a8a3e15e4728d56f --local-dir "$PAITON_DRAFT_DIR"
hf download EliovpAI/Qwen3.8-27B-W3Rot-INT3-Paiton-RDNA4 \
  --revision d74ae7d5f5f1b4b45dd12fb1271e3664283a2ec1 --local-dir "$PAITON_W3ROT_DIR"
(cd "$PAITON_W3ROT_DIR" && sha256sum -c SHA256SUMS)

bash models/Qwen3.8-MXFP4-DFlash2/run-rocm10.sh

When PAITON_W3ROT_DIR is set, the launcher serves these 3-bit weights automatically. --weights mxfp4 forces the MXFP4 path on the same image. This starts the 65K mode (up to eight requests at once), where these weights also enable the image's 4-bit KV cache: 1.7× the attention tokens in about the same peak VRAM (--kv-cache fp8 keeps the FP8 cache). --context 200000 starts the long-context mode on the same image, and --vision adds image input. Once the server is ready, it accepts OpenAI-compatible requests at http://127.0.0.1:18982/v1 with API model name Qwen3.8.

Method

  • Weights. The MLP, attention and Gated DeltaNet projections of all 64 layers (24.3 billion weights) are symmetric INT3 with one BF16 scale per 128 input elements (3.125 bits per weight).
  • Rotation. MLP gate/up, attention q/k/v and Gated DeltaNet in_proj are stored in a block-diagonal Hadamard-128 basis, and the runtime rotates their inputs to match.
  • Calibration. We ran sequential GPTQ in the rotated basis with an MSE scale search on 292,864 tokens of permissively licensed data. CALIBRATION.md lists the sources.
  • Prefill. The rotated projections take 4-bit activations, scaled per token and per 128-element group.
  • Everywhere else. The other projections, and all projections during decode, take 8-bit (FP8) activations.
  • Execution. Native RDNA4 kernels in the Paiton image run these weights. All other tensors come from the pinned target checkpoint.

Accuracy

These results were measured in-model on 2026-09-26 with the Paiton image, using greedy decoding with thinking off. Each benchmark is paired against the MXFP4 image on identical items. Δ is in points with a normal-approximation 95 % CI, and p is the exact McNemar test.

Benchmark MXFP4 This model (W3A4) Δ [95 % CI] McNemar p
GSM8K 5-shot (1,319) 95.68 % 95.30 % −0.38 [−1.44, +0.68] 0.58
HumanEval pass@1 (164) 95.12 % 93.90 % −1.22 [−4.99, +2.56] 0.75
MMLU-Pro subset, 0-shot direct answer (14 × 100) 62.57 % 59.71 % −2.86 [−4.81, −0.90] 0.005
Needle at 61,440 tokens (80) 100 % 100 % 0 1

DFlash2 mean acceptance length relative to MXFP4 changes by +0.1 % on 38 chat prompts, −0.4 % on GSM8K, +2.8 % on HumanEval, +1.8 % on MMLU-Pro and +2.0 % on the needle test.

Math, code and long-context retrieval stay within noise. Knowledge recall drops by about 3 MMLU-Pro points, which is the trade-off of 3-bit weights. The MXFP4 path (--weights mxfp4) remains the choice for maximum knowledge accuracy.

Performance

The benchmark is BetterBench 0.6.0 (quick preset) with the 65K release profile on one R9700, using two fresh server processes per image. The reference is the previous MXFP4 image, qwen38-rocm10-vllm029-65k-20260924-r3. These runs used the earlier v3 calibration of these weights. Speed does not depend on the calibration data, because the tensor format and runtime are identical. Changes are computed from unrounded values.

Metric MXFP4 These weights Change
Weighted decode, tok/s 153.8 184.4 +19.9 %
chat / code / file edit 121.2 / 179.9 / 179.9 135.8 / 226.0 / 195.0 +12.0 / +25.6 / +8.5 %
json / math / prose 217.5 / 183.8 / 78.5 269.0 / 228.9 / 94.7 +23.7 / +24.5 / +20.7 %
reasoning / summarization 117.8 / 138.4 133.5 / 158.4 +13.3 / +14.4 %
Aggregate at 1 / 2 / 4 / 8 requests, tok/s 122.0 / 204.2 / 308.2 / 425.3 148.8 / 249.3 / 368.3 / 492.1 +22.0 / +22.1 / +19.5 / +15.7 %
Prefill at 2K / 8K / 16K, input tok/s 3,689 / 3,834 / 3,871 4,156 / 4,165 / 4,103 +12.7 / +8.6 / +6.0 %
Prefill at 32K / 64K, input tok/s 3,751 / 3,455 3,958 / 3,629 +5.5 / +5.0 %

Time to first token shortens by the same factors as prefill throughput. Model memory falls from 19.18 GiB to 15.89 GiB, and the launcher gives the freed VRAM to the KV cache: 250,578 cached tokens instead of 174,634, or 3.8 concurrent 65K-token requests instead of 2.7. Since 28 September, the 65K mode stores that cache in 4 bits by default with these weights: vLLM reports 393,216 tokens, and six 61K-token requests run together. Its accuracy and long-session checks are in the plugin README.

Files and verification

The repository holds 64 layer shards (w3rot-L00.safetensors … w3rot-L63.safetensors, 1,296 tensors), w3rot-index.json, w3rot-manifest.json, CALIBRATION.md, LICENSE, THIRD_PARTY_NOTICES.md and SHA256SUMS. The manifest records the format, per-shard SHA-256, b/a source hashes and calibration provenance, including 33 dataset entries with revisions and licenses. The accuracy results above used this manifest, which is published unchanged. The paths in its sources section refer to the calibration environment.

sha256sum -c SHA256SUMS
sha256sum w3rot-manifest.json   # 335c5c834fb3f455203908d2b5aff8131aca2fe64ad73d59feddf32247d41a70

Limitations

  • The weights are lossy. MMLU-Pro scores about 3 points below MXFP4, outputs differ from the MXFP4 path, and greedy generations can diverge.
  • The evaluation covers GSM8K, HumanEval, a 1,400-question MMLU-Pro subset, 38 chat prompts and needle retrieval at 61,440 tokens. Multilingual, tool-calling and longer-context quality are not measured.
  • The 28 September image serves these weights in its 65K and long-context modes. They cover the language model; with --vision, image input uses the target checkpoint's vision encoder. One R9700 only; Transformers, stock vLLM and llama.cpp cannot load them.
  • The launcher selects the 4-bit KV cache only in the 65K mode, where it was measured; the long-context mode and other settings use the FP8 cache.
  • The weights are tied to the pinned target revision and to Paiton's rotation and activation formats.

License and attribution

These weights are released under Apache-2.0. They are a quantized derivative of Qwen/Qwen3.8-27B (Apache-2.0), and the rotated b/a rows derive from unsloth/Qwen3.8-27B-NVFP4 (Apache-2.0). The method builds on GPTQ and GSQ (Apache-2.0). The target checkpoint and the DFlash2 drafter (tcclaviger/Qwen3.8-27B-DFlash2-FP8, Apache-2.0) are not redistributed here. The calibration data requires these attributions:

THIRD_PARTY_NOTICES.md has the full notices · Paiton · Public Paiton plugin

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EliovpAI/Qwen3.8-27B-W3Rot-INT3-Paiton-RDNA4

Base model

Qwen/Qwen3.8-27B
Quantized
(1268)
this model