Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000

MXFP4 checkpoint of Qwen/Qwen3.5-27B after 1,000 steps of W4A4 quantization-aware distillation (QAD, weights and activations fake-quantized during training), exported to vLLM compressed-tensors format. One artifact serves both modes: W4A4 (native MXFP4 activations, FlashInfer, SM100+, default load) and W4A16 (--linear-backend marlin).

How this checkpoint was produced

Producing repo tonyzhang-together/QATFactory branch weili/qwen35-27b @ 8e7cf1a
Base model Qwen/Qwen3.5-27B (dense qwen3_5, bf16)
Training QAD distillation from the BF16 base model as teacher: temperature-1.0 pure-KL loss, assistant-only masking, W4A4 fake-quant (MXFP4 weights E2M1 block-32 EVEN-rule E8M0 scales + dynamic MXFP4 activation fake-quant) on all attention + MLP linears. lr 1.0e-5 cosine→0 (1% warmup), global batch 32 (micro 4 × 8 GPU × accum 1), seq 8192, bf16 FSDP2 full-shard, seed 42, 2,640-step (≈1-epoch) schedule. Data: OpenPerfectBlend-derived ChatML distillation corpus, 100k conversations (train jsonl md5 0406bb3a7a482352360716a1bc5e9e04), held-out eval split disjoint. Hardware: 1×8 B200 (jbom-5c-01).
Step Global step 1000 of 2640 (~38% of one epoch).
Provenance caveat The producing run crashed during the optimizer-state save at s1000 (caching-allocator/NCCL issue, fixed in later runs); the model weights were completely written and verified. This artifact is therefore export-complete but not resume-capable. Same-seed reruns of this arm reproduced trajectories within ~1% (verified three times).
Role in the study s1000 waypoint of the MX-W4A4 arm of the stage-3 Qwen3.5-27B 2×2 study (W4A4-training vs W4A16-training × MXFP4 vs NVFP4). PTQ baselines: weili-0234/Qwen3.5-27B-MXFP4-RTN, weili-0234/Qwen3.5-27B-NVFP4-RTN.

Exporter invocation

# on the training node, from the salvaged model-only checkpoint:
python scripts/export_mxfp4_vllm.py \
  --source /scratch/wxu/mxfp4/salvage/mx-w4a4-modelonly-s1000 \
  --model-assets Qwen3.5-27B \
  --output exports/27b-mxw4a4-s1000 --device cuda:0

Quantized modules: all attention + MLP linears (q/k/v/o_proj, gate/up/down_proj); embeddings/lm_head/norms bf16. MXFP4 activation quantization at serve time is fully dynamic — no calibration state exists in this artifact.

Serving

vllm serve weili-0234/Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000                          # W4A4 (SM100+)
vllm serve weili-0234/Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000 --linear-backend marlin  # W4A16

Evaluation results (2026-07-24)

vLLM on 1×B200, identical harness as the RTN baseline rows. GSM8K = 500-question subset; GPQA-Diamond single seed (±2–3 pts noise); MMLU-Pro full. KL = mean top-20 token KL vs the BF16 teacher over a held-out 256-conversation corpus (818,944 scored tokens). All rows W4A4 serving (native MXFP4 activations).

checkpoint KL vs BF16 GSM8K GPQA-D MMLU-Pro
this repo (QAD-W4A4 s1000) 0.0559 92.8 80.3 85.2
MXFP4-RTN (no training) 0.0907 92.0 79.3 83.3
BF16 teacher 0 93.6 83.3 87.1

1,000 steps of W4A4 QAD recover a large share of the RTN→BF16 gap at A4 serving: −38% KL, and improvements on every measured benchmark. A wider 7-benchmark evaluation (RULER, AIME'25, GPQA-D, MMLU-Pro, HumanEval, IFEval, perplexity) against RTN and GPTQ baselines is in progress and will be added to this card.

Downloads last month
25
Safetensors
Model size
16B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for weili-0234/Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000

Base model

Qwen/Qwen3.5-27B
Quantized
(230)
this model