Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000
MXFP4 checkpoint of Qwen/Qwen3.5-27B after 1,000 steps of W4A4
quantization-aware distillation (QAD, weights and activations fake-quantized
during training), exported to vLLM compressed-tensors format. One artifact
serves both modes: W4A4 (native MXFP4 activations, FlashInfer, SM100+, default
load) and W4A16 (--linear-backend marlin).
How this checkpoint was produced
| Producing repo | tonyzhang-together/QATFactory branch weili/qwen35-27b @ 8e7cf1a |
| Base model | Qwen/Qwen3.5-27B (dense qwen3_5, bf16) |
| Training | QAD distillation from the BF16 base model as teacher: temperature-1.0 pure-KL loss, assistant-only masking, W4A4 fake-quant (MXFP4 weights E2M1 block-32 EVEN-rule E8M0 scales + dynamic MXFP4 activation fake-quant) on all attention + MLP linears. lr 1.0e-5 cosine→0 (1% warmup), global batch 32 (micro 4 × 8 GPU × accum 1), seq 8192, bf16 FSDP2 full-shard, seed 42, 2,640-step (≈1-epoch) schedule. Data: OpenPerfectBlend-derived ChatML distillation corpus, 100k conversations (train jsonl md5 0406bb3a7a482352360716a1bc5e9e04), held-out eval split disjoint. Hardware: 1×8 B200 (jbom-5c-01). |
| Step | Global step 1000 of 2640 (~38% of one epoch). |
| Provenance caveat | The producing run crashed during the optimizer-state save at s1000 (caching-allocator/NCCL issue, fixed in later runs); the model weights were completely written and verified. This artifact is therefore export-complete but not resume-capable. Same-seed reruns of this arm reproduced trajectories within ~1% (verified three times). |
| Role in the study | s1000 waypoint of the MX-W4A4 arm of the stage-3 Qwen3.5-27B 2×2 study (W4A4-training vs W4A16-training × MXFP4 vs NVFP4). PTQ baselines: weili-0234/Qwen3.5-27B-MXFP4-RTN, weili-0234/Qwen3.5-27B-NVFP4-RTN. |
Exporter invocation
# on the training node, from the salvaged model-only checkpoint:
python scripts/export_mxfp4_vllm.py \
--source /scratch/wxu/mxfp4/salvage/mx-w4a4-modelonly-s1000 \
--model-assets Qwen3.5-27B \
--output exports/27b-mxw4a4-s1000 --device cuda:0
Quantized modules: all attention + MLP linears (q/k/v/o_proj, gate/up/down_proj); embeddings/lm_head/norms bf16. MXFP4 activation quantization at serve time is fully dynamic — no calibration state exists in this artifact.
Serving
vllm serve weili-0234/Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000 # W4A4 (SM100+)
vllm serve weili-0234/Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000 --linear-backend marlin # W4A16
Evaluation results (2026-07-24)
vLLM on 1×B200, identical harness as the RTN baseline rows. GSM8K = 500-question subset; GPQA-Diamond single seed (±2–3 pts noise); MMLU-Pro full. KL = mean top-20 token KL vs the BF16 teacher over a held-out 256-conversation corpus (818,944 scored tokens). All rows W4A4 serving (native MXFP4 activations).
| checkpoint | KL vs BF16 | GSM8K | GPQA-D | MMLU-Pro |
|---|---|---|---|---|
| this repo (QAD-W4A4 s1000) | 0.0559 | 92.8 | 80.3 | 85.2 |
| MXFP4-RTN (no training) | 0.0907 | 92.0 | 79.3 | 83.3 |
| BF16 teacher | 0 | 93.6 | 83.3 | 87.1 |
1,000 steps of W4A4 QAD recover a large share of the RTN→BF16 gap at A4 serving: −38% KL, and improvements on every measured benchmark. A wider 7-benchmark evaluation (RULER, AIME'25, GPQA-D, MMLU-Pro, HumanEval, IFEval, perplexity) against RTN and GPTQ baselines is in progress and will be added to this card.
- Downloads last month
- 25
Model tree for weili-0234/Qwen3.5-27B-MXFP4-QAD-W4A4-LR1e-5-s1000
Base model
Qwen/Qwen3.5-27B