Qwen3.5-27B-NVFP4-QAD-W4A4 (final, step 2640) — W4A4 serving schema

NVFP4 W4A4-serving export (a4schema: NVFP4 weights + trained activation input scales) of the NV-W4A4 QAD arm final checkpoint from the stage-3 W4A4-vs-W4A16 training study on dense Qwen/Qwen3.5-27B.

Training provenance

  • Recipe: QAD temp-1.0 pure-KL distillation (teacher = BF16 Qwen3.5-27B), QATFactory @ 8e7cf1a + the _save_checkpoint empty_cache+barrier fix.
  • Config: configs/llm/qwen3_5_27b_nvfp4_qad_w4a4.yaml (W4A4 fake-quant training), lr 1.0e-5 cosine (1% warmup), 2640 steps (~1 epoch), global batch 32 (bs 4/GPU x 8 GPUs x accum 1), seq len 8192, seed 42, BF16, FSDP2, save/250, eval/100.
  • Data: openperfectblend_100k Qwen3.5-9B-think train/eval JSONLs (train md5 0406bb3a7a482352360716a1bc5e9e04; disjoint eval split).
  • Hardware: 8x B200 (research-common-b200-jbom-02). Completed 2026-07-25 12:45:02 PDT, exit 0 — the run trained through a 16.7h site network partition without interruption (the W&B "crashed" flag on this run's page reflects lost heartbeats only).

Result headline

Final held-out eval (epoch 1.0, step 2640): fwd-KL 0.01347 vs the BF16 teacher (top-1 agreement 0.963). The KL plateaued at ~0.0135 from step 1700 onward. The matched NV-W4A16-trained twin reached fwd-KL 0.005815 at the same budget — for NVFP4, W4A16 training beats W4A4 training by 2.3x on serving KL, so this artifact primarily serves as the study's comparison point (see Qwen3.5-27B-NVFP4-QAD-W4A16-LR1e-5-s2640 for the better NVFP4 checkpoint).

Reproduction (export step)

python scripts/export_nvfp4_vllm.py \
  --source  /scratch/wxu/mxfp4/outputs/27b-nvfp4-w4a4-lr1.0e-5-2640 \
  --model-assets /scratch/wxu/mxfp4/models/Qwen3.5-27B \
  --output  /scratch/wxu/mxfp4/exports/27b-nvw4a4-s2640-a4schema \
  --device cuda:0

Load-smoked with vLLM 0.25.1: loads clean and completes text correctly. First engine start pays a FlashInfer fp4_gemm autotune (~15-25 min cold).

Benchmarks

7-benchmark protocol results will be added when the arm-final matrix wave completes.

Downloads last month
15
Safetensors
Model size
17B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for weili-0234/Qwen3.5-27B-NVFP4-QAD-W4A4-LR1e-5-s2640

Base model

Qwen/Qwen3.5-27B
Quantized
(230)
this model