weili-0234's picture
Full provenance model card
beb4a79 verified
|
Raw History Blame Contribute Delete
6.11 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3.5-9B
pipeline_tag: text-generation
tags:
  - nvfp4
  - compressed-tensors
  - quantization
  - qad

Qwen3.5-9B-NVFP4-QAD-LR1e-5-s3000

NVFP4 QAD (W4A16-trained (weight-only)) checkpoint of Qwen/Qwen3.5-9B @ c202236235762e1c871ad0ccb60c8ee5ba337b9a, optimizer step 3000 of 4000.

The best NVFP4 arm (stage-2, weight-only trained at the sweep-picked lr 1e-5). Headline finding: A4 serving is nearly free without any W4A4 training — A4 KL 0.0497 vs A16 0.0255 (gap 0.024, vs MXFP4's 0.083), GSM8K 86.6 (A4) / 86.2 (A16), both above the BF16 teacher's 83.8, at ~1.6x A4 throughput. The W4A4-trained sibling (Qwen3.5-9B-NVFP4-QAD-W4A4-LR1e-5-s4000) did NOT improve A4 serving — with an activation gap this small the noisier W4A4 forward costs more than the train/deploy match recovers.

How this checkpoint was produced

item value
training repo QATFactory branch weili/w4a4 @ a9315e4 (PR) — scripts/train_llm_qat.py
method QAD (quantization-aware distillation): student trains with fake-quantized forward; frozen BF16 teacher = the base model itself
objective pure KL at temperature 1.0 (distill_weight 1.0, hard_label_weight 0.0), loss on assistant tokens only
base model / teacher Qwen/Qwen3.5-9B @ c202236235762e1c871ad0ccb60c8ee5ba337b9a
dataset openperfectblend_100k_Qwen3.5-9B_think — ~100K ChatML conversations (OpenPerfectBlend-derived prompts with Qwen3.5-9B think-mode responses; native <think> spans in the assistant turns), prepared in the QATFactory project. Train file from togethercomputer/Qwen3.5-9B-reasonmix @ b88c109 (940,793,581 bytes, md5 0406bb3a7a482352360716a1bc5e9e04; ~84.3k train conversations by the trainer's epoch accounting). Held-out eval = a disjoint 256-conversation split (md5 af10c8c304c146a81ddc35e439d7ac4b), 818,944 scored positions — the same corpus used for the serving-KL rows below
preprocessing model chat template (ChatML), assistant-only loss mask, max_seq_len 8192, right-truncated
this checkpoint optimizer step 3000 of 4000 (24,000 conversations consumed ~= 0.28 epoch, no data repetition)
batch per weight update 8 sequences = 1/GPU x 8 GPUs x grad-accum 1 (<= 8 x 8192 = 65,536 tokens/update)
learning rate peak 1.0e-5, cosine decay to 0 over 4000 steps, linear warmup 1% (40 steps) — picked by a 5-point 500-step sweep (lr in {1e-6, 2e-6, 4e-6, 8e-6, 1e-5}, W&B Qwen3.5-9B-NVFP4-QAD-Sweep500-LR*): held-out eval drop from init 0.02901 at s500 was -0.3% / -6.6% / -29.6% / -43.7% / -46.0% (winner, stable)
optimizer AdamW (adamw_torch, beta1 0.9 / beta2 0.999), weight_decay 0.0, max_grad_norm 1.0
precision / parallelism bf16, FSDP2 full_shard on jbom 8xB200 (single node), gradient checkpointing (non-reentrant), sdpa attention
fake-quantized modules all linear projections (q/k/v/o_proj, gate/up/down_proj, GatedDeltaNet in/out projections); embeddings, lm_head, norms and the vision tower stay BF16
quantization config quant_format: nvfp4, fused_runtime_scales: true, quantize_activations: false
seed / bookkeeping seed 42; held-out eval every 100 steps; checkpoint every 1000 steps; step time 2.2 s/step (2h28m total)
in-loop held-out eval KL 0.02901 (true init) -> 0.01199 (s4000; converged from ~s2400)
W&B 1tpzgfkg (project qatfactory-qat, public)

Training mode: W4A16 (weight-only fake-quant). Weights are fake-quantized in the forward pass; activations stay BF16 during training. The exported artifact can still be served W4A4 — the tables below measure both.

Serving

compressed-tensors NVFP4 artifact carrying the full W4A4 schema (FP4 E2M1 weights, block-16 FP8-E4M3 scales + per-tensor global scales, static input-activation scales learned during training). Default load in vLLM >= 0.25.1 on SM100+ serves W4A4 (measured 12.3-15.2k tok/s single-GPU greedy vs ~7.7-8.3k for W4A16 — ~1.6x). For clean weight-only W4A16 serving use the sibling --weight-only export: Qwen3.5-9B-NVFP4-QAD-LR1e-5-s4000-weight-only.

vllm serve weili-0234/Qwen3.5-9B-NVFP4-QAD-LR1e-5-s3000 --max-model-len 24576

Evaluation (step 3000)

This mid-run checkpoint was uploaded for archival completeness; no serving-eval rows were measured at step 3000. Measured rows exist at s4000. In-loop held-out eval trajectory: 0.02901 (true init) -> 0.01199 (s4000; converged from ~s2400).

Related checkpoints


Part of a monitored QAD experiment series with full bookkeeping (pre-registered predictions, exact SHAs/configs/seeds per run). Produced with AI assistance (Claude).