# Changelog ## 2026-08-29 — initial release (private) Full W4AFP8 quantization of GLM-5.3 from the BF16 master. AWQ calibration on SALT-NLP/SWE-chat coding-agent traces (128 x 2048-token windows, session-disjoint; 24 sessions held out for evaluation). MoE experts INT4 group-128 with AWQ scaling and clipping; non-expert layers FP8 blockwise; DSA indexer weights FP8; MTP draft layer (layer 78) included as weight-only INT4 (an AWQ-calibrated draft variant is staged in aux/ pending acceptance-rate A/B). Structural and dequantization audits passed against the GLM-5.2-W4AFP8 format reference. Held-out teacher-forced NLL and task benchmark results will be appended here as they complete. ## 2026-08-30 — Full evaluation suite + draft calibration study Benchmarks (sglang 0.5.16, tp8, w4afp8, fp8 KV; temp 1.0 top_p 0.95, official request config): - GPQA-Diamond: **91.92%** (182/198) vs AA reference 91.7 (capped-retry protocol: 3 truncated items re-run at 126,976 tokens; per-item diff verified no completed item was re-rolled) - BFCL (45-item suite): 82.2% overall; real function-calling 28/30 — item-identical to GLM-5.2-W4AFP8 - AA-LCR: 73.0 (n=100, uncapped) vs AA reference 76.3 (0.75 sigma at n=100); on GLM-5.2's exact 96-item subset with the same judge (prod GLM-5.2): 74.0 vs 5.2's 76.0 (2-item gap, sign-test p~0.8) - NIAH: 3/3 needles retrieved at ~930k-token prompts (depths 0.1/0.5/0.9) - Teacher-forced dNLL (32 held-out SWE-chat windows): BF16 offline 2.3011 -> served 2.5832 (+0.282 nats, includes fp8-act + fp8-KV serving taxes) MTP draft (layer 78) calibration study — **RTN retained**: - Three AWQ variants (BF16 chain-forward calibration, runtime-captured-activation calibration, fp8-serving-numerics-aware search) ALL regress acceptance ~20-25% vs RTN despite beating RTN on every offline single-step metric (layer relMSE, argmax agreement, t+2 top-1). - Root cause (drafting-depth A/B): at speculative-num-steps=1 AWQ is -3.2% vs RTN; at steps=3 it is -21.4%. EAGLE-style drafting feeds the draft its own hidden output for steps 2-3, and AWQ+clip's biased shrinkage error compounds through that recursion; RTN's unbiased rounding error does not. Draft-layer calibration must include decode-time recursive inputs to beat RTN. - Losing shards and all acceptance runs preserved under aux/ (mtp_layer78_awq{,_rt,_fp8sim}.safetensors, aux/results/accept_*.json). ## 2026-08-30 (later) — Draft calibration study RESOLVED: shard-format defect, not AWQ The ~20-25% acceptance regressions reported above for all AWQ draft variants were caused by a shard-emission defect, not by AWQ calibration: the spliced study shards left the DSA indexer weights (wk/wq_b) as BF16 without weight_scale_inv where the serving convention requires FP8+scale, and downcast gate.e_score_correction_bias F32->BF16. A corrected shard in the exact shipped convention (2337 tensors, verified key/dtype/shape-identical to the RTN reference) ties RTN on held-out acceptance: 2.9309 vs 2.9319 mean accept length (same box, same 24 held-out sessions, EAGLE steps=3). The shipped checkpoint has carried the RTN draft throughout and is unaffected; it remains RTN (tie = no reason to change a validated artifact).