GLM-5.3-Vision-NVFP4-ARVQ-hybrid
Per-expert ARVQ v3 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation. 75/75 MoE layers replaced by the full-corpus sequential PV campaign. All75 routed MoE layers are complete; dense/nonexpert components remain inherited. Each layer's two tensor files and reports are replaced together in one commit.
Training draws sequentially from 18,001,846 text tokens at context 1024. Fixed validation and development-audit sets each contain 16,384 tokens. Adam trains FP4-constrained per-expert books and FP8-constrained per-block scales. The effective batch is 262,144 tokens, accumulated in four 65,536-token passes. Layers 4–26 use 69 updates: book/scale LR .048/.032 through update 45, then .012/.008. From layer27, 69 updates is the maximum: three validation checks more than 0.1% worse than best trigger an earlier LR reduction; 15 updates without improvement after the reduction permit stopping. Layer30 retains audit-qualified update15 after its validation-best update20 failed the development audit. Layer 3 retains its separately qualified lower-LR refinement. Output-gradient index reassignment occurs every 20 updates; validation every five updates retains the best checkpoint. Audit non-regression and export replay gate publication.
Student inputs reflect retained tuned upstream layers. Fixed reference targets use original FP8-source routed experts and the donor backbone. A rolling cache carries both trajectories. Cold arithmetic emulates FP4 activation planes and FP16 boundaries; native SM120 parity and full-model quality remain unverified. The audit set is historical development data, not an untouched final test. Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
Terminal-Bench 2.1
76/89 = 85.4% (89 tasks: 76 pass, 13 fail, no timeouts).
- Harbor 0.21.0, terminus-2 agent, temperature 1.0, top_p 0.95, thinking off, served by vLLM at TP4 on 4× RTX PRO 6000.
- Fails: caffe-cifar-10, db-wal-recovery, dna-assembly, dna-insert, extract-elf, filter-js-from-html, gcode-to-text, install-windows-3.11, make-doom-for-mips, protein-assembly, qemu-startup, raman-fitting, video-processing.
- For comparison, the AQLM-hybrid-1m checkpoint scores 90.7% on the same suite.
Agent trajectories, verifier output and result.json for every trial are in benchmarks/tb2.1/.
Boundary weighting: this v1 recipe uses uniform token loss; the50× next-boundary rule belongs to ARVQ-v2. Original-reference targets and FP8-E4M3 block scales distinguish v1 from v2.
Reproduce this release
The reproducibility guide, stage-by-stage runbook, recipe, data artifact hashes, and source tree now accompany this checkpoint. The package contains202 recovered source files spanning corpus preparation, capture, fitting, PV, export and evaluation. Run python reproduce/verify.py for CPU checks.
Scope: the2026-09-26 audit adds code/documentation and corrects metadata; it does not change model weights. Exact historical retraining is not yet guaranteed: some original local corpus inputs, streamed-dataset revisions and environment locks were not preserved. Recovered artifacts, missing inputs and changed-source qualifications are explicitly documented rather than concealed.
- Downloads last month
- 849