GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid
Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation. 75/75 MoE layers replaced by the full-corpus sequential PV campaign. Other layers retain their previously published weights; see pv_progress.json. Each layer's two tensor files and reports are replaced together in one commit.
The production recipe uses same-input targets, FP16 block scales, and a 50× boundary loss weight. It differs from v1; previous versions of this README accidentally retained the v1 training description.
The cleaned v3.2 training stream contains 15,007,754 tokens, context1024, effective batch262,144 and microbatch65,536. It removes empty think blocks and frames documents with canonical GLM delimiters. Reasoning validation and development audit captures contain16,384 tokens each; audit is WikiText/OOD and was used during development. The maximum one-pass budget is58updates, with adaptive early stopping: LR .048/.032,4× decay by update25 or earlier after three worsening validation checks, then patience10. Reassignment runs every10updates and validation every5. Published receipts show67/75 layers selecting update5; a full corpus is available, but most layers do not consume an entire pass.
Boundary weighting is50 on the position preceding token154842 (</think>) or154820 (<|endoftext|>),1 elsewhere. It multiplies squared output residuals once and normalizes by summed weights; it is not2500×. Window-final positions have weight1. Validation metrics and discrete reassignment checks remain unweighted. Books stay on the FP4 grid; learned FP16 block scales serialize directly. The cold index rate is2bpw, or2.125bpw including block scales, plus codebook/global overhead.
Student inputs are propagated through the selected upstream hybrid layers. The target is the original FP8-source routed experts evaluated on those same student inputs, with frozen hot contribution subtracted. This is not v1's independent original-reference trajectory objective. Two trajectories may still be carried for diagnostic comparison. Reference labels in some historical reports are generic/stale; the reports' target: same_input and fitting code determine the actual objective.
Hot experts, backbone and vision components are inherited; the original bootstrap hot-allocation mismatch was corrected as noted below. Native runtime correctness is a separate qualification from software decoding.
Full-model evaluation (2026-09-22)
Teacher-forced perplexity and per-token KL divergence against the NVFP4 donor (all 75 cold layers decoded and substituted, context 2048, 8-way sharded over four held-out corpora):
| Corpus | Donor ppl | This repo ppl | KLD (nats) | Top-1 agreement |
|---|---|---|---|---|
| calib-domain heldout | 2.9005 | 2.9954 | 0.114 | 89.7% |
| calib-domain heldout-xl | 2.6050 | 2.8028 | 0.150 | 88.7% |
| wikitext | 3.0160 | 5.3726 | 0.735 | 70.4% |
| github code | 2.7471 | 2.7950 | 0.097 | 90.8% |
In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1 ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a possible calibration-domain mismatch. These results do not isolate corpus bias from quantization format, optimization, or trajectory effects; treat general-English quality accordingly.
Repo history note (2026-09-22): the hot tier, aqlm_layer_books and
cold_assignment.json were re-exported to match the cold_manifests allocation.
Snapshots before commit f53a8dcc carry a stale hot tier from the original
bootstrap allocation and will fail the loader's hyb_kind consistency checks.
Terminal-Bench (in progress)
No full Terminal-Bench run has finished on this checkpoint yet.
- Terminal-Bench 2.1, partial: 9 of the 13 tasks the v1 ARVQ repo failed were rerun. 3 now pass (caffe-cifar-10, extract-elf, dna-insert), 3 timed out (dna-assembly, filter-js-from-html, make-doom-for-mips) and 3 were interrupted before finishing.
- Terminal-Bench 4.0, running: concurrency 1 with an 8h agent timeout. 6 of 64 tasks are done: 1 pass (layout-config-recreation2), 4 fail and 1 timeout, so 1/5 excluding timeouts so far. photonic-waveguide-routing hit the timeout but is counted as a fail: it was stuck repeating max-length responses.
This section will be updated when the runs finish.
Serving (vLLM)
Serve with the SM120 fork github.com/jarrelscy/vllm-glm52-sm120, branch
serving/arvq-v4-v5, commit d7ada6d or later (the same commit is the tip
of experiment/arvq-mcbook16; one serving line loads v3, v4 and v5
checkpoints). That build loads the v4 cold-expert format
rvq256_256x8_expert_fp16block used by all PV-replaced layers here (fitted
FP16 block scales serialized directly, applied in FP32 after the FP4 MMA;
experimental ABI — rebuild the ARVQ libraries together with the loader via
arvq/build.sh). The config.json arvq marker declares
format version 4 to match. Earlier branches cannot load this repo:
experiment/arvq-fp16-block-scales (00f57ca51) registers uint8 E4M3 scales
and predates the v4 format string; the glm52-sm120 main branch and the v3
arvq-hybrid-sm120 branch are v3-only; experiment/arvq-fp16-scales-rs4
is for the v2r repo (quarter-scale residual decode), not this one.
Reproduce this release
The reproducibility guide, stage-by-stage runbook, recipe, data artifact hashes, and source tree now accompany this checkpoint. The package contains202 recovered source files spanning corpus preparation, capture, fitting, PV, export and evaluation. Run python reproduce/verify.py for CPU checks.
Scope: the2026-09-26 audit adds code/documentation and corrects metadata; it does not change model weights. Exact historical retraining is not yet guaranteed: some original local corpus inputs, streamed-dataset revisions and environment locks were not preserved. Recovered artifacts, missing inputs and changed-source qualifications are explicitly documented rather than concealed.
- Downloads last month
- 623