jarrelscy's picture
Publish quantization/data reproduction source and correct recipe documentation
cb95231 verified
|
Raw History Blame Contribute Delete
6.34 kB
---
library_name: vllm
tags: [arvq, nvfp4, experimental]
---
# GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid
Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation.
**75/75 MoE layers replaced by the full-corpus sequential PV campaign.**
Other layers retain their previously published weights; see pv_progress.json.
Each layer's two tensor files and reports are replaced together in one commit.
The production recipe uses **same-input targets**, **FP16 block scales**, and a **50× boundary loss weight**. It differs from v1; previous versions of this README accidentally retained the v1 training description.
The cleaned v3.2 training stream contains **15,007,754 tokens**, context1024, effective batch262,144 and microbatch65,536. It removes empty think blocks and frames documents with canonical GLM delimiters. Reasoning validation and development audit captures contain16,384 tokens each; audit is WikiText/OOD and was used during development. The maximum one-pass budget is58updates, with adaptive early stopping: LR .048/.032,4× decay by update25 or earlier after three worsening validation checks, then patience10. Reassignment runs every10updates and validation every5. Published receipts show67/75 layers selecting update5; a full corpus is available, but most layers do **not** consume an entire pass.
Boundary weighting is50 on the **position preceding** token154842 (`</think>`) or154820 (`<|endoftext|>`),1 elsewhere. It multiplies squared output residuals once and normalizes by summed weights; it is not2500×. Window-final positions have weight1. Validation metrics and discrete reassignment checks remain unweighted. Books stay on the FP4 grid; learned FP16 block scales serialize directly. The cold index rate is2bpw, or2.125bpw including block scales, plus codebook/global overhead.
Student inputs are propagated through the selected upstream hybrid layers. The target is the original FP8-source routed experts evaluated on **those same student inputs**, with frozen hot contribution subtracted. This is not v1's independent original-reference trajectory objective. Two trajectories may still be carried for diagnostic comparison. Reference labels in some historical reports are generic/stale; the reports' `target: same_input` and fitting code determine the actual objective.
Hot experts, backbone and vision components are inherited; the original bootstrap hot-allocation mismatch was corrected as noted below. Native runtime correctness is a separate qualification from software decoding.
## Full-model evaluation (2026-09-22)
Teacher-forced perplexity and per-token KL divergence against the NVFP4 donor
(all 75 cold layers decoded and substituted, context 2048, 8-way sharded over
four held-out corpora):
| Corpus | Donor ppl | This repo ppl | KLD (nats) | Top-1 agreement |
|---|---|---|---|---|
| calib-domain heldout | 2.9005 | 2.9954 | 0.114 | 89.7% |
| calib-domain heldout-xl | 2.6050 | 2.8028 | 0.150 | 88.7% |
| wikitext | 3.0160 | 5.3726 | 0.735 | 70.4% |
| github code | 2.7471 | 2.7950 | 0.097 | 90.8% |
In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
possible calibration-domain mismatch. These results do not isolate corpus bias from quantization format, optimization, or trajectory effects; treat general-English quality accordingly.
**Repo history note (2026-09-22):** the hot tier, `aqlm_layer_books` and
`cold_assignment.json` were re-exported to match the cold_manifests allocation.
Snapshots before commit f53a8dcc carry a stale hot tier from the original
bootstrap allocation and will fail the loader's hyb_kind consistency checks.
## Terminal-Bench (in progress)
No full Terminal-Bench run has finished on this checkpoint yet.
- **Terminal-Bench 2.1, partial:** 9 of the 13 tasks the [v1 ARVQ repo](https://huggingface.co/jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-hybrid) failed were rerun. 3 now pass (caffe-cifar-10, extract-elf, dna-insert), 3 timed out (dna-assembly, filter-js-from-html, make-doom-for-mips) and 3 were interrupted before finishing.
- **Terminal-Bench 4.0, running:** concurrency 1 with an 8h agent timeout. 6 of 64 tasks are done: 1 pass (layout-config-recreation2), 4 fail and 1 timeout, so 1/5 excluding timeouts so far. photonic-waveguide-routing hit the timeout but is counted as a fail: it was stuck repeating max-length responses.
This section will be updated when the runs finish.
## Serving (vLLM)
Serve with the SM120 fork `github.com/jarrelscy/vllm-glm52-sm120`, branch
`serving/arvq-v4-v5`, commit `d7ada6d` or later (the same commit is the tip
of `experiment/arvq-mcbook16`; one serving line loads v3, v4 and v5
checkpoints). That build loads the v4 cold-expert format
`rvq256_256x8_expert_fp16block` used by all PV-replaced layers here (fitted
FP16 block scales serialized directly, applied in FP32 after the FP4 MMA;
experimental ABI — rebuild the ARVQ libraries together with the loader via
arvq/build.sh). The config.json `arvq` marker declares
format version 4 to match. Earlier branches cannot load this repo:
`experiment/arvq-fp16-block-scales` (00f57ca51) registers uint8 E4M3 scales
and predates the v4 format string; the `glm52-sm120` main branch and the v3
`arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
is for the v2r repo (quarter-scale residual decode), not this one.
## Reproduce this release
The [reproducibility guide](reproduce/README.md), [stage-by-stage runbook](reproduce/RUNBOOK.md), [recipe](reproduce/recipe.json), [data artifact hashes](reproduce/data_artifacts.json), and [source tree](reproduce/source) now accompany this checkpoint. The package contains202 recovered source files spanning corpus preparation, capture, fitting, PV, export and evaluation. Run `python reproduce/verify.py` for CPU checks.
**Scope:** the2026-09-26 audit adds code/documentation and corrects metadata; it does not change model weights. Exact historical retraining is not yet guaranteed: some original local corpus inputs, streamed-dataset revisions and environment locks were not preserved. Recovered artifacts, missing inputs and changed-source qualifications are explicitly documented rather than concealed.