|
Download README.md from jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid: direct link, hf CLI and curl.
- Browser
- Download file 6.34 kB
-
https://huggingface.co/jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid/resolve/main/README.md
- Command line
-
hf download hf://jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid/README.md
-
curl -L -o README.md https://huggingface.co/jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid/resolve/main/README.md
6.34 kB
| library_name: vllm | |
| tags: [arvq, nvfp4, experimental] | |
| # GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid | |
| Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation. | |
| **75/75 MoE layers replaced by the full-corpus sequential PV campaign.** | |
| Other layers retain their previously published weights; see pv_progress.json. | |
| Each layer's two tensor files and reports are replaced together in one commit. | |
| The production recipe uses **same-input targets**, **FP16 block scales**, and a **50× boundary loss weight**. It differs from v1; previous versions of this README accidentally retained the v1 training description. | |
| The cleaned v3.2 training stream contains **15,007,754 tokens**, context1024, effective batch262,144 and microbatch65,536. It removes empty think blocks and frames documents with canonical GLM delimiters. Reasoning validation and development audit captures contain16,384 tokens each; audit is WikiText/OOD and was used during development. The maximum one-pass budget is58updates, with adaptive early stopping: LR .048/.032,4× decay by update25 or earlier after three worsening validation checks, then patience10. Reassignment runs every10updates and validation every5. Published receipts show67/75 layers selecting update5; a full corpus is available, but most layers do **not** consume an entire pass. | |
| Boundary weighting is50 on the **position preceding** token154842 (`</think>`) or154820 (`<|endoftext|>`),1 elsewhere. It multiplies squared output residuals once and normalizes by summed weights; it is not2500×. Window-final positions have weight1. Validation metrics and discrete reassignment checks remain unweighted. Books stay on the FP4 grid; learned FP16 block scales serialize directly. The cold index rate is2bpw, or2.125bpw including block scales, plus codebook/global overhead. | |
| Student inputs are propagated through the selected upstream hybrid layers. The target is the original FP8-source routed experts evaluated on **those same student inputs**, with frozen hot contribution subtracted. This is not v1's independent original-reference trajectory objective. Two trajectories may still be carried for diagnostic comparison. Reference labels in some historical reports are generic/stale; the reports' `target: same_input` and fitting code determine the actual objective. | |
| Hot experts, backbone and vision components are inherited; the original bootstrap hot-allocation mismatch was corrected as noted below. Native runtime correctness is a separate qualification from software decoding. | |
| ## Full-model evaluation (2026-09-22) | |
| Teacher-forced perplexity and per-token KL divergence against the NVFP4 donor | |
| (all 75 cold layers decoded and substituted, context 2048, 8-way sharded over | |
| four held-out corpora): | |
| | Corpus | Donor ppl | This repo ppl | KLD (nats) | Top-1 agreement | | |
| |---|---|---|---|---| | |
| | calib-domain heldout | 2.9005 | 2.9954 | 0.114 | 89.7% | | |
| | calib-domain heldout-xl | 2.6050 | 2.8028 | 0.150 | 88.7% | | |
| | wikitext | 3.0160 | 5.3726 | 0.735 | 70.4% | | |
| | github code | 2.7471 | 2.7950 | 0.097 | 90.8% | | |
| In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1 | |
| ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext | |
| shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a | |
| possible calibration-domain mismatch. These results do not isolate corpus bias from quantization format, optimization, or trajectory effects; treat general-English quality accordingly. | |
| **Repo history note (2026-09-22):** the hot tier, `aqlm_layer_books` and | |
| `cold_assignment.json` were re-exported to match the cold_manifests allocation. | |
| Snapshots before commit f53a8dcc carry a stale hot tier from the original | |
| bootstrap allocation and will fail the loader's hyb_kind consistency checks. | |
| ## Terminal-Bench (in progress) | |
| No full Terminal-Bench run has finished on this checkpoint yet. | |
| - **Terminal-Bench 2.1, partial:** 9 of the 13 tasks the [v1 ARVQ repo](https://huggingface.co/jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-hybrid) failed were rerun. 3 now pass (caffe-cifar-10, extract-elf, dna-insert), 3 timed out (dna-assembly, filter-js-from-html, make-doom-for-mips) and 3 were interrupted before finishing. | |
| - **Terminal-Bench 4.0, running:** concurrency 1 with an 8h agent timeout. 6 of 64 tasks are done: 1 pass (layout-config-recreation2), 4 fail and 1 timeout, so 1/5 excluding timeouts so far. photonic-waveguide-routing hit the timeout but is counted as a fail: it was stuck repeating max-length responses. | |
| This section will be updated when the runs finish. | |
| ## Serving (vLLM) | |
| Serve with the SM120 fork `github.com/jarrelscy/vllm-glm52-sm120`, branch | |
| `serving/arvq-v4-v5`, commit `d7ada6d` or later (the same commit is the tip | |
| of `experiment/arvq-mcbook16`; one serving line loads v3, v4 and v5 | |
| checkpoints). That build loads the v4 cold-expert format | |
| `rvq256_256x8_expert_fp16block` used by all PV-replaced layers here (fitted | |
| FP16 block scales serialized directly, applied in FP32 after the FP4 MMA; | |
| experimental ABI — rebuild the ARVQ libraries together with the loader via | |
| arvq/build.sh). The config.json `arvq` marker declares | |
| format version 4 to match. Earlier branches cannot load this repo: | |
| `experiment/arvq-fp16-block-scales` (00f57ca51) registers uint8 E4M3 scales | |
| and predates the v4 format string; the `glm52-sm120` main branch and the v3 | |
| `arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4` | |
| is for the v2r repo (quarter-scale residual decode), not this one. | |
| ## Reproduce this release | |
| The [reproducibility guide](reproduce/README.md), [stage-by-stage runbook](reproduce/RUNBOOK.md), [recipe](reproduce/recipe.json), [data artifact hashes](reproduce/data_artifacts.json), and [source tree](reproduce/source) now accompany this checkpoint. The package contains202 recovered source files spanning corpus preparation, capture, fitting, PV, export and evaluation. Run `python reproduce/verify.py` for CPU checks. | |
| **Scope:** the2026-09-26 audit adds code/documentation and corrects metadata; it does not change model weights. Exact historical retraining is not yet guaranteed: some original local corpus inputs, streamed-dataset revisions and environment locks were not preserved. Recovered artifacts, missing inputs and changed-source qualifications are explicitly documented rather than concealed. | |