jarrelscy commited on
Commit
7642c72
·
verified ·
1 Parent(s): dd6ed65

README: full-model eval results; correct serving pin to v4-capable branch

Browse files
Files changed (1) hide show
  1. README.md +33 -10
README.md CHANGED
@@ -2,9 +2,9 @@
2
  library_name: vllm
3
  tags: [arvq, nvfp4, experimental]
4
  ---
5
- # GLM-5.3-Vision-NVFP4-ARVQ-hybrid
6
 
7
- Per-expert ARVQ v3 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation.
8
  **75/75 MoE layers replaced by the full-corpus sequential PV campaign.**
9
  Other layers retain their previously published weights; see pv_progress.json.
10
  Each layer's two tensor files and reports are replaced together in one commit.
@@ -25,17 +25,40 @@ the best checkpoint. Audit non-regression and export replay gate publication.
25
  Student inputs reflect retained tuned upstream layers. Fixed reference targets
26
  use original FP8-source routed experts and the donor backbone. A rolling cache
27
  carries both trajectories. Cold arithmetic emulates FP4 activation planes and
28
- FP16 boundaries; native SM120 parity and full-model quality remain unverified.
 
29
  The audit set is historical development data, not an untouched final test.
30
  Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
31
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
  ## Serving (vLLM)
33
 
34
  Serve with the SM120 fork `github.com/jarrelscy/vllm-glm52-sm120`, branch
35
- `experiment/arvq-fp16-block-scales`, commit `00f57ca51`. That build loads the
36
- v4 cold-expert format `rvq256_256x8_expert_fp16block` used by the PV-replaced
37
- layers here (FP16 block scales converted at load time and applied in FP32
38
- after the FP4 MMA; experimental ABI — rebuild loader and libraries together).
39
- The `glm52-sm120` main branch and the v3 `arvq-hybrid-sm120` branch cannot
40
- load FP16 block-scale layers, and `experiment/arvq-fp16-scales-rs4` is for
41
- the v2r repo (quarter-scale residual decode), not this one.
 
 
 
 
2
  library_name: vllm
3
  tags: [arvq, nvfp4, experimental]
4
  ---
5
+ # GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid
6
 
7
+ Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation.
8
  **75/75 MoE layers replaced by the full-corpus sequential PV campaign.**
9
  Other layers retain their previously published weights; see pv_progress.json.
10
  Each layer's two tensor files and reports are replaced together in one commit.
 
25
  Student inputs reflect retained tuned upstream layers. Fixed reference targets
26
  use original FP8-source routed experts and the donor backbone. A rolling cache
27
  carries both trajectories. Cold arithmetic emulates FP4 activation planes and
28
+ FP16 boundaries; native SM120 kernel execution has not yet been smoke-tested
29
+ (the evaluation below decodes the published weights to BF16 in software).
30
  The audit set is historical development data, not an untouched final test.
31
  Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
32
 
33
+ ## Full-model evaluation (2026-09-22)
34
+
35
+ Teacher-forced perplexity and per-token KL divergence against the NVFP4 donor
36
+ (all 75 cold layers decoded and substituted, context 2048, 8-way sharded over
37
+ four held-out corpora):
38
+
39
+ | Corpus | Donor ppl | This repo ppl | KLD (nats) | Top-1 agreement |
40
+ |---|---|---|---|---|
41
+ | calib-domain heldout | 2.9005 | 2.9954 | 0.114 | 89.7% |
42
+ | calib-domain heldout-xl | 2.6050 | 2.8028 | 0.150 | 88.7% |
43
+ | wikitext | 3.0160 | 5.3726 | 0.735 | 70.4% |
44
+ | github code | 2.7471 | 2.7950 | 0.097 | 90.8% |
45
+
46
+ In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
47
+ ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
48
+ shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
49
+ calibration-domain bias of the ARVQ calibration mix rather than a fitting
50
+ regression; treat general-English quality accordingly.
51
+
52
  ## Serving (vLLM)
53
 
54
  Serve with the SM120 fork `github.com/jarrelscy/vllm-glm52-sm120`, branch
55
+ `experiment/arvq-mcbook16`, commit `d7ada6d` or later. That build loads the
56
+ v4 cold-expert format `rvq256_256x8_expert_fp16block` used by all PV-replaced
57
+ layers here (fitted FP16 block scales serialized directly, applied in FP32
58
+ after the FP4 MMA; experimental ABI — rebuild the ARVQ libraries together
59
+ with the loader via arvq/build.sh). The config.json `arvq` marker declares
60
+ format version 4 to match. Earlier branches cannot load this repo:
61
+ `experiment/arvq-fp16-block-scales` (00f57ca51) registers uint8 E4M3 scales
62
+ and predates the v4 format string; the `glm52-sm120` main branch and the v3
63
+ `arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
64
+ is for the v2r repo (quarter-scale residual decode), not this one.