README: full-model eval results; correct serving pin to v4-capable branch
Browse files
README.md
CHANGED
|
@@ -2,9 +2,9 @@
|
|
| 2 |
library_name: vllm
|
| 3 |
tags: [arvq, nvfp4, experimental]
|
| 4 |
---
|
| 5 |
-
# GLM-5.3-Vision-NVFP4-ARVQ-hybrid
|
| 6 |
|
| 7 |
-
Per-expert ARVQ
|
| 8 |
**75/75 MoE layers replaced by the full-corpus sequential PV campaign.**
|
| 9 |
Other layers retain their previously published weights; see pv_progress.json.
|
| 10 |
Each layer's two tensor files and reports are replaced together in one commit.
|
|
@@ -25,17 +25,40 @@ the best checkpoint. Audit non-regression and export replay gate publication.
|
|
| 25 |
Student inputs reflect retained tuned upstream layers. Fixed reference targets
|
| 26 |
use original FP8-source routed experts and the donor backbone. A rolling cache
|
| 27 |
carries both trajectories. Cold arithmetic emulates FP4 activation planes and
|
| 28 |
-
FP16 boundaries; native SM120
|
|
|
|
| 29 |
The audit set is historical development data, not an untouched final test.
|
| 30 |
Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
|
| 31 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
## Serving (vLLM)
|
| 33 |
|
| 34 |
Serve with the SM120 fork `github.com/jarrelscy/vllm-glm52-sm120`, branch
|
| 35 |
-
`experiment/arvq-
|
| 36 |
-
v4 cold-expert format `rvq256_256x8_expert_fp16block` used by
|
| 37 |
-
layers here (FP16 block scales
|
| 38 |
-
after the FP4 MMA; experimental ABI — rebuild
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
library_name: vllm
|
| 3 |
tags: [arvq, nvfp4, experimental]
|
| 4 |
---
|
| 5 |
+
# GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid
|
| 6 |
|
| 7 |
+
Per-expert ARVQ v4 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation.
|
| 8 |
**75/75 MoE layers replaced by the full-corpus sequential PV campaign.**
|
| 9 |
Other layers retain their previously published weights; see pv_progress.json.
|
| 10 |
Each layer's two tensor files and reports are replaced together in one commit.
|
|
|
|
| 25 |
Student inputs reflect retained tuned upstream layers. Fixed reference targets
|
| 26 |
use original FP8-source routed experts and the donor backbone. A rolling cache
|
| 27 |
carries both trajectories. Cold arithmetic emulates FP4 activation planes and
|
| 28 |
+
FP16 boundaries; native SM120 kernel execution has not yet been smoke-tested
|
| 29 |
+
(the evaluation below decodes the published weights to BF16 in software).
|
| 30 |
The audit set is historical development data, not an untouched final test.
|
| 31 |
Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged.
|
| 32 |
|
| 33 |
+
## Full-model evaluation (2026-09-22)
|
| 34 |
+
|
| 35 |
+
Teacher-forced perplexity and per-token KL divergence against the NVFP4 donor
|
| 36 |
+
(all 75 cold layers decoded and substituted, context 2048, 8-way sharded over
|
| 37 |
+
four held-out corpora):
|
| 38 |
+
|
| 39 |
+
| Corpus | Donor ppl | This repo ppl | KLD (nats) | Top-1 agreement |
|
| 40 |
+
|---|---|---|---|---|
|
| 41 |
+
| calib-domain heldout | 2.9005 | 2.9954 | 0.114 | 89.7% |
|
| 42 |
+
| calib-domain heldout-xl | 2.6050 | 2.8028 | 0.150 | 88.7% |
|
| 43 |
+
| wikitext | 3.0160 | 5.3726 | 0.735 | 70.4% |
|
| 44 |
+
| github code | 2.7471 | 2.7950 | 0.097 | 90.8% |
|
| 45 |
+
|
| 46 |
+
In-domain degradation is small (KLD 0.10–0.15 nats; 35–40% lower than the v1
|
| 47 |
+
ARVQ repo, whose corresponding KLDs are 0.191/0.237/0.733/0.154). Wikitext
|
| 48 |
+
shows a large gap (+78% ppl vs donor) on both this repo and v1, indicating a
|
| 49 |
+
calibration-domain bias of the ARVQ calibration mix rather than a fitting
|
| 50 |
+
regression; treat general-English quality accordingly.
|
| 51 |
+
|
| 52 |
## Serving (vLLM)
|
| 53 |
|
| 54 |
Serve with the SM120 fork `github.com/jarrelscy/vllm-glm52-sm120`, branch
|
| 55 |
+
`experiment/arvq-mcbook16`, commit `d7ada6d` or later. That build loads the
|
| 56 |
+
v4 cold-expert format `rvq256_256x8_expert_fp16block` used by all PV-replaced
|
| 57 |
+
layers here (fitted FP16 block scales serialized directly, applied in FP32
|
| 58 |
+
after the FP4 MMA; experimental ABI — rebuild the ARVQ libraries together
|
| 59 |
+
with the loader via arvq/build.sh). The config.json `arvq` marker declares
|
| 60 |
+
format version 4 to match. Earlier branches cannot load this repo:
|
| 61 |
+
`experiment/arvq-fp16-block-scales` (00f57ca51) registers uint8 E4M3 scales
|
| 62 |
+
and predates the v4 format string; the `glm52-sm120` main branch and the v3
|
| 63 |
+
`arvq-hybrid-sm120` branch are v3-only; `experiment/arvq-fp16-scales-rs4`
|
| 64 |
+
is for the v2r repo (quarter-scale residual decode), not this one.
|