|
Download README.md from jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid: direct link, hf CLI and curl.
- Browser
- Download file 1.82 kB
-
https://huggingface.co/jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid/resolve/fd250ffec866a7b55fdbd059cf9f8255ba345dcf/README.md
- Command line
-
hf download hf://jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid@fd250ffec866a7b55fdbd059cf9f8255ba345dcf/README.md
-
curl -L -o README.md https://huggingface.co/jarrelscy/GLM-5.3-Vision-NVFP4-ARVQ-v2-hybrid/resolve/fd250ffec866a7b55fdbd059cf9f8255ba345dcf/README.md
1.82 kB
| library_name: vllm | |
| tags: [arvq, nvfp4, experimental] | |
| # GLM-5.3-Vision-NVFP4-ARVQ-hybrid | |
| Per-expert ARVQ v3 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation. | |
| **75/75 MoE layers replaced by the full-corpus sequential PV campaign.** | |
| Other layers retain their previously published weights; see pv_progress.json. | |
| Each layer's two tensor files and reports are replaced together in one commit. | |
| Training draws sequentially from 18,001,846 text tokens at context 1024. Fixed validation and | |
| development-audit sets each contain 16,384 tokens. Adam trains FP4-constrained | |
| per-expert books and FP8-constrained per-block scales. The effective batch is | |
| 262,144 tokens, accumulated in four 65,536-token passes. Layers 4–26 use 69 | |
| updates: book/scale LR .048/.032 through update 45, then .012/.008. From layer27, | |
| 69 updates is the maximum: three validation checks more than 0.1% worse than | |
| best trigger an earlier LR reduction; 15 updates without improvement after | |
| the reduction permit stopping. Layer30 retains audit-qualified update15 after | |
| its validation-best update20 failed the development audit. Layer 3 | |
| retains its separately qualified lower-LR refinement. Output-gradient index | |
| reassignment occurs every 20 updates; validation every five updates retains | |
| the best checkpoint. Audit non-regression and export replay gate publication. | |
| Student inputs reflect retained tuned upstream layers. Fixed reference targets | |
| use original FP8-source routed experts and the donor backbone. A rolling cache | |
| carries both trajectories. Cold arithmetic emulates FP4 activation planes and | |
| FP16 boundaries; native SM120 parity and full-model quality remain unverified. | |
| The audit set is historical development data, not an untouched final test. | |
| Hot experts, backbone, BF16 MTP, vision components and allocation are unchanged. | |