brandonmusic commited on
Commit
767b027
·
verified ·
1 Parent(s): 4739eb1

Document SM120 v34 KLD fix and MTP5 runtime

Browse files
Files changed (1) hide show
  1. README.md +73 -4
README.md CHANGED
@@ -1,6 +1,6 @@
1
  # GLM-5.3-Flash-EXL3-4bpw
2
 
3
- Source: `zai-org/GLM-5.3-Flash-BF16@a6c167b62691b2bac901344b65cb651a70f53e43`. All routed experts including MTP45 are uniform four-bit EXL3/TR3 MCG; non-routed tensors retain their official native dtype. The custom TP2 runtime is qualified by actual-runtime BF16-teacher KLD plus byte-identical rank outputs and multi-token generation; its stricter decoded raw-logit parity diagnostic remains failed.
4
 
5
  Five-cold-run mean teacher-to-student KLD: `0.024554564250` over 51,175 sealed causal positions per run. Actual TP2 runtime qualification-window KLD: `0.022750847878` over 2,047 positions (both gates: mean KLD < 0.06). This checkpoint requires the included custom Transformers TP2 adapter and is not a stock vLLM/ExLlamaV3 compatibility claim.
6
 
@@ -39,11 +39,80 @@ which records every payload path, size, and SHA-256. The 25 final windows remain
39
  qualification-only and are excluded from fitting and expert selection.
40
 
41
  ## Minimal TP2 launch
42
- Note this is NOT an optimize launch nearly at all. I woudl recommend trying https://github.com/chriswritescode-dev/glm-5.3-flash-sm120 image referenced
43
- here. that use the MLA which compresses the latent before you get to kv_b_proj. Getting that running you might get 1 million kv cache.
44
- the below run time is probably closer to 32k kv cache max.
 
 
45
  Use Transformers 5.16.1, clone ExLlamaV3 at commit `c5d9c657966ffeeaa9353f0cc899f18629da4a13`, compile its CUDA extension, then run:
46
 
47
  ```bash
48
  PYTHONPATH=runtime/src torchrun --standalone --nproc-per-node=2 runtime/scripts/run_glm53_custom_tp_runtime.py --model . --exllamav3-source /path/to/exllamav3 --prompt 'Hello'
49
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # GLM-5.3-Flash-EXL3-4bpw
2
 
3
+ Source: `zai-org/GLM-5.3-Flash-BF16@a6c167b62691b2bac901344b65cb651a70f53e43`. All routed experts including MTP45 are uniform four-bit EXL3/TR3 MCG; non-routed tensors retain their official native dtype. The original custom Transformers TP2 runtime and the dedicated SM120 vLLM image below are qualified separately against the same BF16 teacher evidence.
4
 
5
  Five-cold-run mean teacher-to-student KLD: `0.024554564250` over 51,175 sealed causal positions per run. Actual TP2 runtime qualification-window KLD: `0.022750847878` over 2,047 positions (both gates: mean KLD < 0.06). This checkpoint requires the included custom Transformers TP2 adapter and is not a stock vLLM/ExLlamaV3 compatibility claim.
6
 
 
39
  qualification-only and are excluded from fitting and expert selection.
40
 
41
  ## Minimal TP2 launch
42
+
43
+ The historical Transformers runtime below is an evidence/qualification path,
44
+ not the optimized daily-driver launch. Use the SM120 container in the next
45
+ section for serving.
46
+
47
  Use Transformers 5.16.1, clone ExLlamaV3 at commit `c5d9c657966ffeeaa9353f0cc899f18629da4a13`, compile its CUDA extension, then run:
48
 
49
  ```bash
50
  PYTHONPATH=runtime/src torchrun --standalone --nproc-per-node=2 runtime/scripts/run_glm53_custom_tp_runtime.py --model . --exllamav3-source /path/to/exllamav3 --prompt 'Hello'
51
  ```
52
+
53
+ ## SM120 TP2 daily-driver image
54
+
55
+ Docker Hub: [`verdictai/glm53-flash-exl3-k4`](https://hub.docker.com/r/verdictai/glm53-flash-exl3-k4)
56
+
57
+ - Version: `r19-sm120-tp2-v34`
58
+ - Immutable digest: `sha256:98df6f97cf5a40513a4bc8deda6a61e5c741430e7e53a71434e2bf22b4e0aa89`
59
+ - Hardware qualified: 2x RTX PRO 6000 Blackwell (SM120), TP2
60
+ - Daily-driver mode: NVFP4 MLA KV, DCP1, CUDA graphs, MTP5, 499,968-token maximum context
61
+ - Alternate modes: FP8 MLA KV and DCP2
62
+ - Sampling defaults from the model generation config: temperature `1.0`, top-p `0.95`
63
+
64
+ The image contains a dedicated vLLM/B12X overlay for GLM-5.3's hybrid linear/full-attention architecture, NoPE MLA, EXL3 K4 routed experts, and the MTP layer. It is not a claim that the checkpoint runs in upstream stock vLLM.
65
+
66
+ ### Corrected actual-runtime KLD
67
+
68
+ The exact 2,048-token `final-0000` qualification window was captured with TP2,
69
+ DCP1, eager execution, `fp8_ds_mla`, no MTP, and full-vocabulary float32
70
+ runtime logits, then compared to the sealed BF16 teacher in float64 chunks.
71
+
72
+ | Metric | Corrected SM120 v34 | Rented B200 custom TP2 | Offline K4 |
73
+ |---|---:|---:|---:|
74
+ | Mean teacher KLD | **0.024864241526** | 0.022750847878 | 0.031831601179 |
75
+ | Top-1 agreement | **0.940400586224** | 0.9384 | — |
76
+
77
+ The earlier local KLD near `0.10` was a runtime scale-decoding defect, not a
78
+ routing-quality result. The cache writer stores GLM's four calibrated
79
+ per-token, per-128-channel scales as arbitrary FP32 values (`amax / 448`). The
80
+ SM120 FlashInfer reader was left at `kv_scale_format="auto"`, which interprets
81
+ inline scales using the DeepSeek-v3.2 power-of-two convention. v34 explicitly
82
+ selects `arbitrary_fp32` in the GLM NoPE adapter. The corrected first-64-row KLD
83
+ is `0.1918669499`; rows 64 onward are `0.0194743407`, and the whole-window
84
+ result reproduces the independently observed server range.
85
+
86
+ The KLD report is published in this repo at
87
+ `runtime-results/v34/fp8-final-0000-kld-report.json`.
88
+
89
+ ### MTP5 measured decode
90
+
91
+ MTP5 is enabled by default with probabilistic rejection sampling. On the NVFP4
92
+ CUDA-graph path, zero-context concurrency-1 decode improved from 81.9 to
93
+ 89.0 tokens/s (+8.7%). A sustained sample accepted 2,113 of 6,195 draft tokens
94
+ (34.1%); acceptance is workload-dependent. The server allocates exactly
95
+ 499,968 KV tokens at the published settings—500,000 crosses one additional
96
+ cache block.
97
+
98
+ ### Docker Compose
99
+
100
+ Download `runtime/compose.sm120-tp2.yaml` from this repo, set the model path if
101
+ needed, and run:
102
+
103
+ ```bash
104
+ GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-EXL3-4bpw \
105
+ docker compose -f compose.sm120-tp2.yaml up -d
106
+ ```
107
+
108
+ ### Serve script
109
+
110
+ The published `runtime/serve-glm53-sm120-tp2.sh` defaults to GPUs 0,1, port
111
+ 8012, NVFP4/DCP1/MTP5, and the immutable image digest:
112
+
113
+ ```bash
114
+ chmod +x serve-glm53-sm120-tp2.sh
115
+ MODEL=/absolute/path/to/GLM-5.3-Flash-EXL3-4bpw ./serve-glm53-sm120-tp2.sh
116
+ ```
117
+
118
+ Use `CACHE=fp8_ds_mla`, `DCP=2`, or `MTP_TOKENS=0` for controlled variants.