File size: 13,315 Bytes
3ac30cc
 
 
5ab363a
64628f0
aba59d2
 
3ac30cc
 
 
64628f0
3ac30cc
 
 
5ab363a
 
aba59d2
b20c49b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3ac30cc
 
5ab363a
3ac30cc
5ab363a
1ae6d70
 
 
 
 
61e26e1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1ae6d70
 
 
 
 
 
 
 
 
 
 
 
3ac30cc
5ab363a
3ac30cc
 
5ab363a
 
3ac30cc
523482f
5ab363a
 
 
 
 
523482f
3ac30cc
 
64628f0
523482f
64628f0
5ab363a
 
523482f
3ac30cc
 
4739eb1
5ab363a
4739eb1
3ac30cc
 
64628f0
3ac30cc
767b027
64628f0
5ab363a
3ac30cc
5ab363a
523482f
767b027
5ab363a
 
 
 
 
3224669
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1ae6d70
 
 
 
 
 
 
 
 
 
 
 
 
 
3224669
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1ae6d70
 
 
 
3224669
5ab363a
 
 
3224669
 
 
 
5ab363a
 
 
3224669
 
 
 
 
 
5ab363a
 
 
 
3224669
 
 
5ab363a
 
3224669
 
 
 
 
 
 
 
 
 
5ab363a
3224669
 
5ab363a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3ac30cc
64628f0
 
 
 
5ab363a
64628f0
3ac30cc
aba59d2
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
---
base_model: zai-org/GLM-5.3-Flash-BF16
library_name: transformers
pipeline_tag: image-text-to-text
license: other
license_name: shapleymcg-license-1.0
license_link: https://github.com/brandonmmusic-max/shapleymcg/blob/main/LICENSE
tags:
  - glm
  - exl3
  - tr3
  - vllm
  - sm120
  - nvfp4
  - dflash2
  - multimodal
  - shapleymcg
model-index:
  - name: GLM-5.3-Flash-tr3-4bpw
    results:
      - task:
          type: text-generation
          name: Distribution fidelity (KL divergence vs BF16 reference)
        dataset:
          type: brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits
          name: sealed 25-window panel, 51,175 scored positions
          revision: 95f4fdd94bf29989db2e0d1054e4931f55edb6aa
        metrics:
          - type: kl_divergence
            name: Mean tokenwise KLD (reference || candidate), nats
            value: 0.024554564249958208
            args:
              units: nats
              higher_is_better: false
              direction: reference_to_candidate
              estimator: full_vocabulary_fp64

x_fidelity:
  spec: https://github.com/malaiwah/glm53-flash-fidelity-suite/blob/main/docs/CARD-ANNOTATION-SPEC.md
  spec_version: fidelity-provenance/v1
  role: quant
  reference_model: zai-org/GLM-5.3-Flash-BF16
  reference_revision: a6c167b62691b2bac901344b65cb651a70f53e43
  fidelity_dataset: null
  registry:
    dataset: malaiwah/quant-fidelity-registry
    measurement_ids:
      - measurement--glm53.brandonmusic-4bpw.brandonmusic-final25
      - measurement--glm53.brandonmusic-4bpw.brandonmusic-final25.clean17
      - measurement--glm53.tr3-4bpw-stream.brandonmusic-final25
  lane: sealed-ep8
  scope: routed_experts_only
  head_bits: 16
  measured_by: first-party (sealed lane; third-party streaming cross-measurement in registry)
---

# GLM-5.3-Flash TR3 4bpw — current SM120 runtime

This is the uniform-K4 EXL3/TR3 routed-expert checkpoint for GLM-5.3-Flash.
The current v84 runtime supports three explicit TP2/EP2/DCP2 profiles on two
SM120 GPUs: multimodal DFlash2, language-only DFlash2, and language-only MTP3.
All use calibrated NVFP4 MLA KV and CUDA graphs. This is a custom vLLM/B12X
build and is not compatible with stock upstream vLLM.

## Encoder reproducibility closure

The repository now contains the complete, hash-verified R10 Python encoder
closure used by the EXL3/MCG adapter, including
`r7_encoder/r10_codec.py` (`R10TrellisCodec`) and the pinned
`lineage/encode_tr3_v31.py` numeric core. It is published under
[`reproducibility/r10/`](reproducibility/r10/) with a per-file SHA-256
manifest and an offline verifier:

```bash
python3 reproducibility/r10/verify_bundle.py
```

The bundle is byte-identical to the immutable prior-control source at Hugging
Face revision `7c73450f05a151439d0f184f216b1eefcc394a31`. It contains the
portable Python/numeric source, not a compiled `exllamav3_ext`; that binary
must still be built for the target PyTorch, CUDA, and SM ABI and is independently
hash-bound by the adapter. See the [bundle README](reproducibility/r10/README.md)
for the exact adapter paths, lineage boundary, and licensing.

## Pick a serving profile

| Goal | Launcher | Extra checkpoint | Measured KV tokens |
|---|---|---|---:|
| Images plus fastest measured C1 decode | `compose.sm120-tp2.yaml` | DFlash2-7 | 129,473 |
| Text-only DFlash2 decode | `compose.sm120-tp2-language-only-dflash2.yaml` | DFlash2-7 | 184,619 |
| **Text-only capacity/default** | **`compose.sm120-tp2-language-only.yaml`** | **none; built-in MTP3** | **1,376,256** |

The MTP3 option means the model's built-in MTP head only: it does not load or
mount the external DFlash checkpoint. Choose DFlash2 when its modest C1 decode
gain matters more than resident context/concurrency; choose MTP3 for the normal
text-only daily driver.

## Run the current image

```text
verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-dflash2
OCI digest: sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692
```

The runtime image does not contain either checkpoint. Download/mount this
EXL3 model and `incoai/GLM-5.3-Flash-DFlash2` separately. The DFlash2
checkpoint is distributed under CC-BY-NC-ND-4.0; review its license before use.

Docker Compose:

```bash
curl -L -o compose.sm120-tp2.yaml \
  https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2.yaml

GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
GLM53_DFLASH_PATH=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
docker compose -f compose.sm120-tp2.yaml up -d

curl http://127.0.0.1:8012/v1/models
```

Standalone serve script:

```bash
curl -L -o serve-glm53-sm120-tp2.sh \
  https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/serve-glm53-sm120-tp2.sh
chmod +x serve-glm53-sm120-tp2.sh

MODEL=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
DFLASH_MODEL=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
GPU_DEVICES=0,1 \
./serve-glm53-sm120-tp2.sh
```

The published profile has a 98,304-token request ceiling and allocated 129,473
KV tokens on the qualified pair. Its hybrid Mamba/DFlash rollback layout has
room for one full resident request; additional requests queue. C2/C4 rows in
the raw benchmark are therefore capacity-limited and are not throughput claims.

## Language-only profile

For text serving, use the language-only profile. It disables the vision tower
and uses the built-in MTP3 head by default, avoiding a second external
checkpoint and leaving substantially more room for KV cache:

```bash
curl -L -o compose.sm120-tp2-language-only.yaml \
  https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2-language-only.yaml

GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
docker compose -f compose.sm120-tp2-language-only.yaml up -d
```

Standalone:

```bash
curl -L -o serve-glm53-sm120-tp2-language-only.sh \
  https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/serve-glm53-sm120-tp2-language-only.sh
chmod +x serve-glm53-sm120-tp2-language-only.sh

MODEL=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
GPU_DEVICES=0,1 \
./serve-glm53-sm120-tp2-language-only.sh
```

To keep DFlash2 while disabling vision, use the separate speed-first launcher:

```bash
curl -L -o compose.sm120-tp2-language-only-dflash2.yaml \
  https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2-language-only-dflash2.yaml

GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
GLM53_DFLASH_PATH=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
docker compose -f compose.sm120-tp2-language-only-dflash2.yaml up -d
```

Its standalone equivalent is
[`serve-glm53-sm120-tp2-language-only-dflash2.sh`](runtime/serve-glm53-sm120-tp2-language-only-dflash2.sh).

The language-only alias points to the same tested v84 code digest:

```text
verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only
OCI digest: sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692
```

Measured capacity on the same two 96 GB GPUs at 300 W each:

| Runtime profile | Vision | Speculator | KV tokens | Concurrency at tested ceiling |
|---|:---:|---|---:|---:|
| Multimodal DFlash2-7, 98,304 max | on | external 7-layer draft | 129,473 | 1.32x |
| Language-only DFlash2-7, 98,304 max | off | external 7-layer draft | 184,619 | 1.88x |
| **Language-only MTP3, 131,072 max** | **off** | **built-in head** | **1,376,256** | **10.50x** |

Turning vision off raises the DFlash KV token pool by 42.6%. The much larger
7.45x language-only gain comes from using the built-in MTP head instead of
keeping the external DFlash2-7 model resident. It is not a vision-only gain.
The language-only profiles do not accept image inputs. The reported KV-token
pool is total allocated capacity, not a promise that every request can use the
entire pool; the configured per-request ceiling and scheduler concurrency still
apply.

## Current measured results

Qualified on two RTX PRO 6000 Blackwell Workstation Edition GPUs (96 GB each),
TP2/EP2/DCP2, NVFP4 MLA KV, prefix cache off, and DFlash2-7. The current quick
speed pass used 600 W limits and +6000 MHz memory offsets. Generation uses the
model defaults (`temperature=1.0`, `top_p=0.95`); the acceptance comparison uses
`reasoning_effort=max`.

| Measurement | Result |
|---|---:|
| Cold prefill, 32K | **6,225 client / 6,277 server tok/s** |
| Cold prefill, 64K | **6,083 client / 6,130 server tok/s** |
| C1 decode, empty context | **145.5 tok/s** |
| C1 decode, 32K context | **147.2 tok/s** |
| C1 decode, 64K context | **151.5 tok/s** |
| DFlash2 acceptance, 5 distinct GSM8K prompts | **5.428 mean / 5.441 token-weighted; 5/5 correct** |
| DFlash2 acceptance, GSM8K first 16 | **5.739 mean / 5.550 token-weighted** |
| DFlash2 acceptance, published reference | 5.78 mean over 128 samples |
| Image smoke | **pass** — correctly identified a mallard |

The clean C1 decode run used a 4,096-token completion budget so the client did
not roll into the next prefill request. A prior 60.1 tok/s row was a harness
rollover artifact and is excluded. The DFlash acceptance fix is material: the partially ported Triton mask scored
1.017 weighted. Restoring the reference semantics—full bidirectional visibility
inside the draft block with a backward-only historical window—raised the same
five-seed probe to 5.068 and the exact GSM8K sample to 5.739. Synthetic padded
long-context decode accepts roughly 2.8–3.0 tokens/step, while five distinct
GSM8K reasoning prompts accepted 4.89–6.03 and all answered correctly; acceptance
is workload-dependent.

Receipts: [600 W prefill JSON](runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.json),
[600 W prefill TUI](runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.tui.log),
[clean 600 W C1 decode JSON](runtime-results/v84/benchmarks/llm-decode-c1-clean-4096-600w.json),
[clean 600 W C1 decode TUI](runtime-results/v84/benchmarks/llm-decode-c1-clean-4096-600w.tui.log),
[earlier C1-C4 benchmark](runtime-results/v84/benchmarks/llm-decode-c1-c4-64k.json),
[acceptance rows](runtime-results/v84/quality/gsm8k-first16-max-acceptance.jsonl),
[distinct-prompt acceptance](runtime-results/v84/quality/gsm8k-distinct5-language-only-acceptance.json),
[language-only capacity](runtime-results/v84/validation/language-only-capacity.json),
and [release validation](runtime-results/v84/validation/release.json).

## Quality and KLD

v84 changes draft speculation, Triton draft-attention semantics, and vision
packaging; it does not change target-model weights, EXL3 kernels, calibrated
MLA KV scales, or target logits. The current target-quality receipts therefore
remain the repeatedly qualified v75 measurements:

| Test | Result |
|---|---:|
| FP8 MLA KV KLD, five-run full 2,047-position mean | **0.024610591221** |
| NVFP4 MLA KV KLD, five-run full 2,047-position mean | **0.054757372223** |
| Estonia 10x, NVFP4 | **10/10** |
| LAVD-low 10x, FP8 | **8/10 accepted** |
| LAVD-low 10x, NVFP4 | **3/10 accepted** — failed quality gate |
| Needle through 500K, NVFP4 | **17/18 raw; final cell passed on longer retry** |

KLD was measured in eager/no-speculation mode against the sealed BF16 teacher
over every causal position in the 2,048-token window. Draft acceptance does not
alter that target-logit measurement. Hotel was explicitly stopped and is not
presented as a current result.

Receipts: [v75 KLD and quality evidence](runtime-results/v75/). Older tuning
history is retained in [the historical model card](docs/HISTORICAL_MODEL_CARD_2026-08-27.md),
not mixed into the current launch path.

## Vision and implementation notes

The image fixes a packaging defect where GLM-5.3 vision RoPE unconditionally
imported `vllm.vllm_flash_attn.layers.rotary` even when a custom wheel shipped
only the compiled flash-attention extensions. It now uses native PyTorch RoPE
as a correctness fallback. Cold multimodal warmup and a real remote-JPEG chat
request both passed.

The target path remains the fused uniform-K4 EXL3 route-128 SMEM/register
kernel. SM120 in this build does not use a TMEM/TCGEN path. DFlash uses Triton
attention because its noncausal sliding-window semantics are now tested there.

## Provenance and attribution

The image embeds `/opt/glm53/PROVENANCE.json` and OCI source, author,
documentation, revision, checkpoint, and validation labels. The manifest binds
the runtime source and benchmark artifacts with SHA-256 hashes. This is a
transparent provenance fingerprint: there is no telemetry, callback, hidden
output watermark, or inference modification.

```bash
curl -L -o verify-provenance.sh \
  https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/verify-provenance.sh
chmod +x verify-provenance.sh
./verify-provenance.sh
```

This checkpoint is distributed under the ShapleyMCG License 1.0
([LICENSE](https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/blob/main/LICENSE)).
Derivatives must declare it with the exact identifier fixed in the LICENSE
appendix: `license: other`, `license_name: shapleymcg-license-1.0`,
`license_link: https://github.com/brandonmmusic-max/shapleymcg/blob/main/LICENSE`.
The runtime foundation comes from Local Inference Lab contributors.