Image-Text-to-Text
Transformers
Safetensors
glm5_next
glm
exl3
tr3
vllm
sm120
nvfp4
dflash2
multimodal
shapleymcg
conversational
Eval Results (legacy)
4-bit precision
Instructions to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="brandonmusic/GLM-5.3-Flash-tr3-4bpw") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw") model = AutoModelForMultimodalLM.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.3-Flash-tr3-4bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
- SGLang
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
File size: 13,315 Bytes
3ac30cc 5ab363a 64628f0 aba59d2 3ac30cc 64628f0 3ac30cc 5ab363a aba59d2 b20c49b 3ac30cc 5ab363a 3ac30cc 5ab363a 1ae6d70 61e26e1 1ae6d70 3ac30cc 5ab363a 3ac30cc 5ab363a 3ac30cc 523482f 5ab363a 523482f 3ac30cc 64628f0 523482f 64628f0 5ab363a 523482f 3ac30cc 4739eb1 5ab363a 4739eb1 3ac30cc 64628f0 3ac30cc 767b027 64628f0 5ab363a 3ac30cc 5ab363a 523482f 767b027 5ab363a 3224669 1ae6d70 3224669 1ae6d70 3224669 5ab363a 3224669 5ab363a 3224669 5ab363a 3224669 5ab363a 3224669 5ab363a 3224669 5ab363a 3ac30cc 64628f0 5ab363a 64628f0 3ac30cc aba59d2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 | ---
base_model: zai-org/GLM-5.3-Flash-BF16
library_name: transformers
pipeline_tag: image-text-to-text
license: other
license_name: shapleymcg-license-1.0
license_link: https://github.com/brandonmmusic-max/shapleymcg/blob/main/LICENSE
tags:
- glm
- exl3
- tr3
- vllm
- sm120
- nvfp4
- dflash2
- multimodal
- shapleymcg
model-index:
- name: GLM-5.3-Flash-tr3-4bpw
results:
- task:
type: text-generation
name: Distribution fidelity (KL divergence vs BF16 reference)
dataset:
type: brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits
name: sealed 25-window panel, 51,175 scored positions
revision: 95f4fdd94bf29989db2e0d1054e4931f55edb6aa
metrics:
- type: kl_divergence
name: Mean tokenwise KLD (reference || candidate), nats
value: 0.024554564249958208
args:
units: nats
higher_is_better: false
direction: reference_to_candidate
estimator: full_vocabulary_fp64
x_fidelity:
spec: https://github.com/malaiwah/glm53-flash-fidelity-suite/blob/main/docs/CARD-ANNOTATION-SPEC.md
spec_version: fidelity-provenance/v1
role: quant
reference_model: zai-org/GLM-5.3-Flash-BF16
reference_revision: a6c167b62691b2bac901344b65cb651a70f53e43
fidelity_dataset: null
registry:
dataset: malaiwah/quant-fidelity-registry
measurement_ids:
- measurement--glm53.brandonmusic-4bpw.brandonmusic-final25
- measurement--glm53.brandonmusic-4bpw.brandonmusic-final25.clean17
- measurement--glm53.tr3-4bpw-stream.brandonmusic-final25
lane: sealed-ep8
scope: routed_experts_only
head_bits: 16
measured_by: first-party (sealed lane; third-party streaming cross-measurement in registry)
---
# GLM-5.3-Flash TR3 4bpw — current SM120 runtime
This is the uniform-K4 EXL3/TR3 routed-expert checkpoint for GLM-5.3-Flash.
The current v84 runtime supports three explicit TP2/EP2/DCP2 profiles on two
SM120 GPUs: multimodal DFlash2, language-only DFlash2, and language-only MTP3.
All use calibrated NVFP4 MLA KV and CUDA graphs. This is a custom vLLM/B12X
build and is not compatible with stock upstream vLLM.
## Encoder reproducibility closure
The repository now contains the complete, hash-verified R10 Python encoder
closure used by the EXL3/MCG adapter, including
`r7_encoder/r10_codec.py` (`R10TrellisCodec`) and the pinned
`lineage/encode_tr3_v31.py` numeric core. It is published under
[`reproducibility/r10/`](reproducibility/r10/) with a per-file SHA-256
manifest and an offline verifier:
```bash
python3 reproducibility/r10/verify_bundle.py
```
The bundle is byte-identical to the immutable prior-control source at Hugging
Face revision `7c73450f05a151439d0f184f216b1eefcc394a31`. It contains the
portable Python/numeric source, not a compiled `exllamav3_ext`; that binary
must still be built for the target PyTorch, CUDA, and SM ABI and is independently
hash-bound by the adapter. See the [bundle README](reproducibility/r10/README.md)
for the exact adapter paths, lineage boundary, and licensing.
## Pick a serving profile
| Goal | Launcher | Extra checkpoint | Measured KV tokens |
|---|---|---|---:|
| Images plus fastest measured C1 decode | `compose.sm120-tp2.yaml` | DFlash2-7 | 129,473 |
| Text-only DFlash2 decode | `compose.sm120-tp2-language-only-dflash2.yaml` | DFlash2-7 | 184,619 |
| **Text-only capacity/default** | **`compose.sm120-tp2-language-only.yaml`** | **none; built-in MTP3** | **1,376,256** |
The MTP3 option means the model's built-in MTP head only: it does not load or
mount the external DFlash checkpoint. Choose DFlash2 when its modest C1 decode
gain matters more than resident context/concurrency; choose MTP3 for the normal
text-only daily driver.
## Run the current image
```text
verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-dflash2
OCI digest: sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692
```
The runtime image does not contain either checkpoint. Download/mount this
EXL3 model and `incoai/GLM-5.3-Flash-DFlash2` separately. The DFlash2
checkpoint is distributed under CC-BY-NC-ND-4.0; review its license before use.
Docker Compose:
```bash
curl -L -o compose.sm120-tp2.yaml \
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2.yaml
GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
GLM53_DFLASH_PATH=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
docker compose -f compose.sm120-tp2.yaml up -d
curl http://127.0.0.1:8012/v1/models
```
Standalone serve script:
```bash
curl -L -o serve-glm53-sm120-tp2.sh \
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/serve-glm53-sm120-tp2.sh
chmod +x serve-glm53-sm120-tp2.sh
MODEL=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
DFLASH_MODEL=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
GPU_DEVICES=0,1 \
./serve-glm53-sm120-tp2.sh
```
The published profile has a 98,304-token request ceiling and allocated 129,473
KV tokens on the qualified pair. Its hybrid Mamba/DFlash rollback layout has
room for one full resident request; additional requests queue. C2/C4 rows in
the raw benchmark are therefore capacity-limited and are not throughput claims.
## Language-only profile
For text serving, use the language-only profile. It disables the vision tower
and uses the built-in MTP3 head by default, avoiding a second external
checkpoint and leaving substantially more room for KV cache:
```bash
curl -L -o compose.sm120-tp2-language-only.yaml \
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2-language-only.yaml
GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
docker compose -f compose.sm120-tp2-language-only.yaml up -d
```
Standalone:
```bash
curl -L -o serve-glm53-sm120-tp2-language-only.sh \
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/serve-glm53-sm120-tp2-language-only.sh
chmod +x serve-glm53-sm120-tp2-language-only.sh
MODEL=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
GPU_DEVICES=0,1 \
./serve-glm53-sm120-tp2-language-only.sh
```
To keep DFlash2 while disabling vision, use the separate speed-first launcher:
```bash
curl -L -o compose.sm120-tp2-language-only-dflash2.yaml \
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/compose.sm120-tp2-language-only-dflash2.yaml
GLM53_MODEL_PATH=/absolute/path/to/GLM-5.3-Flash-tr3-4bpw \
GLM53_DFLASH_PATH=/absolute/path/to/GLM-5.3-Flash-DFlash2 \
docker compose -f compose.sm120-tp2-language-only-dflash2.yaml up -d
```
Its standalone equivalent is
[`serve-glm53-sm120-tp2-language-only-dflash2.sh`](runtime/serve-glm53-sm120-tp2-language-only-dflash2.sh).
The language-only alias points to the same tested v84 code digest:
```text
verdictai/glm53-flash-exl3-k4:r19-sm120-tp2-ep2-dcp2-v84-language-only
OCI digest: sha256:0f1cdcc8891f1cc3a444121eb61d366289a1cbba285f0892dcbb24bc94961692
```
Measured capacity on the same two 96 GB GPUs at 300 W each:
| Runtime profile | Vision | Speculator | KV tokens | Concurrency at tested ceiling |
|---|:---:|---|---:|---:|
| Multimodal DFlash2-7, 98,304 max | on | external 7-layer draft | 129,473 | 1.32x |
| Language-only DFlash2-7, 98,304 max | off | external 7-layer draft | 184,619 | 1.88x |
| **Language-only MTP3, 131,072 max** | **off** | **built-in head** | **1,376,256** | **10.50x** |
Turning vision off raises the DFlash KV token pool by 42.6%. The much larger
7.45x language-only gain comes from using the built-in MTP head instead of
keeping the external DFlash2-7 model resident. It is not a vision-only gain.
The language-only profiles do not accept image inputs. The reported KV-token
pool is total allocated capacity, not a promise that every request can use the
entire pool; the configured per-request ceiling and scheduler concurrency still
apply.
## Current measured results
Qualified on two RTX PRO 6000 Blackwell Workstation Edition GPUs (96 GB each),
TP2/EP2/DCP2, NVFP4 MLA KV, prefix cache off, and DFlash2-7. The current quick
speed pass used 600 W limits and +6000 MHz memory offsets. Generation uses the
model defaults (`temperature=1.0`, `top_p=0.95`); the acceptance comparison uses
`reasoning_effort=max`.
| Measurement | Result |
|---|---:|
| Cold prefill, 32K | **6,225 client / 6,277 server tok/s** |
| Cold prefill, 64K | **6,083 client / 6,130 server tok/s** |
| C1 decode, empty context | **145.5 tok/s** |
| C1 decode, 32K context | **147.2 tok/s** |
| C1 decode, 64K context | **151.5 tok/s** |
| DFlash2 acceptance, 5 distinct GSM8K prompts | **5.428 mean / 5.441 token-weighted; 5/5 correct** |
| DFlash2 acceptance, GSM8K first 16 | **5.739 mean / 5.550 token-weighted** |
| DFlash2 acceptance, published reference | 5.78 mean over 128 samples |
| Image smoke | **pass** — correctly identified a mallard |
The clean C1 decode run used a 4,096-token completion budget so the client did
not roll into the next prefill request. A prior 60.1 tok/s row was a harness
rollover artifact and is excluded. The DFlash acceptance fix is material: the partially ported Triton mask scored
1.017 weighted. Restoring the reference semantics—full bidirectional visibility
inside the draft block with a backward-only historical window—raised the same
five-seed probe to 5.068 and the exact GSM8K sample to 5.739. Synthetic padded
long-context decode accepts roughly 2.8–3.0 tokens/step, while five distinct
GSM8K reasoning prompts accepted 4.89–6.03 and all answered correctly; acceptance
is workload-dependent.
Receipts: [600 W prefill JSON](runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.json),
[600 W prefill TUI](runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.tui.log),
[clean 600 W C1 decode JSON](runtime-results/v84/benchmarks/llm-decode-c1-clean-4096-600w.json),
[clean 600 W C1 decode TUI](runtime-results/v84/benchmarks/llm-decode-c1-clean-4096-600w.tui.log),
[earlier C1-C4 benchmark](runtime-results/v84/benchmarks/llm-decode-c1-c4-64k.json),
[acceptance rows](runtime-results/v84/quality/gsm8k-first16-max-acceptance.jsonl),
[distinct-prompt acceptance](runtime-results/v84/quality/gsm8k-distinct5-language-only-acceptance.json),
[language-only capacity](runtime-results/v84/validation/language-only-capacity.json),
and [release validation](runtime-results/v84/validation/release.json).
## Quality and KLD
v84 changes draft speculation, Triton draft-attention semantics, and vision
packaging; it does not change target-model weights, EXL3 kernels, calibrated
MLA KV scales, or target logits. The current target-quality receipts therefore
remain the repeatedly qualified v75 measurements:
| Test | Result |
|---|---:|
| FP8 MLA KV KLD, five-run full 2,047-position mean | **0.024610591221** |
| NVFP4 MLA KV KLD, five-run full 2,047-position mean | **0.054757372223** |
| Estonia 10x, NVFP4 | **10/10** |
| LAVD-low 10x, FP8 | **8/10 accepted** |
| LAVD-low 10x, NVFP4 | **3/10 accepted** — failed quality gate |
| Needle through 500K, NVFP4 | **17/18 raw; final cell passed on longer retry** |
KLD was measured in eager/no-speculation mode against the sealed BF16 teacher
over every causal position in the 2,048-token window. Draft acceptance does not
alter that target-logit measurement. Hotel was explicitly stopped and is not
presented as a current result.
Receipts: [v75 KLD and quality evidence](runtime-results/v75/). Older tuning
history is retained in [the historical model card](docs/HISTORICAL_MODEL_CARD_2026-08-27.md),
not mixed into the current launch path.
## Vision and implementation notes
The image fixes a packaging defect where GLM-5.3 vision RoPE unconditionally
imported `vllm.vllm_flash_attn.layers.rotary` even when a custom wheel shipped
only the compiled flash-attention extensions. It now uses native PyTorch RoPE
as a correctness fallback. Cold multimodal warmup and a real remote-JPEG chat
request both passed.
The target path remains the fused uniform-K4 EXL3 route-128 SMEM/register
kernel. SM120 in this build does not use a TMEM/TCGEN path. DFlash uses Triton
attention because its noncausal sliding-window semantics are now tested there.
## Provenance and attribution
The image embeds `/opt/glm53/PROVENANCE.json` and OCI source, author,
documentation, revision, checkpoint, and validation labels. The manifest binds
the runtime source and benchmark artifacts with SHA-256 hashes. This is a
transparent provenance fingerprint: there is no telemetry, callback, hidden
output watermark, or inference modification.
```bash
curl -L -o verify-provenance.sh \
https://raw.githubusercontent.com/brandonmmusic-max/glm-5.3-flash-exl3-4bpw/main/runtime/verify-provenance.sh
chmod +x verify-provenance.sh
./verify-provenance.sh
```
This checkpoint is distributed under the ShapleyMCG License 1.0
([LICENSE](https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/blob/main/LICENSE)).
Derivatives must declare it with the exact identifier fixed in the LICENSE
appendix: `license: other`, `license_name: shapleymcg-license-1.0`,
`license_link: https://github.com/brandonmmusic-max/shapleymcg/blob/main/LICENSE`.
The runtime foundation comes from Local Inference Lab contributors.
|