GLM-5.3 MixedK EXL3 3.38 bpw
An EXL3 quantization of zai-org/GLM-5.3, made with SAGE.
SAGE dynamically and intelligently assigns bit widths across the model, making this a MixedK EXL3. It averages 3.38 bits per weight, and every one of the 19,200 routed experts is kept: no pruning and no expert merging.
The pack is sized so the weights plus a full 1,048,576-token KV cache at 4 bits fit in the memory of four NVIDIA DGX Sparks.
Summary
| Base model | zai-org/GLM-5.3 (78 layers, 256 routed experts per MoE layer, 8 active) |
| Format | EXL3 |
| Quantization | SAGE MixedK |
| Body bitrate | 3.38 bpw nominal, 3.39 bpw including scales |
| Output head | 8-bit |
| Routed experts | all 256 per layer kept (19,200 total) |
| MTP / draft layer | not included |
| Max context | 1,048,576 tokens (unchanged from the base model) |
| Total size | 319.0 GB (297.1 GiB), 58 weight shards |
Quality
Measured against the original BF16 model on a held-out evaluation set: text the quantizer never saw.
The held-out set: 10 sequences of 1,024 tokens each (10,240 scored positions), built from the test splits of public benchmarks:
- Web text: WikiText-103 (4 sequences)
- Code: HumanEval (2)
- Math: GSM8K (2)
- Chat: UltraChat-200k (2)
Each sequence is whole documents packed end to end. Math and chat examples are formatted with GLM-5.3's own chat template. None of this text was used while quantizing the model.
Scoring: both models read the same tokens. At every position their next-token predictions are compared over the full 154,880-token vocabulary, in float64.
| Metric | Result |
|---|---|
| Top-1 agreement with BF16 | 92.98% (9,521 / 10,240) |
| BF16 top-1 token within the quant's top 5 | 99.38% |
| Mean KL divergence (BF16 ‖ quant) | 0.0948 |
| Median KL divergence | 0.0018 |
| 99th-percentile KL divergence | 1.61 |
| Top-5 set overlap | 0.827 |
| Mean NLL, BF16 → quant | 1.0092 → 1.0300 |
| Perplexity increase | +2.1% |
Scope of these numbers:
- The BF16 reference is the original weights run through the same ExLlamaV3 runtime, not the vendor's own implementation.
- Scores come from full-sequence forward passes at 1,024 tokens. Long-context and cached generation were not part of this evaluation.
Memory budget: four DGX Sparks at 1M context
| Item | Size |
|---|---|
| Weights | 319.0 GB |
| KV cache, 1,048,576 tokens at Q4 (MLA latent plus indexer keys) | ~39.7 GB |
| Total | ~358.7 GB of 512 GB (4 × 128 GB) |
The rest is left for activations, the runtime and the OS. This is a sizing target only: four-node serving and full 1M-token inference have not been validated yet.
Runtime
GLM-5.3 support (DSA sparse attention, the F32 router bias and MixedK experts) is in a fork of exllamav3, at commit affc194d5476710f167f30729e2508e577759612. Stock exllamav3 cannot load this model yet. Installation and usage instructions will be added here when the runtime is released.
Download
hf download vcruz305/GLM-5.3-EXL3-3.38bpw --local-dir GLM-5.3-EXL3-3.38bpw
License
Same license as the base model; see zai-org/GLM-5.3.
- Downloads last month
- 31
Model tree for vcruz305/GLM-5.3-EXL3-3.38bpw
Base model
zai-org/GLM-5.3