You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

GLM-5.3 MixedK EXL3 3.38 bpw

An EXL3 quantization of zai-org/GLM-5.3, made with SAGE.

SAGE dynamically and intelligently assigns bit widths across the model, making this a MixedK EXL3. It averages 3.38 bits per weight, and every one of the 19,200 routed experts is kept: no pruning and no expert merging.

The pack is sized so the weights plus a full 1,048,576-token KV cache at 4 bits fit in the memory of four NVIDIA DGX Sparks.

Summary

Base model zai-org/GLM-5.3 (78 layers, 256 routed experts per MoE layer, 8 active)
Format EXL3
Quantization SAGE MixedK
Body bitrate 3.38 bpw nominal, 3.39 bpw including scales
Output head 8-bit
Routed experts all 256 per layer kept (19,200 total)
MTP / draft layer not included
Max context 1,048,576 tokens (unchanged from the base model)
Total size 319.0 GB (297.1 GiB), 58 weight shards

Quality

Measured against the original BF16 model on a held-out evaluation set: text the quantizer never saw.

The held-out set: 10 sequences of 1,024 tokens each (10,240 scored positions), built from the test splits of public benchmarks:

  • Web text: WikiText-103 (4 sequences)
  • Code: HumanEval (2)
  • Math: GSM8K (2)
  • Chat: UltraChat-200k (2)

Each sequence is whole documents packed end to end. Math and chat examples are formatted with GLM-5.3's own chat template. None of this text was used while quantizing the model.

Scoring: both models read the same tokens. At every position their next-token predictions are compared over the full 154,880-token vocabulary, in float64.

Metric Result
Top-1 agreement with BF16 92.98% (9,521 / 10,240)
BF16 top-1 token within the quant's top 5 99.38%
Mean KL divergence (BF16 ‖ quant) 0.0948
Median KL divergence 0.0018
99th-percentile KL divergence 1.61
Top-5 set overlap 0.827
Mean NLL, BF16 → quant 1.0092 → 1.0300
Perplexity increase +2.1%

Scope of these numbers:

  • The BF16 reference is the original weights run through the same ExLlamaV3 runtime, not the vendor's own implementation.
  • Scores come from full-sequence forward passes at 1,024 tokens. Long-context and cached generation were not part of this evaluation.

Memory budget: four DGX Sparks at 1M context

Item Size
Weights 319.0 GB
KV cache, 1,048,576 tokens at Q4 (MLA latent plus indexer keys) ~39.7 GB
Total ~358.7 GB of 512 GB (4 × 128 GB)

The rest is left for activations, the runtime and the OS. This is a sizing target only: four-node serving and full 1M-token inference have not been validated yet.

Runtime

GLM-5.3 support (DSA sparse attention, the F32 router bias and MixedK experts) is in a fork of exllamav3, at commit affc194d5476710f167f30729e2508e577759612. Stock exllamav3 cannot load this model yet. Installation and usage instructions will be added here when the runtime is released.

Download

hf download vcruz305/GLM-5.3-EXL3-3.38bpw --local-dir GLM-5.3-EXL3-3.38bpw

License

Same license as the base model; see zai-org/GLM-5.3.

Downloads last month
31
Safetensors
Model size
159B params
Tensor type
BF16
·
F16
·
I16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vcruz305/GLM-5.3-EXL3-3.38bpw

Base model

zai-org/GLM-5.3
Quantized
(70)
this model