--- base_model: zai-org/GLM-5.3-Flash-BF16 library_name: transformers license: mit tags: - glm5 - exl3 - quantized - mixture-of-experts --- # GLM-5.3 Flash EXL3 K3.25 Projection-aware K3/K4 EXL3/MCG for the routed experts in [`zai-org/GLM-5.3-Flash-BF16`](https://huggingface.co/zai-org/GLM-5.3-Flash-BF16), with everything else retained from the official checkpoint. ## What moved—and what did not - **K3.25 target experts:** layers 3…44 start at K3, then exactly 9,072 of 36,288 projections move to K4. The promotion budget is gate/up/down = **1,701 / 2,835 / 4,536** (the 3:5:8 quality allocation). - **Straight-K3 MTP:** all 864 checkpoint-MTP projections in layer 45 stay K3. It remains valid for diagnostics, but the serving recipe defaults to DFlash2, so extra K4 bytes and search time are not spent there. - **37,152 final quantized projections** were produced in one target-plus-MTP `model.quantize(...)` call. - Attention, dense MLPs, shared experts, routers, norms, embeddings, vision, and all other tensors remain native precision. - The release audit compared **1,618 native tensors / 18.01 GiB byte-for-byte** against source revision `f12e0fe1f6b2ea274c11a569582edfd99d993c5e`. - Packed checkpoint payload: **136.16 GiB** across 18 safetensors shards. Quantization used GPTQModel `a053382584fa58cba7bf212ef1b829d08b29b2c0`, EXL3 MCG adjacent tiers, seed 787, `sigma_reg=0.025`, automatic output-scale selection, and a fixed 1,426-record calibration corpus (`sha256:4e569625d97865777da92167b8fbf6fabb4ab7adf55baa5d676c81ce8dd95244`). Within every target layer, K4 promotion follows measured Hessian-weighted relative error times natural gate-squared mass while the global 3:5:8 budget stays exact. Natural GLM router traffic supplied expert Hessians; the committed recovery contract covers any expert below the 1,024-route floor. Allocation provenance and the validation report ship with the model; the full error ledger is retained separately as internal quantization evidence. ## Qualified two-GPU serving result Release `v0.7.0` of the matching B12x/vLLM recipe qualified the public target revision with DFlash2 K5, FP8 MLA, TP2 + EP2 + DCP2, vision up to 16 images, and a 1,048,576-token request limit on 2× RTX PRO 6000 Blackwell at a 400 W power cap per GPU. - Code-agent decode: **213 tok/s at C1** and **832 tok/s at C16**. - 128K cold prefill: **4,572 tok/s**. - Available KV cache: **2,758,919 tokens**. - Exact 1M multi-needle test: **6/6 needles recovered**. - Exact 69-case thinking tool-call suite at parallelism 8: **86/100** (55 pass, 8 partial, 6 fail), versus 88/100 for the matched uniform-K3 control. The full prompts, outputs, scoring, performance curves, and machine-readable receipts are in the recipe's [`v0.7.0` benchmark report](https://github.com/tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx/blob/v0.7.0/benchmarks/RESULTS.md). ## Serving This is intended for the optimized two-GPU GLM-5.3 recipe at [`tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx`](https://github.com/tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx). That runtime carries the GLM/EXL3/MTP and B12x ports; this card does not claim stock-vLLM support. Huge thanks to **Z.ai** for GLM-5.3 Flash, **Brandon** for the earlier K4 quant and qualification work that helped inform this run, **MiaAI-Lab** for the nearby dual-DGX-Spark reference, and the GPTQModel, ExLlamaV3, vLLM, and B12x contributors. ## Audit identity - Source index: `sha256:e6007bd58fb7e07f9fe69544257ee2713f252ef5855bbf685b48c991d524ef0f` - One-shot plan: `sha256:cb8ee0b49c6a41bb31e9813214193fa9426c209dd0a0ca9f52ce8b5ab61465d8` - Validation: `sha256:30e5f8f5fc6c1d10cdd17749b40559a0dca89bec89b5affad219cab0f0260160`