| --- |
| license: mit |
| base_model: |
| - zai-org/GLM-5.2 |
| base_model_relation: quantized |
| language: |
| - ja |
| - en |
| - zh |
| tags: |
| - quantization |
| - vector-quantization |
| - aqlm |
| - mixture-of-experts |
| - glm |
| - vllm |
| pipeline_tag: text-generation |
| --- |
| |
| # GLM-5.2 β mixed-bit VQ (AQLM) ~1.86-bit |
|
|
| A **~180 GiB** quantization of **[GLM-5.2](https://huggingface.co/zai-org/GLM-5.2)** |
| (744B Chinese-native reasoning MoE, MIT) that keeps Japanese/English/Chinese |
| **thinking-mode** quality at **~1.86 bit/weight**, using **vector quantization with |
| GPTQ error compensation (AQLM-style)** instead of scalar rounding. |
|
|
| Runs on **2Γ RTX PRO 6000 (sm_120 Blackwell, ~95 GiB each)** via vLLM. |
| |
| ## Why VQ |
| |
| At the same size, scalar mixed-bit rounding loses too much at 1β2 bit. Replacing the |
| scalar codes with a **shared vector codebook + per-row error compensation** recovers most |
| of it (the two together are *super-additive* β neither alone is enough): |
| |
| | Metric | Scalar mixed-bit (same ~180 GiB) | **This (VQ + compensation)** | |
| |---|---|---| |
| | Calibration KL (fake-quant, iso-size) | baseline | **β47 %** | |
| | Greedy arithmetic eval (JA/EN/ZH, terminate + correct) | 21/22 | **22/22** | |
| | 1-bit experts | collapse (KL β 13) | **survive (KL β 0.42)** | |
| |
| > The β47 % KL is a fake-quant, iso-size comparison; the deployable end-to-end signal is |
| > the greedy eval (multi-digit multiplication and word problems in all three languages, |
| > including held-out items not in calibration). |
| |
| ## Serving |
| |
| This is **not** a plug-and-play GGUF β it needs a matching sm_120 stack: |
| |
| - **vLLM** with `GlmMoeDsaForCausalLM` + sm_120 kernels (reference: `jasl/vllm` PR-41834 sm12x preview). |
| - **transformers 5.12**. |
| - The **VQ serving plugin** from **[mmzz164/OneCompression @ `glm-serving-v1`](https://github.com/mmzz164/OneCompression)** β see [`example/glm-5.2/`](https://github.com/mmzz164/OneCompression/tree/glm-serving-v1/example/glm-5.2) for the launcher and full instructions. |
| - **2Γ ~95 GiB sm_120 GPUs**, EP=1 / TP=2 (VQ codes can't be tensor-parallel-sharded). |
| |
| ```bash |
| GLM_CKPT=/path/to/this/model bash start_glm_api_vq.sh # OpenAI API :8001, served as "glm-5.2" |
| ``` |
| |
| `MixedVQMoEMethod` is auto-selected from the `format:"vq"` markers in `quantization_config`. |
|
|
| ## Performance |
|
|
| - **~16 tok/s** steady-state decode (single stream), **38Γ** over the eager dequant baseline |
| (grouped Triton VQ-GEMM + CUDA graphs; the key win was fixing a shared-memory bank conflict |
| in the codebook gather). |
| - **Context β€ 4096** (dense MLA β sm_120 has no sparse-DSA forward kernel; dense is exact at |
| ctx β€ 2048 and validated functional, incl. >2048 needle retrieval, up to 4096). |
| |
| ## Allocation |
| |
| Mixed **1/2/3-bit per expert** (β 15.8k @1-bit / 34.0k @2-bit / 7.8k @3-bit projections), |
| allocated by activation-frequency-aware AutoBit (arithmetic-routing tuned), then re-encoded |
| to VQ codes + a shared codebook per bit-width. Non-expert "spine" stays scalar 4-bit. |
| |
| ## Limitations |
| |
| - sm_120-specific serving stack (not portable to plain wheels). |
| - Context capped at 4096 on this hardware (sm_120 sparse-attention kernel gap + VRAM); longer |
| context needs smaller weights. |
| - Default serving template is **thinking-on at max reasoning effort** β thorough but verbose |
| for casual chat (pass `chat_template_kwargs={"enable_thinking": false}` for direct answers). |
|
|
| ## License & attribution |
|
|
| - This quantized model: **MIT**. |
| - Base **GLM-5.2**: **MIT**, Β© Zhipu AI β this is a derivative; all rights/attribution to upstream. |
| - Quantization/serving built on **OneCompression** (MIT, Β© Fujitsu Ltd.) and vLLM / transformers (Apache-2.0). |
|
|