| --- |
| license: mit |
| base_model: zai-org/GLM-5.2 |
| pipeline_tag: text-generation |
| tags: |
| - moe |
| - reap |
| - pruning |
| - router-kd |
| - knowledge |
| - nvfp4 |
| - glm |
| - glm-5.2 |
| --- |
| |
| # GLM-5.2-504B-K β knowledge-augmented REAP keep-168 (full-data Router-KD, NVFP4) |
|
|
| The **"K-cut"** sibling of [`0xSero/GLM-5.2-504B`](https://huggingface.co/0xSero/GLM-5.2-504B): the |
| same ~504B / keep-168 budget, but the expert selection is **biased toward knowledge & reasoning** β |
| the winning top-160 core (kept bit-for-bit) **plus the 8 highest-priority knowledge-exclusive experts |
| per layer** that coding-saliency pruning drops. Recovered with gate-only **Router-KD trained on the |
| FULL calibration set (~18.6k real traces)** β 6x the data of the first-pass cuts. |
|
|
| ## Sponsor |
| 8x NVIDIA B200 sponsored by [Lambda](https://lambda.ai). Thank you. |
|
|
| ## Why this variant exists |
| REAP saliency computed from coding traces under-weights experts that fire mainly on |
| reasoning/knowledge. The K-cut deliberately re-includes them β trading a sliver of coding-saliency |
| coverage for broader knowledge coverage. Reach for this on knowledge/reasoning-heavy workloads; use |
| the plain [`GLM-5.2-504B`](https://huggingface.co/0xSero/GLM-5.2-504B) otherwise. |
|
|
| ## Eval (n=2000 held-out real prompts, raw, no max_tokens / no timeout) |
| | metric | GLM-5.2-504B-K (this) | GLM-5.2-504B (plain floor) | |
| |---|---|---| |
| | attractor / loop rate | **0.078** | 0.072 | |
| | natural-EOS rate | 0.923 | 0.928 | |
| | distinct-4 | 0.881 | 0.880 | |
| | median tokens | 1232 | 1267 | |
| |
| On the loop metric this is **at parity with (or better than) the plain-cut floor**. The residual loops are inherent to GLM-5.2 itself (the unpruned teacher |
| loops on the same prompts), so neither cut "fixes" them β this one buys knowledge-expert coverage. |
| |
| ## Serving (vLLM) |
| ```bash |
| vllm serve 0xSero/GLM-5.2-504B-K --tensor-parallel-size 8 \ |
| --quantization modelopt_fp4 --kv-cache-dtype fp8 --trust-remote-code --max-model-len 262144 |
| ``` |
| |
| --- |
| *REAP knowledge-augmented cut + full-data Router-KD. Compute sponsored by [Lambda](https://lambda.ai).* |
| |