| --- |
| license: mit |
| base_model: |
| - zai-org/GLM-5.2 |
| - lukealonso/GLM-5.2-NVFP4 |
| base_model_relation: quantized |
| tags: |
| - glm |
| - moe |
| - nvfp4 |
| - nf3 |
| - mxfp8 |
| - blackwell |
| - quantization |
| --- |
| |
| # GLM-5.2-MXFP8-NVFP4-NF3-Hybrid |
|
|
| > Ready-to-run v20 Compose profiles are included for four 96 GiB SM120 GPUs, |
| > with validated context lengths from 400K through 775K using FP8 or calibrated |
| > NVFP4 MLA KV. Use the pinned v20 image and included launch files because the |
| > checkpoint's custom NF3 expert tensors require its native runtime support. |
|
|
| The full 753B-parameter GLM-5.2, with all 256 routed experts per MoE layer, |
| compressed into 340.9 GiB (366,021,469,560 bytes) without expert pruning. |
| Sixty-four damage-selected expert slots per layer retain Luke Alonso's NVFP4 |
| tensors; the other 192 use the healed NF3 format and its SM120 Grid188 kernel. |
| Attention, dense, and shared-expert tensors remain BF16 in the checkpoint. |
| Eligible linears are converted once to MXFP8 at load according to the selected |
| serving profile. |
|
|
| The weight tensors are v3.6 at immutable weight revision |
| `68babde27a97a4c980c2494e830dd424975cd5a3`. They contain 184 safetensor shards |
| of at most 2 GB. The reviewed release bundle adds only documentation, Compose |
| files, the launcher shim, and the calibrated KV-scale sidecar; it does not |
| rewrite those weight shards. |
|
|
| ## Four v20 serving profiles |
|
|
| | Profile | MTP | Public max-model-len | Shared-ruler KLD ↓ | Launch | |
| |---|---:|---:|---:|---| |
| | FP8 MLA KV | 0 | 400,000 | **0.12715 ± 0.00183** | `docker-compose.fp8.yml` | |
| | FP8 MLA KV, shared-expert MXFP8 | 3 | **450,000** | **0.12821 ± 0.00306** | base FP8 + `docker-compose.fp8-450k-mtp3.yml` | |
| | Calibrated NVFP4 KV, BF16 RoPE | 0 | 700,000 | **0.13234 ± 0.00226** | `docker-compose.nvfp4.yml` | |
| | Calibrated NVFP4 KV, FP8 RoPE | 3 | **775,000** | **0.13386 ± 0.00544** | base NVFP4 + `docker-compose.nvfp4-775k-mtp3.yml` | |
|
|
| All profiles use TP4/DCP4, A16 expert activations, lossless BF16 DCP transport, |
| and InstantTensor `BUFFERED`. The FP8 profile uses the accepted v19 |
| online-MXFP8 membership. The two NVFP4 capacity profiles additionally convert |
| shared experts online; their KLD is therefore certified separately rather than |
| borrowed from the FP8 profile. |
|
|
| The final 775k profile allocated exactly 775,000 KV tokens and completed a |
| token-counted 764,999-token prompt plus 16 generated tokens. It then completed |
| a fresh request returning `POST CAP READY`, with zero engine restarts. It is a |
| tight four-GPU pin and uses a |
| small launcher shim plus two source patches applied only inside the ephemeral |
| container. Read the KV-token count at startup and lower `MAX_MODEL_LEN` if |
| your machine has less runtime headroom. |
|
|
| An exact-stack ceiling canary also completed a 774,999-token prompt against a |
| 785,000-token pool and passed post-cap reuse, but its measured free-memory |
| floor was only 94 MiB per GPU. The portable release profile therefore remains |
| 775,000 rather than advertising the brittle ceiling. |
|
|
| The FP8-KV/MTP3 overlay allocated 450,255 KV tokens for a 450,000-token |
| advertised limit. It completed a 449,935-token prompt with coherent generation |
| and then a fresh 8k request, with zero restarts/OOMs and a 146 MiB measured |
| free-memory floor. The explicit KV-byte pin avoids an automatically sized |
| 467,627-token pool that booted but left only 6 MiB after lazy allocation. |
|
|
| ## Accuracy |
|
|
| GPQA-Diamond used temperature 1.0, top-p 0.95, and maximum reasoning effort. |
| The v3.6 row is this exact checkpoint. |
|
|
| | Precision | GPQA Diamond | |
| |---|---:| |
| | Official FP8 reference | 89.52 | |
| | Full NVFP4 reference | 89.39 | |
| | **This model, v3.6** | **88.89 (176/198)** | |
| | Previous hybrid revision | 88.38 (175/198) | |
| | REAP-594B prune, for contrast | 86.87 | |
|
|
| KLD is a distribution-drift instrument, not an intelligence score. The table |
| above remains the direct quality check. |
|
|
| ## KLD method |
|
|
| The reported KLD uses the current shared GLM-5.2 ruler: |
|
|
| - source: `zai-org/GLM-5.2` BF16 reference logits from 2026-07-08; |
| - one fixed 2,048-token Wikitext window, 2,047 scored positions; |
| - full 154,880-token vocabulary; |
| - direction `KL(BF16 reference || candidate)`; |
| - TP4/DCP4, MTP0, A16, five fresh engine starts; |
| - exact 78-character GLM IndexCache pattern. |
|
|
| The final v20 review reruns each precision posture five times with a fresh |
| engine and an isolated empty cache. KLD is target-only MTP0 even when the |
| corresponding serving profile enables MTP3. The scale file is an audited, |
| per-layer outer-scale artifact; it changes no checkpoint weights and adds no |
| KV bytes. |
|
|
| | Profile | Runs | KLD mean ± sample SD | Min | Max | |
| |---|---:|---:|---:|---:| |
| | FP8 MLA KV | 5 | 0.12715 ± 0.00183 | 0.12564 | 0.12958 | |
| | **FP8 MLA KV, shared-expert MXFP8, direct MXFP8 BMM** | **5** | **0.12821 ± 0.00306** | **0.12318** | **0.13055** | |
| | Calibrated NVFP4 KV, BF16 RoPE | 5 | 0.13234 ± 0.00226 | 0.12888 | 0.13484 | |
| | Calibrated NVFP4 KV, FP8 RoPE, direct BMM off | 5 | 0.13630 ± 0.00641 | 0.13052 | 0.14671 | |
| | **Calibrated NVFP4 KV, FP8 RoPE, direct MXFP8 BMM** | **5** | **0.13386 ± 0.00544** | **0.12769** | **0.14033** | |
|
|
| The machine-readable aggregate is included as `KLD_MATRIX.json`. |
| The published scale sidecar uses host-neutral provenance labels; its |
| `max_abs` and `scales` arrays are byte-for-byte identical to the artifact used |
| for every KLD receipt. `KLD_MATRIX.json` records both the receipt-artifact hash |
| and a canonical numeric-payload hash. |
|
|
| Do not compare these values with the old card's FP8-as-reference, top-8192 |
| drift table. That was a different reference, support, and direction. |
|
|
| ## Serve from the immutable v20 image |
|
|
| Exact image: |
|
|
| ```text |
| voipmonitor/vllm:gilded-gnosis-v20-vllm0c79e41-sie603f74-fi801d57a-cu132-20260726 |
| sha256:10261c7d65101c8aba2ce1fb59eabe73aff9d35eca5043b330cc0ce76d3c98d0 |
| ``` |
|
|
| A matched four-GPU A/B retained this July-26 image over the July-25 final image: |
| the selected image measured 2,291/2,247 prompt tok/s at 64k/128k versus |
| 2,264/2,217, and it contains the scale-file and launcher interfaces used by |
| the reviewed profiles. A later policy A/B found that the July-26 helper's |
| opt-in owner-top-k merge was the remaining prefill regression on this |
| topology. The final 775k overlay disables that experimental route and retains |
| the established replicated global-top-k oracle. |
|
|
| Download without Xet: |
|
|
| ```bash |
| export HF_HUB_DISABLE_XET=1 |
| hf download madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid \ |
| --local-dir ./GLM-5.2-MXFP8-NVFP4-NF3-Hybrid |
| cd ./GLM-5.2-MXFP8-NVFP4-NF3-Hybrid |
| export MODEL_DIR=$PWD |
| ``` |
|
|
| Do not pin the old weight revision for this serving recipe: that historical |
| commit predates the calibrated scale sidecar. After publication, pin the final |
| release-bundle commit if an immutable metadata checkout is required. |
|
|
| Choose one of the standalone Compose files included in this repository: |
|
|
| ```bash |
| # Roughly 400k context, lowest measured KLD. |
| docker compose -f docker-compose.fp8.yml up -d |
| |
| # 450k/MTP3, FP8 KV with shared experts converted online to MXFP8. |
| docker compose \ |
| -f docker-compose.fp8.yml \ |
| -f docker-compose.fp8-450k-mtp3.yml up -d |
| |
| # 700k quality-first long-context pin, calibrated NVFP4 KV and BF16 RoPE. |
| docker compose -f docker-compose.nvfp4.yml up -d |
| |
| # 775k/MTP3 capacity profile, calibrated NVFP4 KV and FP8 RoPE. |
| docker compose \ |
| -f docker-compose.nvfp4.yml \ |
| -f docker-compose.nvfp4-775k-mtp3.yml up -d |
| ``` |
|
|
| Wait for health before sending traffic: |
|
|
| ```bash |
| curl -fsS http://127.0.0.1:5001/health |
| curl -fsS http://127.0.0.1:5001/v1/models |
| ``` |
|
|
| The two base files default to MTP0 so their advertised capacity is |
| reproducible. The 450k and 775k overlays enable MTP3. Keep MTP disabled for |
| target-only decode measurements or KLD work: |
|
|
| ```bash |
| MTP=0 docker compose -f docker-compose.fp8.yml up -d --force-recreate |
| ``` |
|
|
| The v20 helper owns DCP query split, CKV gather, bounded prefetch, |
| InstantTensor, NCCL, graph sizing, sparse attention, and parsers for the two |
| base profiles. Leave those controls on auto there. Both MTP3 overlays pin the |
| measured validation policy to avoid a fresh-cache calibration cleanup defect |
| and to bound their CKV workspace. The 775k NVFP4 overlay additionally pins |
| both `DCP_TOPK_OWNER_MERGE=0` and `VLLM_DCP_TOPK_OWNER_MERGE=0`: the |
| July-26 auto-policy otherwise enables an explicitly opt-in owner merge that |
| regressed 64k/128k prompt throughput by about 19% on the measured topology. |
| The overlays also apply the included MTP online-quant inheritance and |
| sparse-indexer carry-fold patches before launch. No checkpoint tensor is |
| changed. `F8_DMA=0` is intentional: it keeps DCP transport lossless. |
|
|
| ### Why the Compose uses `ONLINE_QUANT=custom` |
| |
| The image's convenient `nf3-mxfp8` shorthand is a faster, more aggressive |
| membership that also converts shared experts and fused q-a/kv-a projections. |
| The FP8 profile uses this measured quality posture: |
| |
| ```json |
| {"linear":{"weight":"mxfp8"},"ignore":["re:.*\\.fused_qkv_a_proj$","re:.*\\.q_a_proj$","re:.*kv_a_proj_with_mqa","re:.*\\.mlp\\.gate$","model.layers.78.eh_proj","lm_head"]} |
| ``` |
| |
| The NVFP4 profiles use the same ignore list plus |
| `"shared_experts":{"weight":"mxfp8"}` to recover the capacity required by their |
| public pins. Their five-run KLD rows measure that exact membership. |
| |
| The NVFP4 Compose also passes |
| `VLLM_NVFP4_MLA_SCALES_FILE` explicitly. The final v20 image contains scale |
| support, but checkpoint-relative auto-discovery is not part of this release. |
| The generic vLLM environment scanner may still label this extension variable |
| unknown; the stock v20 MLA writer reads it directly. Runtime verification |
| observed non-identity `latent_scale_identity=0` kernel signatures, proving that |
| the calibrated values—not the identity fallback—reached the attention path. |
| |
| ## v20 performance |
| |
| Final review measurements on four RTX PRO 6000 Blackwell Max-Q GPUs: |
| |
| | KV profile | MTP | Max len | C1 decode, ctx 0 | C1 decode, 64k | C1 decode, 128k | Prefill 64k | Prefill 128k | Reported KV pool | |
| |---|---:|---:|---:|---:|---:|---:|---:|---:| |
| | FP8 | 0 | 400k | 48.0 | 46.9 | 47.3 | 2,069 | 2,073 | 435,316 | |
| | FP8 | 3 | 180k | 89.7 | 86.6 | 88.7 | — | — | 218,096 | |
| | FP8, shared-expert MXFP8 | 3 | **450k** | — | — | — | **2,172** | **2,060** | **450,255** | |
| | NVFP4, BF16 RoPE | 0 | 700k | 48.8 | 47.8 | 48.3 | 1,994 | 2,027 | 733,272 | |
| | NVFP4, BF16 RoPE | 3 | 360k | 100.4 | 100.8 | 97.9 | — | — | 441,364 | |
| | NVFP4, FP8 RoPE | 0 | 775k | — | — | — | — | — | 842,057 | |
| | **NVFP4, FP8 RoPE** | **3** | **775k** | **106.1** | **106.1** | **102.2** | **2,548** | **2,470** | **775,000** | |
| |
| The final 775k/MTP3 run measured a five-run coding peak of 121.6 tok/s mean / |
| 124.3 tok/s maximum. The final prefill row is the second complete exact-token |
| 64k/128k pass after warmup. The bounded carry-fold measured within about one |
| percent of the unpatched owner-merge-off profile, while avoiding the stock |
| fold's 112 MiB near-capacity allocation and engine OOM. |
| |
| The 450k FP8 row combines the capacity certification with a separate warmed, |
| exact-token 64k/128k prefill measurement. The older 180k row retains the |
| matched decode measurements. No unmeasured decode number is borrowed for the |
| new shared-expert precision posture. |
| |
| `VLLM_B12X_ABSORB_BMM=1` is active in the certified 775k posture. Load ordering |
| matters: after the 184-shard target load, v20 logged that it was serving MLA |
| absorbed projections directly from the B12X MXFP8 pack. The same-checkpoint |
| MTP inheritance patch gives layer 78 the identical online-MXFP8 membership, so |
| the four-shard draft also uses the direct pack and no post-draft fallback is |
| logged. |
| |
| The second patch selects SparkInfer's existing streaming carry fold for |
| logical sparse-indexer output. Exact CUDA tests produced identical selected |
| token sets at 131k and 524k. This removes page-table-sized transient candidate |
| tensors that otherwise caused late OOMs even after a successful 775k boot. |
| These are explicit runtime patches over the immutable image; this profile is |
| not described as unmodified stock v20. |
| |
| ### High-concurrency status |
| |
| The table above intentionally reports C1 decode only. The July-26 v20 NF3 |
| hybrid kernel has a known batch-dispatch crossover under investigation: |
| Grid188 covers the smallest decode bucket and the optimized direct-top-k |
| launch is currently bounded to eight routed rows; larger MTP verifier batches |
| fall back to the packed route used for prefill. This does not affect |
| correctness, KLD, advertised KV capacity, C1 decode, or the prefill results |
| reported here, but it limits C2-C8 scaling. A lossless native dispatch fix is |
| being prepared for an upstream runtime release; no unvalidated concurrency |
| number or private kernel patch is included in this checkpoint release. |
| |
| ## Credits |
| |
| - Base model: [zai-org/GLM-5.2](https://huggingface.co/zai-org/GLM-5.2) |
| - NVFP4 tier: [lukealonso/GLM-5.2-NVFP4](https://huggingface.co/lukealonso/GLM-5.2-NVFP4) |
| - v20 runtime and documentation: |
| [local-inference-lab/rtx6kpro](https://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm5.2_v20.md) |
| - Shared KLD ruler: |
| [GLM-5.2 KLD Evaluation](https://github.com/local-inference-lab/rtx6kpro/blob/master/benchmarks/glm52-kld-evaluation.md) |
| |
| Built and evaluated by madeby561. |
| |