Warn: known KL regression, re-encode pending
Browse files
README.md
CHANGED
|
@@ -9,6 +9,9 @@ tags:
|
|
| 9 |
- exl3
|
| 10 |
---
|
| 11 |
|
|
|
|
|
|
|
|
|
|
| 12 |
> [!TIP]
|
| 13 |
> **[Support this work →](https://donate.sybilsolutions.ai)** · [X](https://x.com/0xsero) · [GitHub](https://github.com/0xsero) · [REAP paper](https://arxiv.org/abs/2510.13999) · [Cerebras REAP](https://huggingface.co/collections/cerebras/cerebras-reap)
|
| 14 |
|
|
|
|
| 9 |
- exl3
|
| 10 |
---
|
| 11 |
|
| 12 |
+
> [!WARNING]
|
| 13 |
+
> **Known KL regression — do not use this checkpoint yet.** Measured KL vs BF16 is **~1.02 nats** (reproduced: 1.0223 then 1.0215; median token KL 0.33), ~10× worse than expected for an experts-only + fp16-sensitive build. The regression is confirmed real (verified-complete checkpoint, measured twice), and traced to a crash-and-resume finalization during the encode. A clean full re-encode is pending; until then use the [3bpw ladder](https://huggingface.co/collections/0xSero/glm-53-reap-exl3-quantization-suite-6a9f03389fe6a96671619ba3) or the [W4A16 cuts](https://huggingface.co/collections/0xSero/glm-53-reap-w4a16-hopper-6a9f128893dbd3dbf0a2fce7).
|
| 14 |
+
|
| 15 |
> [!TIP]
|
| 16 |
> **[Support this work →](https://donate.sybilsolutions.ai)** · [X](https://x.com/0xsero) · [GitHub](https://github.com/0xsero) · [REAP paper](https://arxiv.org/abs/2510.13999) · [Cerebras REAP](https://huggingface.co/collections/cerebras/cerebras-reap)
|
| 17 |
|