Update W4A8 numbers to the full-patch configuration
Browse files
README.md
CHANGED
|
@@ -28,15 +28,15 @@ EOS; numbers are median output tok/s per stream.
|
|
| 28 |
|
| 29 |
| | **W4A16** — [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) (native MXFP4) | **W4A8** — this repo |
|
| 30 |
|---|---|---|
|
| 31 |
-
| tok/s per stream, 50 concurrent | 29.6 |
|
| 32 |
-
| tok/s per stream, 20 concurrent | 40.6 |
|
| 33 |
-
| aggregate tok/s, 50 concurrent | 883 |
|
| 34 |
-
| KL to native, nats/token (95% CI) | +0.0011 (0.0006–0.0016) | +0.
|
| 35 |
| AIME 2026 + HMMT Feb 2026, 8 samples/problem | 82.3% | 81.3% (−1.0 pt, SE 1.6: not significant) |
|
| 36 |
|
| 37 |
-
- **Patches in the measured runs:** W4A16 used patches 02 and 04
|
| 38 |
- **Pick W4A16** when you want the reference model.
|
| 39 |
-
- **Pick W4A8** for about
|
| 40 |
- **Scaling out:** two W4A8 replicas (64 GPUs) behind `sglang-router`, 25 streams each, give 35.7 tok/s per stream.
|
| 41 |
- **KL:** teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187.
|
| 42 |
|
|
|
|
| 28 |
|
| 29 |
| | **W4A16** — [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) (native MXFP4) | **W4A8** — this repo |
|
| 30 |
|---|---|---|
|
| 31 |
+
| tok/s per stream, 50 concurrent | 29.6 | 34.1 |
|
| 32 |
+
| tok/s per stream, 20 concurrent | 40.6 | 43.6 |
|
| 33 |
+
| aggregate tok/s, 50 concurrent | 883 | 949 |
|
| 34 |
+
| KL to native, nats/token (95% CI) | +0.0011 (0.0006–0.0016) | +0.0154 (0.0142–0.0167) |
|
| 35 |
| AIME 2026 + HMMT Feb 2026, 8 samples/problem | 82.3% | 81.3% (−1.0 pt, SE 1.6: not significant) |
|
| 36 |
|
| 37 |
+
- **Patches in the measured runs:** W4A16 used patches 02 and 04 (no `--max-total-tokens` cap). W4A8 used 01, 02, 03 and 05, i.e. exactly the launch command below.
|
| 38 |
- **Pick W4A16** when you want the reference model.
|
| 39 |
+
- **Pick W4A8** for about 15% more throughput at 50 streams.
|
| 40 |
- **Scaling out:** two W4A8 replicas (64 GPUs) behind `sglang-router`, 25 streams each, give 35.7 tok/s per stream.
|
| 41 |
- **KL:** teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187.
|
| 42 |
|