Radioheading commited on
Commit
ab2532f
·
verified ·
1 Parent(s): 2671bac

Update W4A8 numbers to the full-patch configuration

Browse files
Files changed (1) hide show
  1. README.md +6 -6
README.md CHANGED
@@ -28,15 +28,15 @@ EOS; numbers are median output tok/s per stream.
28
 
29
  | | **W4A16** — [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) (native MXFP4) | **W4A8** — this repo |
30
  |---|---|---|
31
- | tok/s per stream, 50 concurrent | 29.6 | 31.9 |
32
- | tok/s per stream, 20 concurrent | 40.6 | 40.7 |
33
- | aggregate tok/s, 50 concurrent | 883 | 932 |
34
- | KL to native, nats/token (95% CI) | +0.0011 (0.0006–0.0016) | +0.0151 (0.0138–0.0163) |
35
  | AIME 2026 + HMMT Feb 2026, 8 samples/problem | 82.3% | 81.3% (−1.0 pt, SE 1.6: not significant) |
36
 
37
- - **Patches in the measured runs:** W4A16 used patches 02 and 04. W4A8 used 01 and 02, plus `--max-total-tokens 524288`. Patches 03 and 05 don't change outputs (03 was tested separately: same KL and speed). The speed gain from 05 alone wasn't measured.
38
  - **Pick W4A16** when you want the reference model.
39
- - **Pick W4A8** for about 8% more throughput at 50 streams.
40
  - **Scaling out:** two W4A8 replicas (64 GPUs) behind `sglang-router`, 25 streams each, give 35.7 tok/s per stream.
41
  - **KL:** teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187.
42
 
 
28
 
29
  | | **W4A16** — [moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3) (native MXFP4) | **W4A8** — this repo |
30
  |---|---|---|
31
+ | tok/s per stream, 50 concurrent | 29.6 | 34.1 |
32
+ | tok/s per stream, 20 concurrent | 40.6 | 43.6 |
33
+ | aggregate tok/s, 50 concurrent | 883 | 949 |
34
+ | KL to native, nats/token (95% CI) | +0.0011 (0.0006–0.0016) | +0.0154 (0.0142–0.0167) |
35
  | AIME 2026 + HMMT Feb 2026, 8 samples/problem | 82.3% | 81.3% (−1.0 pt, SE 1.6: not significant) |
36
 
37
+ - **Patches in the measured runs:** W4A16 used patches 02 and 04 (no `--max-total-tokens` cap). W4A8 used 01, 02, 03 and 05, i.e. exactly the launch command below.
38
  - **Pick W4A16** when you want the reference model.
39
+ - **Pick W4A8** for about 15% more throughput at 50 streams.
40
  - **Scaling out:** two W4A8 replicas (64 GPUs) behind `sglang-router`, 25 streams each, give 35.7 tok/s per stream.
41
  - **KL:** teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187.
42