Capicua25x commited on
Commit
e971741
·
verified ·
1 Parent(s): 937f9be

Correct base MTP acceptance: measured 64% @ MTP-3 (was pahajoki's 81% @ MTP-4)

Browse files
Files changed (1) hide show
  1. README.md +3 -2
README.md CHANGED
@@ -79,7 +79,7 @@ Both columns measured on the same hardware/bench (2× R9700, TP2, MTP-3 each):
79
  | Single-stream, 6k prompt | ~82 tok/s | ~85 |
80
  | Aggregate @128, short prompt | **~1875 tok/s** | ~1683 |
81
  | Concurrency ceiling, short prompt | **~128** | ~128 |
82
- | MTP draft acceptance | ~55% (grafted, MTP-3) | 81.2% (native, MTP-4) |
83
 
84
  Effectively **at parity**: the distill edges the base on short-prompt single-stream (107 vs 101)
85
  and high-concurrency aggregate (1875 vs 1683 @128); 6k single-stream is a wash (82 vs 85). Where
@@ -119,7 +119,8 @@ modules to `quantization_config.ignore`** (else vLLM loads them as quantized →
119
  - Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
120
  - Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
121
  - `--language-model-only` required (see Serving).
122
- - Grafted MTP acceptance ~55% vs ~81% native; a native MTP retrain would lift it. MTP is lossless.
 
123
 
124
  ## Credits & acknowledgments
125
 
 
79
  | Single-stream, 6k prompt | ~82 tok/s | ~85 |
80
  | Aggregate @128, short prompt | **~1875 tok/s** | ~1683 |
81
  | Concurrency ceiling, short prompt | **~128** | ~128 |
82
+ | MTP draft acceptance (MTP-3, measured) | ~56% (grafted) | ~64% (native) |
83
 
84
  Effectively **at parity**: the distill edges the base on short-prompt single-stream (107 vs 101)
85
  and high-concurrency aggregate (1875 vs 1683 @128); 6k single-stream is a wash (82 vs 85). Where
 
119
  - Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
120
  - Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
121
  - `--language-model-only` required (see Serving).
122
+ - Grafted MTP acceptance ~56% vs ~64% native (both measured at MTP-3) — only ~8pp behind; a
123
+ native MTP retrain would close it. MTP is lossless either way.
124
 
125
  ## Credits & acknowledgments
126