YSLAB-ai's picture
Correct and publish full MTP sweep
0bae7d5 verified
|
Raw
History Blame Contribute Delete
4.44 kB
# Benchmarks
## Measured results - Orcarouter checkpoint only
Every runtime number here applies to the pinned
`orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4` revision
`3a3b63161c0745390e5270179af42e46efc70799` on one DGX Spark. The MTP results
also use the exact 31 BF16 MTP tensors from pinned
`RadixArk/Qwen3.8-Flash-Next-NVFP4` revision
`7b719225242aacd3dbd3f9407468c2ee9a9d2594`. Inferact and RadixArk are not
runtime-qualified by these results.
## Single-stream BF16 MTP depth sweep
The load guard verified 31/31 tensors before accepting requests. Each depth used a
32,768-token server profile, `concurrency=1`, one maximum sequence, 0.80 GPU-memory
utilization, one 128-token warm-up, then three fixed 256-token streamed samples.
The primary rate divides all completion tokens—including hidden reasoning—by total
request wall time, so it also includes latency before visible content. The next
column is the median time from request start to the first visible content delta.
Thinking mode used medium reasoning effort,
`temperature=1.0`, `top_p=0.95`, `top_k=20`, `min_p=0.0`,
`presence_penalty=0.0`, and `repetition_penalty=1.0`.
| MTP depth | Median end-to-end completion tok/s | Time to first visible content | Aggregate acceptance |
| ---: | ---: | ---: | ---: |
| 0 | 27.26 | 1.196s | n/a |
| 1 | 33.17 | 1.206s | 305/463 (65.9%) |
| **2** | **44.23** | **1.324s** | **532/722 (73.7%)** |
| 3 | 40.95 | 1.629s | 563/993 (56.7%) |
| 4 | 41.84 | 1.123s | 611/1,148 (53.2%) |
| 5 | 35.45 | 1.249s | 454/1,560 (29.1%) |
| 6 | 35.71 | 1.429s | 516/1,230 (42.0%) |
| 8 | 30.43 | 1.758s | 487/1,600 (30.4%) |
| 10 | 28.31 | 1.648s | 642/2,150 (29.9%) |
`MTP=2` is the qualified selection. It was 62.2% faster end-to-end than `MTP=0`
in this single-stream short-request measurement. For MTP2, retained reasoning-token
accounting also gives a separate median visible-answer phase rate of 46.15 tok/s,
measured from the first through last visible content token. That rate excludes hidden
reasoning and must not be combined with total completion-token counts.
An earlier revision did combine all completion tokens with a visible-content-only
time interval. That mixed metric has been superseded and is not comparable to either
rate above. Higher draft depths lost enough acceptance to cost more verification
work than they saved. Depths five and above require the recipe's 48-token block
alignment for the QSA speculative ring.
The machine-readable record is
[`mtp-bf16-sweep.json`](../results/orcarouter/mtp-bf16-sweep.json).
## Native-context validation with MTP=2
| Measurement | Qualified observation |
| --- | --- |
| vLLM build | `0.1.dev20073+g8e685d198` |
| Model loading | 76.21 GiB |
| BF16 PLE on NVMe | 95.37 GiB |
| BF16 MTP overlay | 4.86 GiB; 31/31 tensors |
| BF16 KV available | 17.37 GiB |
| KV token capacity | 627,960 tokens |
| Native-window concurrency estimate | 2.40 at 262,144 tokens |
| Exact retrieval | pass at 240,051 prompt tokens |
| Retrieval TTFT / prefill | 116.80s / 2,055.17 prompt tok/s |
The full-profile model-load phase took 576.3 seconds. Qualified API readiness is
longer because runtime initialization follows weight loading; allow about 15 minutes
for a cold-start timeout rather than treating model-load time as readiness time.
## Baseline concurrency without the overlay
These older results use the original Orcarouter checkpoint at `MTP=0`, an 8K
request profile, and report whole-wave completion throughput rather than per-request
decode rates.
| Streams | Whole-wave completion throughput |
| ---: | ---: |
| 1 | 13.06 completion tok/s |
| 2 | 17.76 completion tok/s |
| 4 | 19.69 completion tok/s |
| 8 | 21.66 completion tok/s |
They are retained as baseline evidence. No claim is made that this is an MTP=2
concurrency sweep.
## Functional, vision, and stability evidence
The baseline qualification includes factual, code, JSON, tool, and vision checks.
The vision fixture contained three red squares and two blue circles and passed. The
MTP overlay additionally passed text generation, load-guard, acceptance, and exact
long-context retrieval checks.
The earlier file named `stability-2h.json` records only 904 seconds and is not a
completed two-hour run. Separately, `summary.json` records 8,159 seconds of
uninterrupted uptime at a gate, zero restarts, and eleven correct health probes.
These remain separate evidence and are not combined into a two-hour claim.