YSLAB-ai's picture
Correct and publish full MTP sweep
0bae7d5 verified
|
Raw
History Blame Contribute Delete
4.44 kB

Benchmarks

Measured results - Orcarouter checkpoint only

Every runtime number here applies to the pinned orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 revision 3a3b63161c0745390e5270179af42e46efc70799 on one DGX Spark. The MTP results also use the exact 31 BF16 MTP tensors from pinned RadixArk/Qwen3.8-Flash-Next-NVFP4 revision 7b719225242aacd3dbd3f9407468c2ee9a9d2594. Inferact and RadixArk are not runtime-qualified by these results.

Single-stream BF16 MTP depth sweep

The load guard verified 31/31 tensors before accepting requests. Each depth used a 32,768-token server profile, concurrency=1, one maximum sequence, 0.80 GPU-memory utilization, one 128-token warm-up, then three fixed 256-token streamed samples. The primary rate divides all completion tokens—including hidden reasoning—by total request wall time, so it also includes latency before visible content. The next column is the median time from request start to the first visible content delta. Thinking mode used medium reasoning effort, temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, and repetition_penalty=1.0.

MTP depth Median end-to-end completion tok/s Time to first visible content Aggregate acceptance
0 27.26 1.196s n/a
1 33.17 1.206s 305/463 (65.9%)
2 44.23 1.324s 532/722 (73.7%)
3 40.95 1.629s 563/993 (56.7%)
4 41.84 1.123s 611/1,148 (53.2%)
5 35.45 1.249s 454/1,560 (29.1%)
6 35.71 1.429s 516/1,230 (42.0%)
8 30.43 1.758s 487/1,600 (30.4%)
10 28.31 1.648s 642/2,150 (29.9%)

MTP=2 is the qualified selection. It was 62.2% faster end-to-end than MTP=0 in this single-stream short-request measurement. For MTP2, retained reasoning-token accounting also gives a separate median visible-answer phase rate of 46.15 tok/s, measured from the first through last visible content token. That rate excludes hidden reasoning and must not be combined with total completion-token counts.

An earlier revision did combine all completion tokens with a visible-content-only time interval. That mixed metric has been superseded and is not comparable to either rate above. Higher draft depths lost enough acceptance to cost more verification work than they saved. Depths five and above require the recipe's 48-token block alignment for the QSA speculative ring.

The machine-readable record is mtp-bf16-sweep.json.

Native-context validation with MTP=2

Measurement Qualified observation
vLLM build 0.1.dev20073+g8e685d198
Model loading 76.21 GiB
BF16 PLE on NVMe 95.37 GiB
BF16 MTP overlay 4.86 GiB; 31/31 tensors
BF16 KV available 17.37 GiB
KV token capacity 627,960 tokens
Native-window concurrency estimate 2.40 at 262,144 tokens
Exact retrieval pass at 240,051 prompt tokens
Retrieval TTFT / prefill 116.80s / 2,055.17 prompt tok/s

The full-profile model-load phase took 576.3 seconds. Qualified API readiness is longer because runtime initialization follows weight loading; allow about 15 minutes for a cold-start timeout rather than treating model-load time as readiness time.

Baseline concurrency without the overlay

These older results use the original Orcarouter checkpoint at MTP=0, an 8K request profile, and report whole-wave completion throughput rather than per-request decode rates.

Streams Whole-wave completion throughput
1 13.06 completion tok/s
2 17.76 completion tok/s
4 19.69 completion tok/s
8 21.66 completion tok/s

They are retained as baseline evidence. No claim is made that this is an MTP=2 concurrency sweep.

Functional, vision, and stability evidence

The baseline qualification includes factual, code, JSON, tool, and vision checks. The vision fixture contained three red squares and two blue circles and passed. The MTP overlay additionally passed text generation, load-guard, acceptance, and exact long-context retrieval checks.

The earlier file named stability-2h.json records only 904 seconds and is not a completed two-hour run. Separately, summary.json records 8,159 seconds of uninterrupted uptime at a gate, zero restarts, and eleven correct health probes. These remain separate evidence and are not combined into a two-hour claim.