miketung's picture
Add files using upload-large-folder tool
c0696d0 verified
|
Raw
History Blame Contribute Delete
3.12 kB

Detailed results

Hardware: 2× RTX PRO 6000 Blackwell Max-Q (sm_120, 96 GB, 300 W cap), PCIe, TP2. A side process held 2.4–2.9 GB per GPU throughout. Raw outputs are in raw/ and Nsight Compute summaries in ncu/.

Dense MXFP8 backend

Server A/B with CUDA graphs and Engram in pinned RAM, before any MoE kernel changes (raw/ab-marlin.txt, raw/ab-humming.txt).

Dense backend Decode, 2K, 1 stream Decode, 2K, 4 streams Prefill, 46K, 1 stream Decode, 46K, 1 stream Prefill, 46K, 4 streams GSM8K-200
FlashInfer W8A8 (vLLM default) 4,026 68.0 97.5%
Marlin W8A16 101.6 210.2 3,804 90.5 7,617 97.0%
Humming 0.1.12 87.2 227.6 3,937 87.6 7,940 97.5%

Kernel profiles in the server

One 33,863-token real-code prompt plus 64 decode tokens under the torch profiler (bench/profile-run.sh, raw/prof-*.txt). Times are summed GPU kernel time.

Prefill, 33.9K tokens Marlin, stock MoE + multi-row prefill + flat scheduler
Total 9,938 ms 8,731 ms 8,719 ms
MoE 5,771 ms 4,455 ms 4,470 ms
Dense GEMM 2,002 ms 2,003 ms 2,038 ms
All-reduce 749 ms 857 ms 783 ms
Sparse attention 555 ms 555 ms 568 ms
Decode, per step Multi-row prefill kernel + flat scheduler
Total 28.9 ms 23.5 ms
MoE 16.1 ms 8.9 ms
Dense GEMM 8.0 ms 8.2 ms
NCCL all-reduce 0.6 ms 1.6 ms

All-reduce time per call grew with the flat scheduler, most likely one rank waiting for the other. Dense GEMMs now cost as much as the MoE per decode step.

Nsight Compute on the stock kernel

kernels/ncu runs came from ncu --set full on single exl3_moe launches (ncu/*.summary.txt).

Stock exl3_moe 6 tokens 4,096 tokens
DRAM throughput 18% 10%
Issue slots busy 51% 56%
Registers per thread 128 128
Achieved occupancy 33% 33%
Integer ALU pipe 31% 34%
Tensor pipe 11% 13%
Top stall, cycles per issued instruction barrier 1.61 wait 1.38

Barrier stalls were 36–38% of stall samples in both cases. In decode, a serial scan over 384 expert counts was the single hottest spot, about 8% of samples.

Things that did not help

Change Effect
Shared-memory pipeline depth 3 → 5, 8, 12 stages At most 4% faster (raw/stage-sweep.txt)
Expert group count 4–23 instead of 23 Between 13% faster and 61% slower depending on batch size (raw/bench-na.txt)
4× multi-row prefill (64 rows per tile decode) Slower than 2× at every size: register spills at the 128-register limit
Wider N tiles on the decode path Deadlocked in split-K

Hybrid dense GEMM, controlled A/B

Both runs use the flat scheduler (raw/ab-flat.txt, raw/ab-flathyb.txt).

Default DENSE_HYBRID_M=128, 6 GiB KV
Prefill, 46K, 1 / 4 streams 4,044 / 8,085 4,260 / 8,503
Decode, 46K, 1 stream 116.6 114.9
DSpark mean acceptance length 3.66 3.59
GSM8K-200 98.5% 95.5%