# Detailed results Hardware: 2× RTX PRO 6000 Blackwell Max-Q (sm_120, 96 GB, 300 W cap), PCIe, TP2. A side process held 2.4–2.9 GB per GPU throughout. Raw outputs are in `raw/` and Nsight Compute summaries in `ncu/`. ## Dense MXFP8 backend Server A/B with CUDA graphs and Engram in pinned RAM, before any MoE kernel changes (`raw/ab-marlin.txt`, `raw/ab-humming.txt`). | Dense backend | Decode, 2K, 1 stream | Decode, 2K, 4 streams | Prefill, 46K, 1 stream | Decode, 46K, 1 stream | Prefill, 46K, 4 streams | GSM8K-200 | |---|---:|---:|---:|---:|---:|---:| | FlashInfer W8A8 (vLLM default) | | | 4,026 | 68.0 | | 97.5% | | Marlin W8A16 | 101.6 | 210.2 | 3,804 | 90.5 | 7,617 | 97.0% | | Humming 0.1.12 | 87.2 | 227.6 | 3,937 | 87.6 | 7,940 | 97.5% | ## Kernel profiles in the server One 33,863-token real-code prompt plus 64 decode tokens under the torch profiler (`bench/profile-run.sh`, `raw/prof-*.txt`). Times are summed GPU kernel time. | Prefill, 33.9K tokens | Marlin, stock MoE | + multi-row prefill | + flat scheduler | |---|---:|---:|---:| | Total | 9,938 ms | 8,731 ms | 8,719 ms | | MoE | 5,771 ms | 4,455 ms | 4,470 ms | | Dense GEMM | 2,002 ms | 2,003 ms | 2,038 ms | | All-reduce | 749 ms | 857 ms | 783 ms | | Sparse attention | 555 ms | 555 ms | 568 ms | | Decode, per step | Multi-row prefill kernel | + flat scheduler | |---|---:|---:| | Total | 28.9 ms | 23.5 ms | | MoE | 16.1 ms | 8.9 ms | | Dense GEMM | 8.0 ms | 8.2 ms | | NCCL all-reduce | 0.6 ms | 1.6 ms | All-reduce time per call grew with the flat scheduler, most likely one rank waiting for the other. Dense GEMMs now cost as much as the MoE per decode step. ## Nsight Compute on the stock kernel `kernels/ncu` runs came from `ncu --set full` on single `exl3_moe` launches (`ncu/*.summary.txt`). | Stock `exl3_moe` | 6 tokens | 4,096 tokens | |---|---:|---:| | DRAM throughput | 18% | 10% | | Issue slots busy | 51% | 56% | | Registers per thread | 128 | 128 | | Achieved occupancy | 33% | 33% | | Integer ALU pipe | 31% | 34% | | Tensor pipe | 11% | 13% | | Top stall, cycles per issued instruction | barrier 1.61 | wait 1.38 | Barrier stalls were 36–38% of stall samples in both cases. In decode, a serial scan over 384 expert counts was the single hottest spot, about 8% of samples. ## Things that did not help | Change | Effect | |---|---| | Shared-memory pipeline depth 3 → 5, 8, 12 stages | At most 4% faster (`raw/stage-sweep.txt`) | | Expert group count 4–23 instead of 23 | Between 13% faster and 61% slower depending on batch size (`raw/bench-na.txt`) | | 4× multi-row prefill (64 rows per tile decode) | Slower than 2× at every size: register spills at the 128-register limit | | Wider N tiles on the decode path | Deadlocked in split-K | ## Hybrid dense GEMM, controlled A/B Both runs use the flat scheduler (`raw/ab-flat.txt`, `raw/ab-flathyb.txt`). | | Default | `DENSE_HYBRID_M=128`, 6 GiB KV | |---|---:|---:| | Prefill, 46K, 1 / 4 streams | 4,044 / 8,085 | 4,260 / 8,503 | | Decode, 46K, 1 stream | 116.6 | 114.9 | | DSpark mean acceptance length | 3.66 | 3.59 | | GSM8K-200 | 98.5% | 95.5% |