occamy-1.0-MTP / NATIVE-MTP3.md
Eang's picture
Add validated native MTP3 runtime and H200 results
e563cc5 verified
|
Raw History Blame Contribute Delete
3.78 kB

Native three-step MTP

The released head now has a validated native three-step path on H200 TP1 with BF16 Occamy. It uses a patched SGLang runtime; the head weights have not changed. This is separate from the older MTP1 hook-based path.

Results

  • Two fresh process pairs, each with a cold and warm request: 4/4 exact output tokens and token logprobs.
  • Ten fixed prompts, repeated twice: 20/20 exact output tokens and token logprobs, maximum logprob difference zero.
  • Shared-prefix protection, duplicate-slot release, rejected requests and recovery after abort passed their checks.

Both sides use the same patched runtime, with MTP disabled only for the baseline. CUDA Graphs, overlap scheduling, radix caching, server warmup and BF16 remain enabled. These are bounded greedy-decoding checks, not a proof for arbitrary workloads. The case named long-context has 422 input tokens; it is not a full-context benchmark.

The final cache-boundary fix changes cold-baseline logprobs relative to the previous runtime while preserving this case's output tokens. Exact parity here is against the corrected baseline, not every earlier runtime's outputs.

What changed

The runtime patch and regression tests address three differences:

  1. Retain FP32 sigmoid gates and make GDN state-update rounding consistent between decode and verification.
  2. Fix FA3 KV splitting to one split when CUDA Graphs are enabled.
  3. Publish pending tracked-Mamba prefill cache mappings before scheduling dependent decode or verification. This keeps both paths on the same canonical KV version without overwriting shared cache values.

The scheduler change is limited to the non-DP/non-MLP-sync overlap loop with Mamba extra_buffer and radix caching. The DP/MLP-sync path bypasses this new boundary and is not validated here. The boundary can reduce prefill-to-decode overlap; performance after this final fix has not been measured. Earlier B200 timing observations are not a benchmark of this combined H200 patch.

Apply to an isolated SGLang checkout

Use the SGLang source revision recorded in validation/native3-h200.json. Apply the patch before starting the server or capturing CUDA Graphs. Do not apply it blindly to a different release or an already modified checkout.

# MTP_REPO is this downloaded model repository.
# SGLANG_SRC is an isolated checkout of the recorded SGLang revision.
cd "$SGLANG_SRC"
git apply --check "$MTP_REPO/runtime/native3-h200.patch"
git apply "$MTP_REPO/runtime/native3-h200.patch"

Use an environment with SGLang's matching dependencies installed. Assemble the unchanged published head with the BF16 base:

cd "$MTP_REPO"
MTP_BASE_DIR=/path/to/occamy-bf16
python assemble_head.py --base "$MTP_BASE_DIR" \
  --head ./mtp-trained.safetensors --out ./occamy-with-mtp

SGLANG_ENABLE_SPEC_V2=1 PYTHONPATH="$SGLANG_SRC/python${PYTHONPATH:+:$PYTHONPATH}" \
python -m sglang.launch_server \
  --model-path ./occamy-with-mtp --tokenizer-path "$MTP_BASE_DIR" \
  --host 127.0.0.1 --port 30000 --dtype bfloat16 \
  --context-length 2048 --max-running-requests 1 --max-total-tokens 2048 \
  --max-mamba-cache-size 16 --mem-fraction-static 0.50 \
  --random-seed 42 --mm-attention-backend sdpa \
  --mamba-scheduler-strategy extra_buffer \
  --speculative-algorithm NEXTN --speculative-num-steps 3 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 4

Use greedy requests with thinking disabled for the recorded checks. Do not add the MTP1 canonical_attention hooks to this path. Memory settings depend on hardware. Native MTP3 with FP8, NVFP4, GGUF, sampling, multimodal input, or distributed serving is not established by these H200 BF16 TP1 results.