occamy-1.0-MTP / VLLM-COMMUNITY.md
Eang's picture
Preserve BF16 MTP head in NVFP4 assembly and document community vLLM setup
2fd68d3 verified
|
Raw History Blame Contribute Delete
2.35 kB

NVFP4 + MTP on vLLM: community setup

miter37's DGX Spark recipe documents a working single-request NVFP4 + MTP2 setup. Thanks for identifying the quantization configuration gap and sharing the setup and logs.

This is a community deployment reference, not an official vLLM output-parity or speedup result. The author's measurements use different context lengths and generation settings; they are not a controlled A/B. Sampling produced mixed-script continuations in some runs, with the cause unresolved.

Assembly

Use assemble_head.py with the NVFP4 base and the BF16 head. It now writes an independent hf_quant_config.json and excludes mtp* from NVFP4 in both quantization configurations. Base files and weights remain unchanged. Reassemble into a new output directory if you used an older helper; do not edit through an existing configuration symlink.

Backend settings

The community recipe uses vLLM 0.29 with Marlin for the NVFP4 target and Triton for the BF16 drafter:

vllm serve /models/occamy-with-mtp \
  --moe-backend marlin \
  --speculative-config '{"method":"mtp","num_speculative_tokens":2,"moe_backend":"triton"}' \
  --max-num-seqs 1 --max-model-len 2048 --enforce-eager

This is a starting configuration, not a hardware-independent memory prescription. Start with MTP1 before increasing draft depth or context. The SGLang hooks shipped here do not attach to vLLM; omitting them does not establish equivalent correctness.

Containers

The assembly helper links weight files to their resolved absolute paths. Mount the base, head and assembled directories at the same absolute paths inside the container. Mounting only the assembled directory leaves broken weight links. Check before launching:

find -L /models/occamy-with-mtp -maxdepth 1 -type l -print

No output means no broken top-level links were found.

Before relying on it

Compare MTP off/on with identical prompts, tokenizer, precision and greedy settings, including cold and warm requests. Check output token sequences before measuring speed. Then measure matched request lengths and cache states; report acceptance separately from end-to-end latency. Concurrency and sampling need their own checks. The existing SGLang validation does not cover these vLLM combinations.