# NVFP4 + MTP on vLLM: community setup [miter37's DGX Spark recipe](https://github.com/miter37/occamy-1.0-nvfp4-mtp2-recipe) documents a working single-request NVFP4 + MTP2 setup. Thanks for identifying the quantization configuration gap and sharing the setup and logs. This is a community deployment reference, not an official vLLM output-parity or speedup result. The author's measurements use different context lengths and generation settings; they are not a controlled A/B. Sampling produced mixed-script continuations in some runs, with the cause unresolved. ## Assembly Use `assemble_head.py` with the NVFP4 base and the BF16 head. It now writes an independent `hf_quant_config.json` and excludes `mtp*` from NVFP4 in both quantization configurations. Base files and weights remain unchanged. Reassemble into a new output directory if you used an older helper; do not edit through an existing configuration symlink. ## Backend settings The community recipe uses vLLM 0.29 with Marlin for the NVFP4 target and Triton for the BF16 drafter: ```bash vllm serve /models/occamy-with-mtp \ --moe-backend marlin \ --speculative-config '{"method":"mtp","num_speculative_tokens":2,"moe_backend":"triton"}' \ --max-num-seqs 1 --max-model-len 2048 --enforce-eager ``` This is a starting configuration, not a hardware-independent memory prescription. Start with MTP1 before increasing draft depth or context. The SGLang hooks shipped here do not attach to vLLM; omitting them does not establish equivalent correctness. ## Containers The assembly helper links weight files to their resolved absolute paths. Mount the base, head and assembled directories at the same absolute paths inside the container. Mounting only the assembled directory leaves broken weight links. Check before launching: ```bash find -L /models/occamy-with-mtp -maxdepth 1 -type l -print ``` No output means no broken top-level links were found. ## Before relying on it Compare MTP off/on with identical prompts, tokenizer, precision and greedy settings, including cold and warm requests. Check output token sequences before measuring speed. Then measure matched request lengths and cache states; report acceptance separately from end-to-end latency. Concurrency and sampling need their own checks. The existing SGLang validation does not cover these vLLM combinations.