MTP head fails to load in vLLM: mtp.* linears missing from quantization_config.ignore

#2
by greglechin - opened

Thanks for putting this up — the recipe note pointing at dbirks was useful.

One issue: the model can't be served with speculative decoding enabled. vLLM dies at
worker start:

ValueError: There is no module or parameter named 'fc.weight' in Qwen3_5MultiTokenPredictor.
The available parameters belonging to fc (ColumnParallelLinear) are:
{'fc.weight_packed', 'fc.weight_shape', 'fc.weight_scale'}

(vllm/model_executor/models/qwen3_5_mtp.py → load_weights, via
spec_decode/mtp/speculator.py::load_draft_model)

Cause

The 15 mtp.* tensors in model_extra_tensors.safetensors are BF16 — correct, and
what you want:

mtp.fc.weight                                  BF16 [5120, 10240]
mtp.layers.0.self_attn.{q,k,v,o}_proj.weight   BF16
mtp.layers.0.mlp.{gate,up,down}_proj.weight    BF16
mtp.norm / pre_fc_norm_* / layernorms          BF16

But quantization_config.ignore doesn't mention mtp anywhere, while
config_groups.group_0.targets is ["Linear"]. So vLLM matches mtp.fc as a
compressed-tensors W4A16 layer, allocates weight_packed/weight_scale/weight_shape,
and the BF16 mtp.fc.weight on disk has no parameter to land on.

It looks like it came in with the copied recipe. Diffing the two ignore lists:

ignore entries mtp.* present
dbirks/Qwen3.8-27B-W4A16-AutoRound 215 yes — all 8 linears
this repo 207 none

The set difference is exactly those 8 entries, nothing else — so this is the ignore
list minus the MTP block, not a different quantization decision.

This is why it tests fine without --speculative-config: the draft model is never
instantiated and the mtp.* tensors are simply never loaded.

Fix

Config-only, no requantization — the weights are already BF16 on disk. Add to
ignore in both config.json and quantization_config.json:

"mtp.fc",
"mtp.layers.0.self_attn.q_proj",
"mtp.layers.0.self_attn.k_proj",
"mtp.layers.0.self_attn.v_proj",
"mtp.layers.0.self_attn.o_proj",
"mtp.layers.0.mlp.gate_proj",
"mtp.layers.0.mlp.up_proj",
"mtp.layers.0.mlp.down_proj"

Patching those in locally gets it past load_draft_model.

Worth keeping BF16 there deliberately rather than quantizing on a future pass — the
kernelogic/Qwen3.8-27B-…-MTP-BF16 author reports that quantizing mtp.* along with
everything else collapses draft acceptance and costs roughly half the decode speed,
which is presumably why dbirks ignored them in the first place.

I'll post this from another quant that did the same thing:

Wanted to report a specific issue when serving this model with Multi-Token Prediction (MTP) speculative decoding in vLLM (--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'), along with the verified fix.

The Problem
During startup at the Loading drafter model... stage, vLLM fails to bind the drafter weights and crashes across all worker ranks with:

ValueError: There is no module or parameter named 'layers.0.mlp.down_proj.weight' in Qwen3_5MultiTokenPredictor.
The available parameters belonging to layers.0.mlp.down_proj (RowParallelLinear) are: {'layers.0.mlp.down_proj.qweight', 'layers.0.mlp.down_proj.scales', 'layers.0.mlp.down_proj.qzeros', 'layers.0.mlp.down_proj.g_idx'}

Root Cause
Quantized MTP Layers: vLLM’s speculative execution runner (Qwen3_5MTP / llm_base_proposer.py) expects dense, unquantized parameters (.weight) for the draft transformer layer. Because mtp.layers.0.* was targeted by the 4-bit GPTQ quantization rule ("+:.mtp."), the engine cannot bind the packed .qweight tensors to the drafter graph.

Missing config.json Directives: quantization_config in config.json lacks explicit exclusion keys (modules_to_not_convert, block_name_to_quantize restriction), causing vLLM's model loader to instantiate draft modules as GPTQ containers rather than standard linear layers.

Recommended Recipe Fix for Future Builds
In future AutoRound runs or updates, keeping the mtp layers in native BF16 alongside linear_attn.in_proj_a/b and lm_head resolves this completely.

Because MTP consists of just a single transformer layer (<1.5% of total model weights), leaving it in BF16 adds virtually zero VRAM overhead while ensuring 100% draft acceptance fidelity and zero loader crashes.

Verified Workaround for Users
If anyone wants to run this specific Group 64 build with MTP speculative decoding right now:

Graft the BF16 MTP weights into the checkpoint:
Extract the native mtp.* tensors from base Qwen/Qwen3.8-27B (or dbirks/Qwen3.8-27B-W4A16-AutoRound) and replace the 4-bit mtp.* .qweight/.scales tensors inside the safetensors shards, then update model.safetensors.index.json.

Patch config.json to instruct vLLM to leave MTP unquantized:

import json

with open("config.json", "r") as f:
config = json.load(f)

q_cfg = config.get("quantization_config", {})
q_cfg["block_name_to_quantize"] = "model.language_model.layers"
q_cfg["modules_to_not_convert"] = ["mtp", "visual", "lm_head"]

if "dynamic" not in q_cfg:
q_cfg["dynamic"] = {}
q_cfg["dynamic"]["-:.mtp."] = {}

if "extra_config" not in q_cfg:
q_cfg["extra_config"] = {}
q_cfg["extra_config"][".mtp."] = {"bits": 16, "data_type": "fp"}

with open("config.json", "w") as f:
json.dump(config, f, indent=2)

Once patched, the Group 64 backbone runs at full speed via Marlin W4A16 while the unquantized draft head delivers full speculative decoding speedups without errors.

Sign up or log in to comment