AssertionError: Weight shape mismatch in VLLM

#1
by bulatovv - opened

I am encountering an error while attempting to run inference on a vLLM.
The engine fails during the weight loading phase, within the linear layer's weight_loader.

Environment

  • vLLM image: vllm-openai/gemma4
  • GPU: NVIDIA A40 (Driver: 560.35.05, CUDA: 12.9)

Logs

INFO 04-08 08:04:55 [model.py:549] Resolved architecture: Gemma4ForConditionalGeneration
INFO 04-08 08:04:55 [model.py:1680] Using max model len 262144
INFO 04-08 08:04:55 [config.py:99] Gemma4 model has heterogeneous head dimensions (head_dim=256, global_head_dim=512). Forcing TRITON_ATTN backend to prevent mixed-backend numerical divergence.
INFO 04-08 08:05:12 [core.py:105] Initializing a V1 LLM engine with config: dtype=torch.bfloat16, tensor_parallel_size=1, quantization=bitsandbytes
...
INFO 04-08 08:05:25 [bitsandbytes_loader.py:786] Loading weights with BitsAndBytes quantization. May take a while ...
Loading safetensors checkpoint shards: 100% Completed | 5/5 [00:00<00:00,  5.86it/s]
...
ERROR 04-08 08:05:35 [core.py:1108]   File "vllm/model_executor/models/gemma4.py", line 993, in load_weights
ERROR 04-08 08:05:35 [core.py:1108]     weight_loader(param, loaded_weight, shard_id)
ERROR 04-08 08:05:35 [core.py:1108]   File "vllm/model_executor/layers/linear.py", line 1358, in weight_loader
ERROR 04-08 08:05:35 [core.py:1108]     assert param_data.shape == loaded_weight.shape
ERROR 04-08 08:05:35 [core.py:1108] AssertionError

I noticed that the model has been updated since my last download, I'll try the latest version.

I also face the same issue (downloaded weights yesterday, running on A100). I would be interested if anyone has a fix.

I noticed that the model has been updated since my last download, I'll try the latest version.

I've checked the new weights and the problem persists

Unsloth AI org

Bnb models are not that well supported for vllm, especially newer models.

same issue here

Sign up or log in to comment