Image-Text-to-Text
Safetensors
gemma4
unsloth
gemma
google
conversational
4-bit precision
bitsandbytes
Instructions to use unsloth/gemma-4-31B-it-unsloth-bnb-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
AssertionError: Weight shape mismatch in VLLM
#1
by bulatovv - opened
I am encountering an error while attempting to run inference on a vLLM.
The engine fails during the weight loading phase, within the linear layer's weight_loader.
Environment
- vLLM image:
vllm-openai/gemma4 - GPU: NVIDIA A40 (Driver: 560.35.05, CUDA: 12.9)
Logs
INFO 04-08 08:04:55 [model.py:549] Resolved architecture: Gemma4ForConditionalGeneration
INFO 04-08 08:04:55 [model.py:1680] Using max model len 262144
INFO 04-08 08:04:55 [config.py:99] Gemma4 model has heterogeneous head dimensions (head_dim=256, global_head_dim=512). Forcing TRITON_ATTN backend to prevent mixed-backend numerical divergence.
INFO 04-08 08:05:12 [core.py:105] Initializing a V1 LLM engine with config: dtype=torch.bfloat16, tensor_parallel_size=1, quantization=bitsandbytes
...
INFO 04-08 08:05:25 [bitsandbytes_loader.py:786] Loading weights with BitsAndBytes quantization. May take a while ...
Loading safetensors checkpoint shards: 100% Completed | 5/5 [00:00<00:00, 5.86it/s]
...
ERROR 04-08 08:05:35 [core.py:1108] File "vllm/model_executor/models/gemma4.py", line 993, in load_weights
ERROR 04-08 08:05:35 [core.py:1108] weight_loader(param, loaded_weight, shard_id)
ERROR 04-08 08:05:35 [core.py:1108] File "vllm/model_executor/layers/linear.py", line 1358, in weight_loader
ERROR 04-08 08:05:35 [core.py:1108] assert param_data.shape == loaded_weight.shape
ERROR 04-08 08:05:35 [core.py:1108] AssertionError
I noticed that the model has been updated since my last download, I'll try the latest version.
I also face the same issue (downloaded weights yesterday, running on A100). I would be interested if anyone has a fix.
I noticed that the model has been updated since my last download, I'll try the latest version.
I've checked the new weights and the problem persists
Bnb models are not that well supported for vllm, especially newer models.
same issue here