Instructions to use jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound") model = AutoModelForMultimodalLM.from_pretrained("jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound
- SGLang
How to use jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound with Docker Model Runner:
docker model run hf.co/jamesbrunet/Swift-Qwen3.8-27b-W4A16-AutoRound
MTP head fails to load in vLLM: mtp.* linears missing from quantization_config.ignore
Thanks for putting this up — the recipe note pointing at dbirks was useful.
One issue: the model can't be served with speculative decoding enabled. vLLM dies at
worker start:
ValueError: There is no module or parameter named 'fc.weight' in Qwen3_5MultiTokenPredictor.
The available parameters belonging to fc (ColumnParallelLinear) are:
{'fc.weight_packed', 'fc.weight_shape', 'fc.weight_scale'}
(vllm/model_executor/models/qwen3_5_mtp.py → load_weights, viaspec_decode/mtp/speculator.py::load_draft_model)
Cause
The 15 mtp.* tensors in model_extra_tensors.safetensors are BF16 — correct, and
what you want:
mtp.fc.weight BF16 [5120, 10240]
mtp.layers.0.self_attn.{q,k,v,o}_proj.weight BF16
mtp.layers.0.mlp.{gate,up,down}_proj.weight BF16
mtp.norm / pre_fc_norm_* / layernorms BF16
But quantization_config.ignore doesn't mention mtp anywhere, whileconfig_groups.group_0.targets is ["Linear"]. So vLLM matches mtp.fc as a
compressed-tensors W4A16 layer, allocates weight_packed/weight_scale/weight_shape,
and the BF16 mtp.fc.weight on disk has no parameter to land on.
It looks like it came in with the copied recipe. Diffing the two ignore lists:
| ignore entries | mtp.* present |
|
|---|---|---|
dbirks/Qwen3.8-27B-W4A16-AutoRound |
215 | yes — all 8 linears |
| this repo | 207 | none |
The set difference is exactly those 8 entries, nothing else — so this is the ignore
list minus the MTP block, not a different quantization decision.
This is why it tests fine without --speculative-config: the draft model is never
instantiated and the mtp.* tensors are simply never loaded.
Fix
Config-only, no requantization — the weights are already BF16 on disk. Add toignore in both config.json and quantization_config.json:
"mtp.fc",
"mtp.layers.0.self_attn.q_proj",
"mtp.layers.0.self_attn.k_proj",
"mtp.layers.0.self_attn.v_proj",
"mtp.layers.0.self_attn.o_proj",
"mtp.layers.0.mlp.gate_proj",
"mtp.layers.0.mlp.up_proj",
"mtp.layers.0.mlp.down_proj"
Patching those in locally gets it past load_draft_model.
Worth keeping BF16 there deliberately rather than quantizing on a future pass — thekernelogic/Qwen3.8-27B-…-MTP-BF16 author reports that quantizing mtp.* along with
everything else collapses draft acceptance and costs roughly half the decode speed,
which is presumably why dbirks ignored them in the first place.
I'll post this from another quant that did the same thing:
Wanted to report a specific issue when serving this model with Multi-Token Prediction (MTP) speculative decoding in vLLM (--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'), along with the verified fix.
The Problem
During startup at the Loading drafter model... stage, vLLM fails to bind the drafter weights and crashes across all worker ranks with:
ValueError: There is no module or parameter named 'layers.0.mlp.down_proj.weight' in Qwen3_5MultiTokenPredictor.
The available parameters belonging to layers.0.mlp.down_proj (RowParallelLinear) are: {'layers.0.mlp.down_proj.qweight', 'layers.0.mlp.down_proj.scales', 'layers.0.mlp.down_proj.qzeros', 'layers.0.mlp.down_proj.g_idx'}
Root Cause
Quantized MTP Layers: vLLM’s speculative execution runner (Qwen3_5MTP / llm_base_proposer.py) expects dense, unquantized parameters (.weight) for the draft transformer layer. Because mtp.layers.0.* was targeted by the 4-bit GPTQ quantization rule ("+:.mtp."), the engine cannot bind the packed .qweight tensors to the drafter graph.
Missing config.json Directives: quantization_config in config.json lacks explicit exclusion keys (modules_to_not_convert, block_name_to_quantize restriction), causing vLLM's model loader to instantiate draft modules as GPTQ containers rather than standard linear layers.
Recommended Recipe Fix for Future Builds
In future AutoRound runs or updates, keeping the mtp layers in native BF16 alongside linear_attn.in_proj_a/b and lm_head resolves this completely.
Because MTP consists of just a single transformer layer (<1.5% of total model weights), leaving it in BF16 adds virtually zero VRAM overhead while ensuring 100% draft acceptance fidelity and zero loader crashes.
Verified Workaround for Users
If anyone wants to run this specific Group 64 build with MTP speculative decoding right now:
Graft the BF16 MTP weights into the checkpoint:
Extract the native mtp.* tensors from base Qwen/Qwen3.8-27B (or dbirks/Qwen3.8-27B-W4A16-AutoRound) and replace the 4-bit mtp.* .qweight/.scales tensors inside the safetensors shards, then update model.safetensors.index.json.
Patch config.json to instruct vLLM to leave MTP unquantized:
import json
with open("config.json", "r") as f:
config = json.load(f)
q_cfg = config.get("quantization_config", {})
q_cfg["block_name_to_quantize"] = "model.language_model.layers"
q_cfg["modules_to_not_convert"] = ["mtp", "visual", "lm_head"]
if "dynamic" not in q_cfg:
q_cfg["dynamic"] = {}
q_cfg["dynamic"]["-:.mtp."] = {}
if "extra_config" not in q_cfg:
q_cfg["extra_config"] = {}
q_cfg["extra_config"][".mtp."] = {"bits": 16, "data_type": "fp"}
with open("config.json", "w") as f:
json.dump(config, f, indent=2)
Once patched, the Group 64 backbone runs at full speed via Marlin W4A16 while the unquantized draft head delivers full speculative decoding speedups without errors.