Instructions to use btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit") model = AutoModelForMultimodalLM.from_pretrained("btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit
- SGLang
How to use btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit with Docker Model Runner:
docker model run hf.co/btbtyler09/Qwen3.5-35B-A3B-GPTQ-4bit
Calibration data: sampling strategy, mix ratio and seqlen?
Thanks for sharing the detailed GPTQ config! A few follow-ups on the calibration data:
- Is the 2,048 samples the total across both datasets, or 2,048 per dataset?
- What's the mixing ratio between evol-codealpaca-v1 (code) and C4 (general text)
— e.g. 1:1, 8:2, or something else? - Were the 2,048 samples drawn directly from the full datasets (random / stratified
sampling), or did you first curate/filter a smaller pool and then sample 2,048 from it? - What sequence length were the samples truncated / packed to (2048, 4096, ...)?
Trying to reproduce a matching recipe on an internal Qwen3.5-35B-A3B checkpoint. Thanks!
The ratio is not explicitly set in the config I used. It ends up being ~70/30 code/c4, but depends on the max token length selected. I don't truncate anything, I sample the training datasets for complete pairs within the token length bins. If you have the hardware for it sampling at even higher token lengths may be justified. I used 4 bins with the longest token lengths being 2048. This was a hardware limitation for me as I was running into oom issues on anything longer.
Hi again!
I am currently trying to run the GPTQ quantization (4-bit, group_size 32) for Qwen3.5-35B-A3B on dual A100 (80GB) GPUs, but the process is extremely slow. It takes about 12 minutes per layer, estimating over 8 hours in total.
Here is the command I am running:
python quantize_gptq_qwen35.py \
--model_path Qwen3.5-35B-A3B \
--calib_data mixed_evol_codealpaca_c4.json \
--bits 4 --group_size 32 --calib_samples 2048

I used the en variant.
The timing sounds about right. You are essentially training the quantized model and the moe variants were slow. I think it took me around 13 hours on my setup for 3.5-35B.
As the quantization progressed to Layer 1, the estimated time indeed jumped to over 1 day (currently showing "1 day, 1:47:47" with 1 hour 17 minutes elapsed). It seems the initial 8-hour estimate was a bit optimistic because the self-attention layers processed very quickly, but once it hit the 256 sequential MoE experts, the speed dropped significantly.
I also noticed a few fallback(rtn) warnings in the logs (for example, on mlp.experts.152.gate_proj) due to Hessian instability.
Just to double check:
- Did you also experience these
fallback(rtn)warnings during your 13-hour run? - Is 24+ hours considered normal for a dual-A100 setup with 2048 calibration samples, or did you use any specific settings to avoid the slow sequential expert loops?
Yes rtn fallback is normal for rare experts. I don't know about A100s tbh. I used 4 x Mi100s. It's possible that there are settings which could speed up the quantization on A100s, but I haven't used A100s for quantization.

