Image-Text-to-Text
Transformers
Safetensors
qwen3_5_moe
conversational
8-bit precision
compressed-tensors
Instructions to use ig1/Qwen3.6-35B-A3B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ig1/Qwen3.6-35B-A3B-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ig1/Qwen3.6-35B-A3B-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ig1/Qwen3.6-35B-A3B-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("ig1/Qwen3.6-35B-A3B-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ig1/Qwen3.6-35B-A3B-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ig1/Qwen3.6-35B-A3B-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ig1/Qwen3.6-35B-A3B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ig1/Qwen3.6-35B-A3B-NVFP4
- SGLang
How to use ig1/Qwen3.6-35B-A3B-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ig1/Qwen3.6-35B-A3B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ig1/Qwen3.6-35B-A3B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ig1/Qwen3.6-35B-A3B-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ig1/Qwen3.6-35B-A3B-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ig1/Qwen3.6-35B-A3B-NVFP4 with Docker Model Runner:
docker model run hf.co/ig1/Qwen3.6-35B-A3B-NVFP4
Update README.md
Browse files
README.md
CHANGED
|
@@ -200,7 +200,7 @@ docker run --rm --name 'Qwen3.6-35B-A3B-NVFP4' \
|
|
| 200 |
| `--max-num-seqs` | `≤32` | Limit concurrent sequences, can not be higher than max-cudagraph-capture-size | Included above |
|
| 201 |
| `--max-num-batched-tokens` | `2048` | Reduce activation buffers | Saves ~350 MiB |
|
| 202 |
|
| 203 |
-
About the `--speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'`: this will speed up generation in a low concurrent users/requests scenario but the MTP layers eat ~
|
| 204 |
|
| 205 |
#### Windows with Docker for WSL
|
| 206 |
|
|
|
|
| 200 |
| `--max-num-seqs` | `≤32` | Limit concurrent sequences, can not be higher than max-cudagraph-capture-size | Included above |
|
| 201 |
| `--max-num-batched-tokens` | `2048` | Reduce activation buffers | Saves ~350 MiB |
|
| 202 |
|
| 203 |
+
About the `--speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'`: this will speed up generation in a low concurrent users/requests scenario but the MTP layers eat ~1.69G of your VRAM decreasing the space available for the KV cache (equivalent to roughly ~26,400 tokens). Because the purpose here was to save VRAM we did not include it. But in the end, it is up to you to decide if you prefer faster inference with a lower KV cache and slower inference with a bigger KV cache.
|
| 204 |
|
| 205 |
#### Windows with Docker for WSL
|
| 206 |
|