Instructions to use GotoAI-Inc/Qwen3.6-27B-W8A16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GotoAI-Inc/Qwen3.6-27B-W8A16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GotoAI-Inc/Qwen3.6-27B-W8A16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("GotoAI-Inc/Qwen3.6-27B-W8A16") model = AutoModelForMultimodalLM.from_pretrained("GotoAI-Inc/Qwen3.6-27B-W8A16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GotoAI-Inc/Qwen3.6-27B-W8A16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GotoAI-Inc/Qwen3.6-27B-W8A16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GotoAI-Inc/Qwen3.6-27B-W8A16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GotoAI-Inc/Qwen3.6-27B-W8A16
- SGLang
How to use GotoAI-Inc/Qwen3.6-27B-W8A16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GotoAI-Inc/Qwen3.6-27B-W8A16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GotoAI-Inc/Qwen3.6-27B-W8A16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GotoAI-Inc/Qwen3.6-27B-W8A16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GotoAI-Inc/Qwen3.6-27B-W8A16", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GotoAI-Inc/Qwen3.6-27B-W8A16 with Docker Model Runner:
docker model run hf.co/GotoAI-Inc/Qwen3.6-27B-W8A16
Qwen3.6-27B-W8A16
Int8 weight-only quantization of Qwen/Qwen3.6-27B, in compressed-tensors format for vLLM. 31.59 GB, down from 55.56 GB — it fits a 48 GB card with room for a long context window, or a 40 GB card with the KV budget discussed under Usage.
This is the fidelity-first build. Int8 round-to-nearest stays far closer to the bfloat16 weights than int4 does, at ~1.6x the footprint of the int4 W4A16 sibling (19.42 GB). If you are targeting a 24 GB card, use that one; use this one when you have the VRAM and want the least quality loss quantization can give without calibration.
Unofficial and unaffiliated with Alibaba/Qwen. All model capabilities, evaluations and limitations belong to the original model card — see the base model for those.
What was changed
Weights were quantized from bfloat16 to int8, group size 128, symmetric, weight-only
(activations stay 16-bit) using llmcompressor.model_free_ptq. No calibration data was
used and the model was never loaded — the quantizer operates directly on the safetensors.
Architecture, tokenizer, chat template and processor configs are the vendor's, unmodified.
496 Linear modules were converted, covering 78.3% of the checkpoint's bytes:
| component | precision | size |
|---|---|---|
| language-model linears (64 layers) | int8 g128 | 24.73 GB (78.3%) |
embed_tokens + lm_head (untied) |
bfloat16 | 5.09 GB (16.1%) |
vision tower (model.visual, 27 layers) |
bfloat16 | 0.92 GB (2.9%) |
MTP speculator head (mtp.*) |
bfloat16 | 0.85 GB (2.7%) |
| conv1d kernels, norms, biases | bfloat16 | 0.004 GB |
| total | 31.59 GB |
Four things are deliberately left at 16-bit:
model.visual.*— vLLM builds multimodal towers withquant_config=None, so a checkpoint carrying quantized vision weights cannot be loaded.mtp.*— the built-in multi-token-prediction speculator head (mtp_num_hidden_layers: 1), loaded through vLLM's speculative-decoding path rather than the main stack.linear_attn.conv1d— 3-D causal-convolution kernels in the gated-DeltaNet blocks, shape(10240, 1, 4). Not Linear layers, and quantizers reject them outright.lm_head+embed_tokens— precision-sensitive, andlm_headis untied here.
The gated-DeltaNet projections (in_proj_qkv, in_proj_a, in_proj_b, in_proj_z,
out_proj) are quantized; only the convolution kernels beside them are excluded, along
with the 1-D A_log and dt_bias state-space parameters, which any quantizer skips
automatically.
Usage
Runs on released vLLM. Qwen3.6 reuses the Qwen3.5 architecture
(Qwen3_5ForConditionalGeneration, model_type: qwen3_5), which has been supported since
0.25.1 — no nightly build is required:
vllm serve GotoAI-Inc/Qwen3.6-27B-W8A16 \
--max-model-len 65536 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder \
--reasoning-parser qwen3
Do not pass --quantization; compressed-tensors is detected from config.json. The int8
W8A16 scheme uses Marlin kernels and runs on compute capability 7.5 and above.
--tool-call-parser qwen3_coderis what the base model card specifies. In current vLLMqwen3_xmlis an alias for the same parser class, so either name works; without one, the<tool_call><function=…><parameter=…>XML the chat template asks for is returned as plain text.--reasoning-parser qwen3splits<think>…</think>intoreasoning_content.--language-model-onlyskips the vision tower and its multimodal profiling, freeing ~0.92 GB of weights plus the profiling headroom, at the cost of image and video input.- MTP speculative decoding uses the head already in this checkpoint — no draft model
to download:
--speculative-config '{"method": "mtp", "num_speculative_tokens": 2}'. The base card writes"method": "qwen3_next_mtp"; current vLLM deprecates the per-family names and normalizes them tomtp, resolvingqwen3_5toQwen3_5MTPfrom the checkpoint's own config. vLLM aligns the draft's quantization with the target's, and there:.*mtp.*entry in this checkpoint's ignore list keeps those tensors bf16. Not smoke-tested here.
Fitting the card
Only 16 of the 64 layers use full attention (every 4th; the other 48 are gated DeltaNet
with constant-size recurrent state). With 4 KV heads at head_dim 256, the KV cache costs
~64 KB/token — about 2 GB at 32k and 4 GB at 64k. Against 31.59 GB of weights that is
comfortable at 64k on a 48 GB card; on a 40 GB card, budget for roughly 32k of context, or
add --language-model-only for more. This is arithmetic from config.json, not a measured
deployment.
Controlling thinking
Qwen3.6 thinks by default and does not support the /think and /nothink soft
switches. It also has no reasoning_effort knob — the two template variables it does
accept are passed through chat_template_kwargs:
{"chat_template_kwargs": {"enable_thinking": false}} // instruct / non-thinking mode
{"chat_template_kwargs": {"preserve_thinking": true}} // keep earlier turns' thinking
preserve_thinking is the feature this release adds: by default only the thinking from
the latest user message is retained, and turning it on keeps historical reasoning traces
in context — the base model card recommends it for agentic use, where it improves decision
consistency and KV-cache reuse. Set a server-wide default with
--default-chat-template-kwargs '{"preserve_thinking": true}'; request-level values still
win.
Sampling, per the base model card: temperature=1.0, top_p=0.95, top_k=20 for thinking
mode, temperature=0.6 for precise coding, and temperature=0.7, top_p=0.80, presence_penalty=1.5 in non-thinking mode.
Context
262144 tokens natively. The base model card documents a YaRN recipe reaching 1,010,000
tokens via --hf-overrides plus VLLM_ALLOW_LONG_MAX_MODEL_LEN=1; rope_type is left at
default here, and static YaRN costs quality at short contexts, so enable it only if you
need it.
Reproducing this checkpoint
Built with llm-quantizer:
./llmq.py run --profile qwen3.6-27b
W8A16 is the profile's default scheme, so no --scheme flag is needed. The command is
equivalent to:
# llmcompressor==0.13.1a20260814, compressed-tensors==0.18.1a20260818,
# transformers==5.15.1, torch==2.13.0
from llmcompressor import model_free_ptq
model_free_ptq(
model_stub="Qwen/Qwen3.6-27B",
save_directory="Qwen3.6-27B-W8A16",
scheme="W8A16",
ignore=["re:.*visual.*", "re:.*mtp.*", "re:.*\\.conv1d$",
"lm_head", "re:.*embed_tokens.*"],
device="cuda:0",
)
The source ships as 15 shards of ~4 GB, and a job holds one shard at a time, so the build peaks at a few GB of VRAM — no re-sharding needed and no large GPU required.
Evaluation
No benchmarks have been run. Data-free round-to-nearest quantization degrades quality more than a calibrated (GPTQ/AWQ) or QAT build; how much, for your task, is unmeasured here. Int8 degrades far less than int4 — that is the reason this build exists — but "less" is not "none". Treat the published Qwen3.6 numbers as describing the bfloat16 model, not this one.
For an agentic model the informative checks are well-formed reasoning_content and clean
multi-step tool calls rather than perplexity: structured emission degrades before fluency
does.
License
Apache 2.0, inherited from the base model — the vendor's LICENSE is included unmodified.
"Qwen" is Alibaba's mark; this repository is not endorsed by or affiliated with Alibaba.
- Downloads last month
- 26
Model tree for GotoAI-Inc/Qwen3.6-27B-W8A16
Base model
Qwen/Qwen3.6-27B