Quantized Open Models
Collection
Quantized open-weight models, reproducible recipes. • 31 items • Updated
Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled quantized to AWQ-W4A16 (4-bit weights).
Same footprint as W4A16 but calibrates faster and usually holds up better on instruction-tuned and multilingual models. Start here.
Caveat. Asymmetric; a few older vLLM kernels prefer symmetric W4A16.
| Source | Jackrong/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled |
| Scheme | AWQ-W4A16 (4-bit) |
| Format | compressed-tensors |
| Parameters | 9.7B |
| Size on disk | 8.6 GB |
| Compression | 2.24x smaller than the 19.3 GB source |
| Calibration | HuggingFaceH4/ultrachat_200k, 256 samples |
| Left unquantized | lm_head, re:.*visual.*, re:.*vision_tower.*, re:.*vision_model.*, re:.*vision.*, re:.*multi_modal_projector.*, re:.*merger.* |
| Quantized on | RTX A6000 |
| Quantized by | Sohailhosseini |
vllm serve Sohailhosseini/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-AWQ-W4A16 \
--max-model-len 32768
from vllm import LLM, SamplingParams
if __name__ == "__main__":
llm = LLM("Sohailhosseini/Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled-AWQ-W4A16", max_model_len=32768)
out = llm.chat(
[{"role": "user", "content": "What is quantization? Answer in one sentence."}],
SamplingParams(temperature=0.6, max_tokens=512),
)
print(out[0].outputs[0].text)
Produced with HF-quantized. recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers, calibration set and hardware are in the table above.
Licence is inherited from the source model. Quantization does not change what you are permitted to do with the weights.
Base model
Qwen/Qwen3.5-9B-Base