Quantized Open Models
Collection
Quantized open-weight models, reproducible recipes. • 31 items • Updated
tencent/UI-Mate-27B quantized to AWQ-W4A16 (4-bit weights).
Same footprint as W4A16 but calibrates faster and usually holds up better on instruction-tuned and multilingual models. Start here.
Caveat. Asymmetric; a few older vLLM kernels prefer symmetric W4A16.
| Source | tencent/UI-Mate-27B |
| Scheme | AWQ-W4A16 (4-bit) |
| Format | compressed-tensors |
| Parameters | 27.4B |
| Size on disk | 18.7 GB |
| Compression | 2.92x smaller than the 54.7 GB source |
| Calibration | HuggingFaceH4/ultrachat_200k, 256 samples |
| Left unquantized | lm_head, re:.*visual.*, re:.*vision_tower.*, re:.*vision_model.*, re:.*vision.*, re:.*multi_modal_projector.*, re:.*merger.* |
| Quantized on | A100 SXM |
| Quantized by | Sohailhosseini |
vllm serve Sohailhosseini/UI-Mate-27B-AWQ-W4A16 \
--max-model-len 32768
from vllm import LLM, SamplingParams
if __name__ == "__main__":
llm = LLM("Sohailhosseini/UI-Mate-27B-AWQ-W4A16", max_model_len=32768)
out = llm.chat(
[{"role": "user", "content": "What is quantization? Answer in one sentence."}],
SamplingParams(temperature=0.6, max_tokens=512),
)
print(out[0].outputs[0].text)
Produced with HF-quantized. recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers, calibration set and hardware are in the table above.
Licence is inherited from the source model. Quantization does not change what you are permitted to do with the weights.