DeepSeek-R1-Distill-Qwen-1.5B (IQ3_XXS GGUF)
A heavily compressed GGUF quantization of DeepSeek-R1-Distill-Qwen-1.5B โ the smallest official DeepSeek reasoning model.
This card covers the single file:
| Property | Value |
|---|---|
| File | DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf |
| Size | 733.44 MiB (769.07 MB) |
| Quantization | IQ3_XXS (โ 3.06 bits/weight) |
| Parameters | 1.78B (total) / ~1.5B (active) |
| Architecture | Qwen2 |
| Context length | 32768 tokens |
| GGUF version | 3 (latest) |
Why this quantization
IQ3_XXS is the smallest IQ type that reliably completes the model's
[Start thinking] / [End thinking] reasoning chain. At ~0.77 GB the weights
load on any 4 GB-class GPU (e.g. GTX 1650) while still producing final answers
instead of degenerating into an infinite thinking loop.
Note: even lighter quantizations exist (
iq2_xxsat 2.06 bpw, ~0.56 GB) but on this reasoning model they frequently get stuck in the thinking loop and never emit the final answer.
It was produced from the F16 source with llama.cpp using an importance
matrix computed with llama-imatrix over a wikitext-2 calibration set. The
advertised context was capped at 32768 tokens so the KV cache fits a 4 GB GPU
when fully offloaded (-ngl 999).
Usage
llama.cpp (CLI)
llama-cli \
-m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
-p "What is the capital of France?" \
-n 512 \
-ngl 999 \
-c 32768
llama-server (OpenAI-compatible API)
llama-server \
-m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
-ngl 999 \
-c 32768 \
--port 8080
# curl http://localhost:8080/v1/chat/completions ...
Docker Model Runner (docker model)
docker model package \
--gguf DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
--context-size 32768 \
deepseek-r1
docker model run deepseek-r1 "What is the capital of France?"
Ollama
ollama create deepseek-r1-iq3-xxs -f Modelfile # FROM ./DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf
ollama run deepseek-r1-iq3-xxs
Quality expectations
IQ3_XXS is an aggressive compression, but it stays coherent enough to finish
its reasoning and answer. Expect occasional arithmetic or factual slips and
less fluency than F16 / Q4_K_M. Give the model enough output tokens: the
DeepSeek-R1 chain-of-thought can use several hundred tokens before the answer
(e.g. request max_tokens >= 1024 from an API). For higher quality on the
same hardware, prefer Q4_K_M (1.0 GB); for a smaller file, 0.67 GB)
also works but is slightly less reliable.IQ2_M (
Credits & license
- Base model: deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B (MIT)
- Source GGUF: bartowski/DeepSeek-R1-Distill-Qwen-1.5B-GGUF (MIT)
- Quantized with llama.cpp +
llama-imatrix - This file is released under the MIT license.
- Downloads last month
- 38
3-bit
Model tree for manalejandro/DeepSeek-R1-Distill-Qwen-1-5B-f16-iq3_xxs-GGUF
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B