DeepSeek-R1-Distill-Qwen-1.5B (IQ3_XXS GGUF)

A heavily compressed GGUF quantization of DeepSeek-R1-Distill-Qwen-1.5B โ€” the smallest official DeepSeek reasoning model.

This card covers the single file:

Property Value
File DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf
Size 733.44 MiB (769.07 MB)
Quantization IQ3_XXS (โ‰ˆ 3.06 bits/weight)
Parameters 1.78B (total) / ~1.5B (active)
Architecture Qwen2
Context length 32768 tokens
GGUF version 3 (latest)

Why this quantization

IQ3_XXS is the smallest IQ type that reliably completes the model's [Start thinking] / [End thinking] reasoning chain. At ~0.77 GB the weights load on any 4 GB-class GPU (e.g. GTX 1650) while still producing final answers instead of degenerating into an infinite thinking loop.

Note: even lighter quantizations exist (iq2_xxs at 2.06 bpw, ~0.56 GB) but on this reasoning model they frequently get stuck in the thinking loop and never emit the final answer.

It was produced from the F16 source with llama.cpp using an importance matrix computed with llama-imatrix over a wikitext-2 calibration set. The advertised context was capped at 32768 tokens so the KV cache fits a 4 GB GPU when fully offloaded (-ngl 999).

Usage

llama.cpp (CLI)

llama-cli \
  -m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
  -p "What is the capital of France?" \
  -n 512 \
  -ngl 999 \
  -c 32768

llama-server (OpenAI-compatible API)

llama-server \
  -m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
  -ngl 999 \
  -c 32768 \
  --port 8080
# curl http://localhost:8080/v1/chat/completions ...

Docker Model Runner (docker model)

docker model package \
  --gguf DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
  --context-size 32768 \
  deepseek-r1
docker model run deepseek-r1 "What is the capital of France?"

Ollama

ollama create deepseek-r1-iq3-xxs -f Modelfile   # FROM ./DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf
ollama run deepseek-r1-iq3-xxs

Quality expectations

IQ3_XXS is an aggressive compression, but it stays coherent enough to finish its reasoning and answer. Expect occasional arithmetic or factual slips and less fluency than F16 / Q4_K_M. Give the model enough output tokens: the DeepSeek-R1 chain-of-thought can use several hundred tokens before the answer (e.g. request max_tokens >= 1024 from an API). For higher quality on the same hardware, prefer Q4_K_M (1.0 GB); for a smaller file, IQ2_M (0.67 GB) also works but is slightly less reliable.

Credits & license

Downloads last month
38
GGUF
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for manalejandro/DeepSeek-R1-Distill-Qwen-1-5B-f16-iq3_xxs-GGUF

Quantized
(266)
this model