🦬 Zubr Mini 1.9 3B

Repository: alekringtonnn-ai/zubr-mini-1.9-3b


📖 What This Model Is

zubr-mini-1.9-3b is a multimodal (Vision) language model with 3 billion parameters, fine-tuned from Mistral 3 3B (Ministral-3-3B-Instruct-2512 by Mistral AI). Developed by alekringtonnn-ai.

The model is distributed in GGUF format — a universal format for quantized weights designed for local inference. It runs on regular PCs (CPU), laptops, mobile devices, and servers without expensive GPUs.

💡 For Macs with Apple Silicon (M1–M4), there is a separate MLX version of this same model: alekringtonnn-ai/zubr-mini-1.9-3b-mlx.

Key Features

  • Multimodal (Vision): the model supports image recognition and Image-Text-to-Text tasks (image + text → text).
  • Multilingual: language support expanded from 26 to 48 languages (including Russian).
  • Performance: improved vocabulary, high inference speed, and deep instruction tuning.
  • Local execution: the GGUF format runs on CPU and GPU via Ollama, llama.cpp, vLLM, LM Studio, and other runtimes.
  • License: Apache 2.0 — free to use, including for commercial purposes.

Model Tree

mistralai/Ministral-3-3B-Base-2512
        │ (quantization / instruct version)
mistralai/Ministral-3-3B-Instruct-2512
        │
mistralai/Ministral-3-3B-Instruct-2512-GGUF
        │ (fine-tune by alekringtonnn-ai)
alekringtonnn-ai/zubr-mini-1.9-3b          ← GGUF version (THIS repository)
        │ (conversion to MLX)
alekringtonnn-ai/zubr-mini-1.9-3b-mlx       ← MLX version for Apple Silicon

📊 Versions (Quantizations) in the Repository

The repository offers 8 GGUF quantization options — from an ultra-compact 2-bit one to a nearly original 8-bit one:

Version File Size What it is What it's for
Q2_K zubr-mini-1.9.Q2_K.gguf ~1.46 GB The most aggressive 2-bit quantization Devices with minimal memory (~4 GB RAM). Quality suffers noticeably
Q2_K_L zubr-mini-1.9.Q2_K_L.gguf ~1.56 GB An improved 2-bit quantization with larger high-precision layers Same as Q2_K, but with slightly higher quality in certain layers
IQ3_XS zubr-mini-1.9.IQ3_XS.gguf ~1.58 GB Intelligent 3-bit quantization (IQ) — more compact than classic 3-bit Low-power devices (~6 GB RAM): better than Q2, smaller than Q3_K_L
Q3_K_L zubr-mini-1.9.Q3_K_L.gguf ~1.93 GB 3-bit quantization with large high-precision layers Devices with limited memory (~8 GB RAM) — a quality/size balance
Q4_0 zubr-mini-1.9.Q4_0.gguf ~2.05 GB Classic 4-bit quantization (fast, simple format) A well-established legacy format: high speed, ~8 GB RAM
Q4_K_M ⭐ zubr-mini-1.9.Q4_K_M.gguf ~2.15 GB 4-bit quantization with medium high-precision layers The recommended choice for most users: near-full quality at only ~2.15 GB
Q6_K zubr-mini-1.9.Q6_K.gguf ~2.82 GB 6-bit quantization — high quality For tasks where quality matters (e.g., RAG, precise answers). Requires ~12 GB RAM
Q8_0 zubr-mini-1.9.Q8_0.gguf ~3.65 GB 8-bit quantization — virtually lossless quality Maximum quality with no performance loss. Requires ~16 GB RAM

💡 How to choose: the standard recommendation is Q4_K_M (quality/size balance). If memory is tight — Q3_K_L or IQ3_XS; if you need maximum quality — Q6_K or Q8_0.


📥 How to Download

Requirements

  • A CPU with AVX/AVX2 support (almost any modern x86) or an ARM device; a GPU is optional.
  • RAM: at least twice the size of the quantization file (see the table above).
  • Free disk space: 1.5–3.7 GB depending on the version.
  • For vLLM — an NVIDIA GPU (or Apple Silicon) and Linux/macOS.

Method 1: Ollama (the easiest)

# Install Ollama: https://ollama.com
# The recommended version:
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M

# Other quantizations — just change the tag:
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q2_K
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q8_0

Ollama will automatically download the required GGUF file and start a chat in your terminal.

Method 2: llama.cpp

# The recommended version — download and run in one command:
llama-cli -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M

# Or download the file and specify it explicitly:
llama-cli -m zubr-mini-1.9.Q4_K_M.gguf -p "Hello! Tell me about yourself."

Method 3: huggingface-cli (direct file downloads)

# Install the CLI
pip install -U "huggingface_hub[cli]"

# Download a specific quantization:
huggingface-cli download alekringtonnn-ai/zubr-mini-1.9-3b \
  zubr-mini-1.9.Q4_K_M.gguf

# Download the entire repository (all 8 quantizations, ~17 GB):
huggingface-cli download alekringtonnn-ai/zubr-mini-1.9-3b

By default, files are saved to ~/.cache/huggingface/hub/.

Method 4: Direct links (browser / wget / curl)

curl -L -o zubr-mini-1.9.Q4_K_M.gguf \
  "https://huggingface.co/alekringtonnn-ai/zubr-mini-1.9-3b/resolve/main/zubr-mini-1.9.Q4_K_M.gguf"

The same links also work in a browser: open the file page and click "Download", or append ?download=true to the URL.

Method 5: LM Studio / Jan (no command line)

  1. Install LM Studio or Jan.
  2. In the search tab, type zubr-mini-1.9-3b (or paste the repository link).
  3. Select a quantization and click Download.
  4. Start chatting in the built-in chat.

Method 6: Docker

docker model run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M

🚀 How to Run

Via vLLM (API server)

pip install vllm
vllm serve "alekringtonnn-ai/zubr-mini-1.9-3b"

Once started, the model is available through an OpenAI-compatible API at http://localhost:8000/v1.

Via Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="alekringtonnn-ai/zubr-mini-1.9-3b",
    filename="*Q4_K_M.gguf",  # Or any other quantization
    verbose=False,
)

response = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello! Write a short poem about a bison."},
    ]
)
print(response["choices"][0]["message"]["content"])

Via Ollama in Python

import ollama

response = ollama.chat(
    model="hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M",
    messages=[
        {"role": "user", "content": "Hello! Tell me about yourself."}
    ]
)
print(response["message"]["content"])

⚙️ Recommended Generation Parameters

Parameter Value Comment
temperature 0.7 A balance between creativity and accuracy
max_tokens 512–2048 Depends on the task
top_p 0.9 A standard value for instruct models
num_ctx (context) 4096+ Context window size; affects RAM usage

Downloads last month
719
GGUF
Model size
3B params
Architecture
mistral3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alekringtonnn-ai/zubr-mini-1.9-3b