Instructions to use alekringtonnn-ai/zubr-mini-1.9-3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alekringtonnn-ai/zubr-mini-1.9-3b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="alekringtonnn-ai/zubr-mini-1.9-3b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("alekringtonnn-ai/zubr-mini-1.9-3b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use alekringtonnn-ai/zubr-mini-1.9-3b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M # Run inference directly in the terminal: llama cli -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M # Run inference directly in the terminal: llama cli -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Use Docker
docker model run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use alekringtonnn-ai/zubr-mini-1.9-3b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "alekringtonnn-ai/zubr-mini-1.9-3b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alekringtonnn-ai/zubr-mini-1.9-3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
- SGLang
How to use alekringtonnn-ai/zubr-mini-1.9-3b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "alekringtonnn-ai/zubr-mini-1.9-3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alekringtonnn-ai/zubr-mini-1.9-3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "alekringtonnn-ai/zubr-mini-1.9-3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alekringtonnn-ai/zubr-mini-1.9-3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use alekringtonnn-ai/zubr-mini-1.9-3b with Ollama:
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
- Unsloth Desktop
- Pi
How to use alekringtonnn-ai/zubr-mini-1.9-3b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use alekringtonnn-ai/zubr-mini-1.9-3b with Docker Model Runner:
docker model run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
- Lemonade
How to use alekringtonnn-ai/zubr-mini-1.9-3b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Run and chat with the model
lemonade run user.zubr-mini-1.9-3b-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use alekringtonnn-ai/zubr-mini-1.9-3b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use alekringtonnn-ai/zubr-mini-1.9-3b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
🦬 Zubr Mini 1.9 3B
Repository: alekringtonnn-ai/zubr-mini-1.9-3b
📖 What This Model Is
zubr-mini-1.9-3b is a multimodal (Vision) language model with 3 billion parameters, fine-tuned from Mistral 3 3B (Ministral-3-3B-Instruct-2512 by Mistral AI). Developed by alekringtonnn-ai.
The model is distributed in GGUF format — a universal format for quantized weights designed for local inference. It runs on regular PCs (CPU), laptops, mobile devices, and servers without expensive GPUs.
💡 For Macs with Apple Silicon (M1–M4), there is a separate MLX version of this same model: alekringtonnn-ai/zubr-mini-1.9-3b-mlx.
Key Features
- Multimodal (Vision): the model supports image recognition and Image-Text-to-Text tasks (image + text → text).
- Multilingual: language support expanded from 26 to 48 languages (including Russian).
- Performance: improved vocabulary, high inference speed, and deep instruction tuning.
- Local execution: the GGUF format runs on CPU and GPU via Ollama, llama.cpp, vLLM, LM Studio, and other runtimes.
- License: Apache 2.0 — free to use, including for commercial purposes.
Model Tree
mistralai/Ministral-3-3B-Base-2512
│ (quantization / instruct version)
mistralai/Ministral-3-3B-Instruct-2512
│
mistralai/Ministral-3-3B-Instruct-2512-GGUF
│ (fine-tune by alekringtonnn-ai)
alekringtonnn-ai/zubr-mini-1.9-3b ← GGUF version (THIS repository)
│ (conversion to MLX)
alekringtonnn-ai/zubr-mini-1.9-3b-mlx ← MLX version for Apple Silicon
📊 Versions (Quantizations) in the Repository
The repository offers 8 GGUF quantization options — from an ultra-compact 2-bit one to a nearly original 8-bit one:
| Version | File | Size | What it is | What it's for |
|---|---|---|---|---|
| Q2_K | zubr-mini-1.9.Q2_K.gguf | ~1.46 GB | The most aggressive 2-bit quantization | Devices with minimal memory (~4 GB RAM). Quality suffers noticeably |
| Q2_K_L | zubr-mini-1.9.Q2_K_L.gguf | ~1.56 GB | An improved 2-bit quantization with larger high-precision layers | Same as Q2_K, but with slightly higher quality in certain layers |
| IQ3_XS | zubr-mini-1.9.IQ3_XS.gguf | ~1.58 GB | Intelligent 3-bit quantization (IQ) — more compact than classic 3-bit | Low-power devices (~6 GB RAM): better than Q2, smaller than Q3_K_L |
| Q3_K_L | zubr-mini-1.9.Q3_K_L.gguf | ~1.93 GB | 3-bit quantization with large high-precision layers | Devices with limited memory (~8 GB RAM) — a quality/size balance |
| Q4_0 | zubr-mini-1.9.Q4_0.gguf | ~2.05 GB | Classic 4-bit quantization (fast, simple format) | A well-established legacy format: high speed, ~8 GB RAM |
| Q4_K_M ⭐ | zubr-mini-1.9.Q4_K_M.gguf | ~2.15 GB | 4-bit quantization with medium high-precision layers | The recommended choice for most users: near-full quality at only ~2.15 GB |
| Q6_K | zubr-mini-1.9.Q6_K.gguf | ~2.82 GB | 6-bit quantization — high quality | For tasks where quality matters (e.g., RAG, precise answers). Requires ~12 GB RAM |
| Q8_0 | zubr-mini-1.9.Q8_0.gguf | ~3.65 GB | 8-bit quantization — virtually lossless quality | Maximum quality with no performance loss. Requires ~16 GB RAM |
💡 How to choose: the standard recommendation is Q4_K_M (quality/size balance). If memory is tight — Q3_K_L or IQ3_XS; if you need maximum quality — Q6_K or Q8_0.
📥 How to Download
Requirements
- A CPU with AVX/AVX2 support (almost any modern x86) or an ARM device; a GPU is optional.
- RAM: at least twice the size of the quantization file (see the table above).
- Free disk space: 1.5–3.7 GB depending on the version.
- For vLLM — an NVIDIA GPU (or Apple Silicon) and Linux/macOS.
Method 1: Ollama (the easiest)
# Install Ollama: https://ollama.com
# The recommended version:
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
# Other quantizations — just change the tag:
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q2_K
ollama run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q8_0
Ollama will automatically download the required GGUF file and start a chat in your terminal.
Method 2: llama.cpp
# The recommended version — download and run in one command:
llama-cli -hf alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
# Or download the file and specify it explicitly:
llama-cli -m zubr-mini-1.9.Q4_K_M.gguf -p "Hello! Tell me about yourself."
Method 3: huggingface-cli (direct file downloads)
# Install the CLI
pip install -U "huggingface_hub[cli]"
# Download a specific quantization:
huggingface-cli download alekringtonnn-ai/zubr-mini-1.9-3b \
zubr-mini-1.9.Q4_K_M.gguf
# Download the entire repository (all 8 quantizations, ~17 GB):
huggingface-cli download alekringtonnn-ai/zubr-mini-1.9-3b
By default, files are saved to ~/.cache/huggingface/hub/.
Method 4: Direct links (browser / wget / curl)
curl -L -o zubr-mini-1.9.Q4_K_M.gguf \
"https://huggingface.co/alekringtonnn-ai/zubr-mini-1.9-3b/resolve/main/zubr-mini-1.9.Q4_K_M.gguf"
The same links also work in a browser: open the file page and click "Download", or append ?download=true to the URL.
Method 5: LM Studio / Jan (no command line)
- Install LM Studio or Jan.
- In the search tab, type
zubr-mini-1.9-3b(or paste the repository link). - Select a quantization and click Download.
- Start chatting in the built-in chat.
Method 6: Docker
docker model run hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M
🚀 How to Run
Via vLLM (API server)
pip install vllm
vllm serve "alekringtonnn-ai/zubr-mini-1.9-3b"
Once started, the model is available through an OpenAI-compatible API at http://localhost:8000/v1.
Via Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="alekringtonnn-ai/zubr-mini-1.9-3b",
filename="*Q4_K_M.gguf", # Or any other quantization
verbose=False,
)
response = llm.create_chat_completion(
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello! Write a short poem about a bison."},
]
)
print(response["choices"][0]["message"]["content"])
Via Ollama in Python
import ollama
response = ollama.chat(
model="hf.co/alekringtonnn-ai/zubr-mini-1.9-3b:Q4_K_M",
messages=[
{"role": "user", "content": "Hello! Tell me about yourself."}
]
)
print(response["message"]["content"])
⚙️ Recommended Generation Parameters
| Parameter | Value | Comment |
|---|---|---|
temperature |
0.7 | A balance between creativity and accuracy |
max_tokens |
512–2048 | Depends on the task |
top_p |
0.9 | A standard value for instruct models |
num_ctx (context) |
4096+ | Context window size; affects RAM usage |
- Downloads last month
- 719
2-bit
3-bit
4-bit
6-bit
8-bit