Instructions to use alekringtonnn-ai/zubr-tiny-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use alekringtonnn-ai/zubr-tiny-2b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M # Run inference directly in the terminal: llama cli -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M # Run inference directly in the terminal: llama cli -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Use Docker
docker model run hf.co/alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use alekringtonnn-ai/zubr-tiny-2b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "alekringtonnn-ai/zubr-tiny-2b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alekringtonnn-ai/zubr-tiny-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
- Ollama
How to use alekringtonnn-ai/zubr-tiny-2b with Ollama:
ollama run hf.co/alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
- Unsloth Desktop
- Pi
How to use alekringtonnn-ai/zubr-tiny-2b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "alekringtonnn-ai/zubr-tiny-2b:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use alekringtonnn-ai/zubr-tiny-2b with Docker Model Runner:
docker model run hf.co/alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
- Lemonade
How to use alekringtonnn-ai/zubr-tiny-2b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Run and chat with the model
lemonade run user.zubr-tiny-2b-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use alekringtonnn-ai/zubr-tiny-2b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use alekringtonnn-ai/zubr-tiny-2b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alekringtonnn-ai/zubr-tiny-2b:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "alekringtonnn-ai/zubr-tiny-2b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
zubr-tiny-2b
Model Overview
zubr-tiny-2b is a lightweight, fine-tuned conversational language model based on the state-of-the-art Qwen 3.5 (and Qwen 2.5 architecture ecosystem) developed by alekringtonnn-ai.
By leveraging the powerful foundations of the Qwen series, this 2-billion parameter model offers exceptional multi-lingual capabilities, reasoning, and instruction-following proficiency while maintaining an incredibly small hardware footprint. It is highly optimized for fast local inference, low RAM/VRAM consumption, and efficient deployment on consumer-grade hardware such as laptops and edge devices.
How to Download via Terminal
You can easily download the model weights and configuration files directly from Hugging Face using your terminal. Choose one of the methods below:
Method 1: Using Hugging Face CLI (Recommended)
This is the most efficient method to download the repository or specific model shards.
Install or update the Hugging Face Hub CLI:
pip install -U huggingface_hubDownload the complete repository:
huggingface-cli download alekringtonnn-ai/zubr-tiny-2bDownload to a specific local directory:
huggingface-cli download alekringtonnn-ai/zubr-tiny-2b --local-dir ./zubr-tiny-2b
Method 2: Using Git LFS
If you prefer working with standard Git workflows, make sure Git Large File Storage is installed.
Initialize Git LFS:
git lfs installClone the repository:
git clone https://huggingface.co
Method 3: Direct Download via cURL
To pull specific configurations or single files without Python dependencies:
curl -L -O https://huggingface.co/resolve/main/config.json
Quick Start (Python)
Since the model is based on Qwen, it is fully compatible with the standard transformers library. You can run it locally using the following snippet:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "alekringtonnn-ai/zubr-tiny-2b"
# Load the tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto"
)
# Format your prompt
prompt = "Привет! Расскажи о себе."
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
# Generate response
inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=512)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
Features & Limitations
- Base Architecture: Qwen 3.5 / Qwen 2.5 2B.
- Context Length: Inherits the extended context window support from the base Qwen architecture.
- Target Use Case: Perfect for private local chatbots, text summarization, and embedded tasks where resources are highly constrained.
- Downloads last month
- 498
2-bit
4-bit