Instructions to use Abdusin/gemma-3-4b-est-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Abdusin/gemma-3-4b-est-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M # Run inference directly in the terminal: llama cli -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M # Run inference directly in the terminal: llama cli -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Abdusin/gemma-3-4b-est-v1:Q4_K_M
Use Docker
docker model run hf.co/Abdusin/gemma-3-4b-est-v1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Abdusin/gemma-3-4b-est-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Abdusin/gemma-3-4b-est-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Abdusin/gemma-3-4b-est-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Abdusin/gemma-3-4b-est-v1:Q4_K_M
- Ollama
How to use Abdusin/gemma-3-4b-est-v1 with Ollama:
ollama run hf.co/Abdusin/gemma-3-4b-est-v1:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Abdusin/gemma-3-4b-est-v1 with Docker Model Runner:
docker model run hf.co/Abdusin/gemma-3-4b-est-v1:Q4_K_M
- Lemonade
How to use Abdusin/gemma-3-4b-est-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Abdusin/gemma-3-4b-est-v1:Q4_K_M
Run and chat with the model
lemonade run user.gemma-3-4b-est-v1-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Gemma 3 4B Estonian (v1)
A Gemma 3 4B base model adapted for the Estonian language. The stock Gemma 3 4B knows surprisingly little Estonian; this model was trained to fix that while staying small enough to run comfortably on a phone.
What was done
Two stages, both with QLoRA on a single consumer GPU (RTX 5070 12 GB, ~26 hours total):
- Continued pretraining on 161.5M tokens of Estonian web text (the Estonian subset of FineWeb-2), packed to 2048-token blocks, one epoch, learning rate 5e-5.
- Supervised fine-tuning on ~30.5k examples (one epoch, lr 1e-4):
- 15k Estonian instructions from tartuNLP/magpie-gemma-3-12b-it-100k-et
- ~2.9k responses distilled from the open EstLLM-8B model on native Estonian prompts (dictionary explanations, inflection, reading comprehension)
- ~11k pairs built from public Estonian language resources: EKI dictionary definitions → word (word_meanings_et), noun-phrase inflection (inflection_et), grammar correction with gold targets (grammar_et), extractive QA (EstQA), and news summarization (ERRnews)
Evaluation
Measured with the official Estonian LLM benchmark harness (LREC 2026, Lillepalu & Alumäe), full test sets, zero-shot with the chat template. Published numbers for reference models are from the benchmark paper (arXiv:2510.21193).
| Task (exact match) | Gemma-3-4B-it (base) | This model | EstLLM-8B (published) |
|---|---|---|---|
| Inflection (1,400) | 0.107 | 0.779 | 0.811 |
| Word meanings (1,000) | 0.133 | 0.252 | 0.327 |
| Grammar correction (1,000) | 0.083 | 0.223 | 0.275 |
| Trivia (800) | — | 0.276 | 0.586 |
| News summarization, ROUGE-L (523) | ~0.07 | 0.162 | 0.152 |
| National exam (1,614, macro over 8 subjects) | — | 0.445 | 0.575 |
| 6-task average | — | 0.356 | 0.454 |
For context: Qwen3-4B-Instruct scores 0.212 and Llama-3.1-8B-Instruct 0.244 on the published version of this benchmark. This model beats the summarization score of EstLLM-8B and stays competitive with it on inflection, while being a 4B model trained on less than 2% of the Estonian pretraining data EstLLM used.
GGUF files
| File | Size | Notes |
|---|---|---|
gguf/gemma-3-4b-est-v1-q4_k_m.gguf |
2.49 GB | Recommended; runs on 6 GB+ phones |
gguf/gemma-3-4b-est-v1-q5_k_m.gguf |
2.83 GB | Slightly better quality |
gguf/gemma-3-4b-est-v1-f16.gguf |
7.77 GB | Reference |
Works out of the box with llama.cpp, Ollama, LM Studio, and on Android/iOS via PocketPal AI or ChatterUI (context up to 4096 recommended).
Usage
from transformers import AutoTokenizer, Gemma3ForConditionalGeneration
import torch
tok = AutoTokenizer.from_pretrained("Abdusin/gemma-3-4b-est-v1")
model = Gemma3ForConditionalGeneration.from_pretrained(
"Abdusin/gemma-3-4b-est-v1", torch_dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Tere! Räägi veidi Tartu linna ajaloost."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=300)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
Limitations
- Trained primarily for Estonian; English and other languages were not evaluated after training and may have degraded somewhat.
- World-knowledge tasks (trivia, national exams) still trail much larger Estonian models — the continued-pretraining corpus here is 161M tokens, which is small.
- Standard LLM caveats apply: it can hallucinate confidently, and it should not be used for medical, legal, or other high-stakes decisions.
License
This model inherits the Gemma Terms of Use. By using or redistributing it you agree to those terms.
Credits
- Base model: Gemma 3 by Google
- Training recipe follows the published Estonian adaptation work: EstLLM (teacher model), Llammas, and the Estonian LLM benchmark by TartuNLP / TalTech
- Datasets by TartuNLP, TalTechNLP, the Institute of the Estonian Language (EKI), ERR, and the FineWeb-2 project — thank you for making them public
- Downloads last month
- 966
Model tree for Abdusin/gemma-3-4b-est-v1
Base model
google/gemma-3-4b-pt