--- license: gemma base_model: google/gemma-3-4b-pt tags: - estonian - gemma - llama.cpp - continued-pretraining - summarization language: - et - en datasets: - HuggingFaceFW/fineweb-2 - tartuNLP/magpie-gemma-3-12b-it-100k-et - TalTechNLP/word_meanings_et - TalTechNLP/inflection_et - TalTechNLP/grammar_et - TalTechNLP/EstQA - TalTechNLP/ERRnews pipeline_tag: text-generation --- # Gemma 3 4B Estonian (v1) A Gemma 3 4B base model adapted for the **Estonian language**. The stock Gemma 3 4B knows surprisingly little Estonian; this model was trained to fix that while staying small enough to run comfortably on a phone. ## What was done Two stages, both with QLoRA on a single consumer GPU (RTX 5070 12 GB, ~26 hours total): 1. **Continued pretraining** on 161.5M tokens of Estonian web text (the Estonian subset of [FineWeb-2](https://huggingface.co/datasets/HuggingFaceFW/fineweb-2)), packed to 2048-token blocks, one epoch, learning rate 5e-5. 2. **Supervised fine-tuning** on ~30.5k examples (one epoch, lr 1e-4): - 15k Estonian instructions from [tartuNLP/magpie-gemma-3-12b-it-100k-et](https://huggingface.co/datasets/tartuNLP/magpie-gemma-3-12b-it-100k-et) - ~2.9k responses distilled from the open [EstLLM-8B](https://huggingface.co/tartuNLP/Llama-3.1-EstLLM-8B-Instruct-1125) model on native Estonian prompts (dictionary explanations, inflection, reading comprehension) - ~11k pairs built from public Estonian language resources: EKI dictionary definitions → word ([word_meanings_et](https://huggingface.co/datasets/TalTechNLP/word_meanings_et)), noun-phrase inflection ([inflection_et](https://huggingface.co/datasets/TalTechNLP/inflection_et)), grammar correction with gold targets ([grammar_et](https://huggingface.co/datasets/TalTechNLP/grammar_et)), extractive QA ([EstQA](https://huggingface.co/datasets/TalTechNLP/EstQA)), and news summarization ([ERRnews](https://huggingface.co/datasets/TalTechNLP/ERRnews)) ## Evaluation Measured with the official [Estonian LLM benchmark harness](https://github.com/taltechnlp/lm-eval-harness-tasks-estonian) (LREC 2026, Lillepalu & Alumäe), full test sets, zero-shot with the chat template. Published numbers for reference models are from the benchmark paper (arXiv:2510.21193). | Task (exact match) | Gemma-3-4B-it (base) | **This model** | EstLLM-8B (published) | |---|---|---|---| | Inflection (1,400) | 0.107 | **0.779** | 0.811 | | Word meanings (1,000) | 0.133 | **0.252** | 0.327 | | Grammar correction (1,000) | 0.083 | **0.223** | 0.275 | | Trivia (800) | — | **0.276** | 0.586 | | News summarization, ROUGE-L (523) | ~0.07 | **0.162** | 0.152 | | National exam (1,614, macro over 8 subjects) | — | **0.445** | 0.575 | | **6-task average** | — | **0.356** | 0.454 | For context: Qwen3-4B-Instruct scores 0.212 and Llama-3.1-8B-Instruct 0.244 on the published version of this benchmark. This model beats the summarization score of EstLLM-8B and stays competitive with it on inflection, while being a 4B model trained on less than 2% of the Estonian pretraining data EstLLM used. ## GGUF files | File | Size | Notes | |---|---|---| | `gguf/gemma-3-4b-est-v1-q4_k_m.gguf` | 2.49 GB | Recommended; runs on 6 GB+ phones | | `gguf/gemma-3-4b-est-v1-q5_k_m.gguf` | 2.83 GB | Slightly better quality | | `gguf/gemma-3-4b-est-v1-f16.gguf` | 7.77 GB | Reference | Works out of the box with llama.cpp, Ollama, LM Studio, and on Android/iOS via PocketPal AI or ChatterUI (context up to 4096 recommended). ## Usage ```python from transformers import AutoTokenizer, Gemma3ForConditionalGeneration import torch tok = AutoTokenizer.from_pretrained("Abdusin/gemma-3-4b-est-v1") model = Gemma3ForConditionalGeneration.from_pretrained( "Abdusin/gemma-3-4b-est-v1", torch_dtype=torch.bfloat16, device_map="auto") messages = [{"role": "user", "content": "Tere! Räägi veidi Tartu linna ajaloost."}] inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) out = model.generate(inputs, max_new_tokens=300) print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True)) ``` ## Limitations - Trained primarily for Estonian; English and other languages were not evaluated after training and may have degraded somewhat. - World-knowledge tasks (trivia, national exams) still trail much larger Estonian models — the continued-pretraining corpus here is 161M tokens, which is small. - Standard LLM caveats apply: it can hallucinate confidently, and it should not be used for medical, legal, or other high-stakes decisions. ## License This model inherits the [Gemma Terms of Use](https://ai.google.dev/gemma/terms). By using or redistributing it you agree to those terms. ## Credits - Base model: [Gemma 3](https://huggingface.co/google/gemma-3-4b-pt) by Google - Training recipe follows the published Estonian adaptation work: [EstLLM](https://arxiv.org/abs/2603.02041) (teacher model), [Llammas](https://arxiv.org/abs/2310.03389), and the [Estonian LLM benchmark](https://arxiv.org/abs/2510.21193) by TartuNLP / TalTech - Datasets by TartuNLP, TalTechNLP, the Institute of the Estonian Language (EKI), ERR, and the FineWeb-2 project — thank you for making them public