--- base_model: unsloth/gemma-4-E2B-it library_name: gguf license: gemma pipeline_tag: text-generation language: - en - zh - ar tags: - gemma - gemma-4 - gguf - quantized - q4-k-m - llama-cpp - ollama - fine-tuned - rag - offlineaid - australian-consumer-safety - anti-scam - disaster-response datasets: - helenkwok/offlineaid --- # OfflineAid — Gemma 4 E2B fine-tune (Q4_K_M GGUF) Q4_K_M-quantized GGUF of the [OfflineAid](https://github.com/helenkwok/offlineaid) Stage-1 fine-tune (E2B variant). Drop-in for `llama.cpp` and Ollama. Smaller sibling of [`helenk/gemma-4-E4B-finetune-GGUF`](https://huggingface.co/helenk/gemma-4-E4B-finetune-GGUF) — same training corpus and recipe, lower memory footprint. - **Base:** [`unsloth/gemma-4-E2B-it`](https://huggingface.co/unsloth/gemma-4-E2B-it) - **LoRA:** [`helenk/gemma-4-E2B-lora`](https://huggingface.co/helenk/gemma-4-E2B-lora) — Unsloth fine-tune on Kaggle T4 - **Merged fp16 source:** [`helenk/gemma-4-E2B-finetune`](https://huggingface.co/helenk/gemma-4-E2B-finetune) - **Quantize chain:** `peft.merge_and_unload` → `llama.cpp/convert_hf_to_gguf.py` → `llama-quantize Q4_K_M` - **File size:** ~3.4 GB ## Use with Ollama ``` ollama pull hf.co/helenk/gemma-4-E2B-finetune-GGUF ``` Or via a local `Modelfile`: ``` FROM /path/to/gemma-4-E2B-offlineaid-Q4_K_M.gguf RENDERER gemma4 PARSER gemma4 PARAMETER num_ctx 32768 PARAMETER stop "" PARAMETER temperature 0.0 ``` ``` ollama create offlineaid-e2b -f Modelfile ollama run offlineaid-e2b ``` ## Use with llama.cpp ``` ./llama-cli \ -m gemma-4-E2B-offlineaid-Q4_K_M.gguf \ -p "Answer in Simplified Chinese.\n\nQUESTION: ..." \ --temp 0.0 -n 256 ``` ## Tier A held-out eval Tier A held-out numbers for the writeup are reported on the larger E4B variant — see [`helenk/gemma-4-E4B-finetune-GGUF`](https://huggingface.co/helenk/gemma-4-E4B-finetune-GGUF) for the full methodology and the headline result (fine-tune + RAG raises overall multilingual format-OK from 53.2% → 70.3% vs stock + RAG, with AR format-OK +32.4 pp / 2.5×). Both E2B and E4B share identical training data, recipe, and quantization chain. ## Intended use Lower-memory variant of the OfflineAid Stage-1 fine-tune. Pixel 7 production deployment uses *stock* Gemma 4 E2B + retrieval, not this fine-tune; see [the project writeup](https://github.com/helenkwok/offlineaid) for the architectural rationale. ## License Inherits Google's [Gemma Terms of Use](https://ai.google.dev/gemma/terms). Training data ([`helenkwok/offlineaid`](https://huggingface.co/datasets/helenkwok/offlineaid)) is CC-BY-4.0. ## Sibling repos - Merged fp16 safetensors (~9.5 GB): [`helenk/gemma-4-E2B-finetune`](https://huggingface.co/helenk/gemma-4-E2B-finetune) - LoRA adapter (~30 MB): [`helenk/gemma-4-E2B-lora`](https://huggingface.co/helenk/gemma-4-E2B-lora) - Larger E4B variant: [`helenk/gemma-4-E4B-finetune-GGUF`](https://huggingface.co/helenk/gemma-4-E4B-finetune-GGUF)