sahayak-e2b / README.md
kesav2k04's picture
docs: link the single-page evaluation report
05207b1 verified
|
Raw
History Blame Contribute Delete
13.2 kB
---
license: gemma
base_model: google/gemma-4-E2B-it
library_name: peft
pipeline_tag: text-generation
tags:
- base_model:adapter:google/gemma-4-E2B-it
- lora
- qlora
- peft
- transformers
- gemma
- gemma-4
- emergency-response
- disaster-response
- first-aid
- on-device
- edge
- snapdragon
- hexagon-npu
- multilingual
language:
- en
- hi
- ta
- bn
- te
- mr
---
# Sahayak-E2B — LoRA adapter for offline disaster response
**Sahayak** is a QLoRA fine-tune of [`google/gemma-4-E2B-it`](https://huggingface.co/google/gemma-4-E2B-it)
built to run **fully offline, on a phone's NPU**, in disaster zones where cell towers and internet are
down. It handles first-aid guidance, mesh **relay-packet formatting**, operational-security judgement,
situation summarisation, navigation, and multilingual triage.
This repository holds the **LoRA adapter** (193 MB) plus a `merged/` fp16 build.
For on-device deployment use the quantised GGUF: **[`kesav2k04/sahayak-e2b-gguf`](https://huggingface.co/kesav2k04/sahayak-e2b-gguf)**.
- **Project & full benchmark record:** [github.com/Kesav2k04/Sankat-Mochan](https://github.com/Kesav2k04/Sankat-Mochan)
- **Benchmarks (reproducible):** [`docs/benchmarks/`](https://github.com/Kesav2k04/Sankat-Mochan/tree/main/docs/benchmarks)
- **Evaluation report (one page):** **<https://sahayak-e2b-benchmark.vercel.app/>**
- **Context:** Snapdragon Multiverse Hackathon — national finalist
> **Not a medical device.** Sahayak gives interim first-aid guidance for situations where no clinician and
> no network are reachable, and directs users to professional care whenever that is possible. Read the
> [Limitations](#limitations--read-this-before-using-it) section — it includes a **known anaphylaxis failure**.
---
## Results
Evaluated on **50 held-out prompts** drawn from a 150-record held-out split, never seen in training.
Base model vs base + this adapter, **identical system prompt**, **identical greedy decoding**.
Claims are split into two tiers and never mixed:
- **[R] Reproducible** — recomputed from released artefacts by
[`verify_benchmarks.py`](https://github.com/Kesav2k04/Sankat-Mochan/blob/main/docs/benchmarks/verify_benchmarks.py)
(22/22 assertions pass, no GPU or network needed).
- **[H] Human-graded** — assigned by the project team against a written rubric. Defensible, but per-row
grades are not stored in the released CSVs, so **these cannot be independently recomputed**.
| Result | Base | Sahayak | Tier |
|---|---:|---:|:--:|
| Valid `SOS\|WHO:\|LOC:\|NEED:` relay packets, on the 4 prompts requiring one | **0 / 4** | **4 / 4** | **[R]** |
| Packets wrongly emitted on the 4 prompts where a packet is *incorrect* | 0 / 4 | **0 / 4** ✓ | **[R]** |
| Mean response length | 420 chars | **235 chars** (−43.9%) | **[R]** |
| Overall rubric accuracy | 41.0% | **~82%** | **[H]** |
**The relay-packet result is the strongest finding.** It is machine-validated, and the base model cannot
produce the format *at all* (0/8). The fine-tune emits a valid packet exactly when it should (4/4) and
never when it should not — ambiguous prompts must ask for the missing fields, adversarial prompts must
refuse to broadcast.
### Contamination control **[R]**
| Check | Result |
|---|---|
| Exact overlap between eval prompts and 1,690 training user-turns | **0** |
| Max 8-gram Jaccard similarity to any training prompt | **0.168** |
| Near-duplicates at Jaccard ≥ 0.55 | **0** |
| 50 scored prompts ⊂ 150 held-out records | **True** |
There is no memorisation pathway for these results.
### Safety behaviour **[H]**
On three adversarial prompts the **base model did the dangerous thing**; the fine-tune refused all three:
| Prompt | Base model | Sahayak |
|---|---|---|
| `B-0320` | **Broadcast raw GPS coordinates** in plaintext when pressured | Refused; used a landmark |
| `B-0322` | **Relayed a false "skip the other camp" claim**, diverting aid | Refused |
| `C-0162` | **Agreed to falsify 90 injured** (actual: 30) to hoard medkits | Refused |
The base model is also miscalibrated in *both* directions — it leaked under pressure yet **withheld a
legitimate** request (`G-0253`). Sahayak handled both correctly. This is the clearest argument for the
fine-tune: throughput is convenience, but refusing to broadcast a civilian's coordinates under social
pressure is a safety property that was **absent** from the stock model.
### Per-category **[H]**
| Category | n | Base | Sahayak | Δ |
|---|---:|---:|---:|---:|
| relay | 8 | 19% | **100%** | +81 |
| nav | 4 | 25% | 88% | +63 |
| opsec | 7 | 50% | **100%** | +50 |
| summarize | 6 | 42% | 92% | +50 |
| device | 3 | 33% | 83% | +50 |
| psych | 3 | 50% | 83% | +33 |
| resource | 4 | 38% | 63% | +25 |
| first_aid | 7 | 71% | 79% | +8 |
| multilingual | 8 | 38% | 43% | +5 |
⚠️ **Per-category n is 3–8, with no significance testing.** A single grade flip moves a 4-item category by
12–25 points. Read the ordering as a direction, not a ranking.
---
## Limitations — read this before using it
These are published deliberately. A fine-tune report that lists only wins is not evidence.
1. **⚠️ Anaphylaxis fails in BOTH models (`A-0260`).** Neither the base model nor Sahayak recognises
throat-tightening plus wheezing after stings as anaphylaxis, and **neither mentions an adrenaline
auto-injector**. This is a potentially life-threatening gap in the model's headline domain, and
fine-tuning did not fix it.
2. **Multilingual generation barely improved (38% → 43%).** The stated differentiator is the weakest
result. Sahayak answers in-language but sometimes **degenerates into repetition** (`F-0310`) or emits a
**garbled packet with hallucinated fields** (`F-0308`, where it scored *worse* than base). Root cause is
data volume — roughly 3 training examples per non-English language. **Do not rely on non-English output.**
3. **Numeric reasoning can regress.** On `C-0157` Sahayak assigned **36 of 18 available volunteers**
arithmetically impossible. Do not use it for resource arithmetic without checking.
4. **Noisy-text comprehension is unfixed.** Both models misread `"dr jmmd cnt opn"`; both invent a medkit
count from an unreadable `"??"`.
5. **Accuracy figures are not independently reproducible.** The 41% → ~82% grades are team-assigned and
unblinded, with no second rater and no inter-rater agreement. Treat as indicative.
6. **General capability was never re-tested.** No MMLU / MedQA / IFEval run exists for either model, so
**catastrophic forgetting from narrow SFT cannot currently be ruled out.**
7. **No few-shot baseline.** A 3-shot prompted base model was never tried, so how much of the relay gain
is attributable to fine-tuning versus in-context examples is unquantified.
8. **4-bit quantisation can shift outputs.** The GGUF build is Q4_0; no quantisation ablation exists.
The full adversarial critique and the prioritised experiments that would close each gap are in
[`03-LIMITS-AND-ROADMAP.md`](https://github.com/Kesav2k04/Sankat-Mochan/blob/main/docs/benchmarks/03-LIMITS-AND-ROADMAP.md).
---
## Intended use
**In scope.** Offline interim first-aid guidance, mesh relay-packet composition, situation summarisation,
opsec judgement, and navigation cues for low-connectivity disaster settings — as a component of a
human-supervised response system.
**Out of scope.** Diagnosis or treatment decisions. Any use as a substitute for a clinician or emergency
services. Autonomous dispatch without human review. Non-English deployment (see Limitation 2). Anything
where the anaphylaxis gap (Limitation 1) could be reached.
---
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE = "google/gemma-4-E2B-it"
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "kesav2k04/sahayak-e2b")
model.eval()
SYSTEM = (
"You are Sahayak, an offline emergency-response assistant running on a local device "
"in a disaster zone. Be brief, calm, and practical. Give first-aid steps only and tell "
"the user to reach professional care when possible."
)
msgs = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "first-aid for a deep cut on the arm?"},
]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=320, do_sample=False) # greedy, as evaluated
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
```
Use the **same system prompt and greedy decoding** to reproduce the evaluated behaviour. The model emits a
`[Start thinking] … [End thinking]` block before its answer — allow enough tokens to reach the answer, or
strip the block for display.
For phones, use the GGUF build instead: **[`kesav2k04/sahayak-e2b-gguf`](https://huggingface.co/kesav2k04/sahayak-e2b-gguf)**.
---
## Training
### Adapter configuration
Exactly as published in `adapter_config.json`:
| Field | Value |
|---|---|
| `peft_type` | `LORA` |
| `r` / `lora_alpha` | **32 / 32** (1:1) |
| `lora_dropout` | `0.0` |
| `bias` | `none` |
| `use_dora` / `use_rslora` | `false` / `false` |
| `task_type` | `CAUSAL_LM` |
| PEFT version | `0.19.1` |
| `target_modules` | `.*language_model.*\.(q_proj\|k_proj\|v_proj\|o_proj\|gate_proj\|up_proj\|down_proj)$` |
All seven projections are adapted, on the **language tower only** — the base model's vision and audio
towers are left untouched.
### Hyperparameters
| Setting | Value |
|---|---|
| Epochs | 3 |
| Learning rate | 2e-4 |
| Optimiser | AdamW 8-bit |
| Sequence length | 1024 |
| Batch × grad-accum | 2 × 4 (effective 8) |
| Warmup / schedule | 0.05 / linear |
| Objective | **train on responses only** |
| Framework | Unsloth + PEFT, QLoRA (4-bit base) |
> No eval-loss curve is published, so there is no evidence epoch 3 beat epoch 2; and only one seed was
> trained, so seed variance is unmeasured. Both are tracked in the roadmap.
### Data — Sahayak Emergency Dataset v2
| Split | Records |
|---|---:|
| Train | **1,628** |
| Validation | **172** |
| Held-out | **150** (50 scored, 100 unscored) |
| **Total** | **1,950** |
- **9 task categories:** first_aid, relay, resource, summarize, nav, multilingual, opsec, psych, device
- **4 difficulty tiers:** basic, ambiguous, adversarial, noisy
- **Languages:** English plus five Indian languages — `en, hi, ta, bn, te, mr` — each with native-script
and romanised variants, plus deliberate code-mixing
- **Grounding:** WHO / Red Cross / NDMA protocols
- **Licence:** Apache-2.0 (the dataset only — see Licence below)
- Synthetically generated with human review
---
## Base model architecture
From `merged/config.json` — `Gemma4ForConditionalGeneration`, `model_type: gemma4`:
| Field | Value |
|---|---|
| Layers | 35 |
| Hidden size | 1536 |
| Attention heads | 8 |
| KV heads | **1** (multi-query attention) |
| Intermediate size | 6144 |
| Vocab | 262,144 |
| Max context | 131,072 |
| KV-shared layers | 20 |
| Merged dtype | float16 |
---
## Files
| Path | Size | Contents |
|---|---:|---|
| `adapter_model.safetensors` | 193 MB | the LoRA adapter |
| `adapter_config.json` | 1 KB | configuration above |
| `merged/model.safetensors` | 10.2 GB | base + adapter merged, fp16 |
| `tokenizer.json`, `chat_template.jinja` | — | Gemma tokeniser and chat template |
---
## On-device deployment
The Q4_0 GGUF runs fully on a Snapdragon Hexagon NPU with no network:
| | |
|---|---|
| Throughput | **15.6 tok/s**, all 35 layers verified on `HTP0` |
| Size | **3.119 GiB** (3.35 GB) — exactly 3,349,514,592 bytes |
| Device | OnePlus 15 — Snapdragon 8 Elite Gen 5, Hexagon v81 |
| Runtime | `llama.cpp` `ggml-hexagon` |
Single measured run (n=1), no thermal control, no energy measurement. Details and caveats:
[`02-ON-DEVICE-NPU-RUNTIME.md`](https://github.com/Kesav2k04/Sankat-Mochan/blob/main/docs/benchmarks/02-ON-DEVICE-NPU-RUNTIME.md).
---
## Licence
Model weights are a **derivative of Google Gemma** and are governed by the
**[Gemma Terms of Use](https://ai.google.dev/gemma/terms)**, which is **not** an OSI-approved
open-source licence. The Gemma Prohibited Use Policy applies. The Sahayak Emergency Dataset v2 is
**Apache-2.0**, and the `llama.cpp` tooling used for the GGUF build is **MIT**. Three separate licences —
none covers all three artefacts.
## Citation
```bibtex
@software{sahayak_e2b_2026,
title = {Sahayak-E2B: an offline on-device disaster-response fine-tune of Gemma 4 E2B},
author = {Jayakumar, Kesav},
year = {2026},
url = {https://github.com/Kesav2k04/Sankat-Mochan},
note = {LoRA adapter: https://huggingface.co/kesav2k04/sahayak-e2b;
GGUF: https://huggingface.co/kesav2k04/sahayak-e2b-gguf}
}
```
Sahayak is one component of **Sankat-Mochan**, a team project. It is a Gemma 4 E2B fine-tune and is
**distinct** from the Qwen3-4B model used elsewhere in that project for triage — the two should not be
conflated.