Krishi Sathi 2B (Sarvam) — Offline Kannada/English Farming Assistant for Karnataka

⚠️ Disclaimer — experimental / educational use only

Krishi Sathi is an experimental research and education project, provided as-is with no warranty of any kind. Large language models can produce incorrect or outdated information. This model is not professional agricultural advice; the authors accept no responsibility or liability for decisions made or losses incurred based on its outputs. Always verify any chemical, dosage, variety or practice with the product label and your local Raitha Samparka Kendra / agriculture department before acting.

Krishi Sathi ("farmer's companion") is a small bilingual farmer-advisory model built to run fully offline on low-end Android phones. This is the Sarvam-1-based release — chosen after a bake-off against Qwen3-1.7B because Sarvam-1's Indic pretraining and tokenizer give dramatically better Kannada fluency and ~3.8× fewer tokens per Kannada sentence (≈4× faster on-device generation).

  • Base model: sarvamai/sarvam-1 (2B, LoRA fine-tune, merged)
  • Quantized build: Q4_K_M GGUF (~1.5 GB), Llama-2 chat format ([INST]/<>)
  • Languages: Kannada (ಕನ್ನಡ) and English
  • Recommended sampling: temperature 0.3, repeat_penalty 1.1, stop on "[INST]"

License — read this first

Sarvam-1 and its derivatives (including this model) are governed by the Sarvam AI Research License: non-commercial and research use only. Commercial deployment requires a license from Sarvam AI. This model card distribution includes the license per its terms. This model is a modified (fine-tuned) derivative of Sarvam-1.

Training data

  1. Kisan Call Centre (KCC) Karnataka advisory queries — ~2.9k unique frequency-weighted query groups distilled from ~42k real farmer calls (2009–2023 public transcript data); terse operator notes rewritten into grounded answers by gemma-3-27b-it with dosages preserved from source.
  2. India-adapted general agronomy Q&A (filtered KisanVaani, teacher-rewritten).
  3. ~1,650 offline-honesty examples (teacher-generated, diverse phrasings) teaching refusal of live-data questions: prices, weather, scheme status, shop availability — with referral to e-NAM/APMC, Meghdoot/IMD, and Raitha Samparka Kendra channels.
  4. Kannada via IndicTrans2 (rotary en-indic-1B); training data decontaminated of any example that states current market prices.

Evaluated behavior (v4 smoke eval)

  • Kannada answers are fluent and practical; English answers competent.
  • Correctly refuses weather/price/scheme-status questions in both languages without inventing numbers.
  • Known issues: pesticide dosages in free-form English answers can drift from source values; occasional verbose safety-boilerplate tails; rare language-mixing at the start of a Kannada answer; occasionally suggests legacy chemicals whose registration has changed (e.g. Monocrotophos — banned for vegetables in India).

Deployment guidance — dosage safety

Do NOT surface free-form model dosages directly to farmers without the companion lookup table (dosage_table.json, shipped alongside): the intended app design is model-generated explanation + table-verified numbers. Always retain the "verify with your local Raitha Samparka Kendra" guidance. Not a substitute for soil testing, expert diagnosis, or official advisories.

Prompt format

Llama-2 style, single turn:

<s>[INST] <<SYS>>
{system prompt}
<</SYS>>

{farmer's question} [/INST]

Recommended system prompt: see Modelfile in this repo.

Roadmap

Telugu/Tamil/Marathi/Hindi via the same pipeline (Sarvam-1 covers all); teacher-verified dosage table v1; DPO polish for verbosity/stopping; pest-photo diagnosis via a small VLM.

Downloads last month
51
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ganmoor-ai-labs/krishi-sathi-sarvam-2b

Quantized
(19)
this model