File size: 2,896 Bytes
6c37197
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
# Sanchari β€” Technical Blueprint (Full β€” Investor Grade)

> Version: v0.1  
> Purpose: Complete technical blueprint to develop Sanchari-S β†’ Sanchari-M β†’ Sanchari-L  
> Target audience: engineers, infrastructure teams, investors

---

## Summary (one line)
Build a practical, India-focused multilingual instruction-following LM family (S: ~200–350M, M: ~1–3B, L: 7B+) using modern efficient training (PyTorch + DeepSpeed/Accelerate + FlashAttention), with explicit data provenance, safety audits, and production deployment targets.

---

## 1. Design principles (what we deliver and why)
1. **Practicality first** β€” Sanchari-S must be cheap to train and fast to run for API/mobile use.  
2. **Indian language competence** β€” prioritize English (Indian), Hindi, Telugu, mixed-script.  
3. **Safety & governance** β€” every training stage includes PII scrubbing and red-team testing before any checkpoint release.  
4. **Modular pipeline** β€” tokenizer, preprocessing, training config, adapter-based instruction tuning.  
5. **Deliverables at each phase** β€” private checkpoints for investors (under NDA), evaluation reports, HF repo + demo.

---

## 2. Model architecture & approach

### 2.1 Base model family
- **Sanchari-S:** ~200–350M parameters. Decoder-only transformer (GPT-like). Primary use: fast inference, API & mobile clients.
- **Sanchari-M:** ~1–3B parameters. Better instruction following and multi-turn coherence.
- **Sanchari-L:** ~7B+ parameters. Full foundation for enterprise applications.

### 2.2 Model topology (recommended)
- **Type:** Decoder-only transformer.
- **Layer scaling (example recipes):**
  - Sanchari-S: 24 layers Γ— 8 heads Γ— 2048 hidden (example β‰ˆ 300M)
  - Sanchari-M: 36 layers Γ— 16 heads Γ— 4096 hidden (example β‰ˆ 1.3B)
  - Sanchari-L: 48–80 layers Γ— 32 heads Γ— 6144 hidden (approx 7B)
- **Attention:** FlashAttention2 compatible (for GPU memory and speed).
- **Norms and embeddings:** RMSNorm or LayerNorm, rotary positional embeddings (RoPE) for stable long-range.
- **Efficiency:** Provide LoRA adapters and QLoRA options for cost-effective instruction tuning.

> Rationale: decoder-only is industry standard for instruction-following and easier to deploy as an API.

---

## 3. Tokenizer & text processing
### 3.1 Tokenizer strategy
- **Tokenizer:** SentencePiece Unigram or BPE with ~50k vocabulary, trained on mixed Indic + English corpus.
- **Features required:**
  - Mixed script normalization (normalize Unicode, NFKC)
  - Preserve whitespace tokens for code-like text
  - Subword segmentation for Indic scripts
- **Tooling:** `sentencepiece` or `tokenizers` (Hugging Face), with training script.

### 3.2 Tokenizer commands (example)
```bash
# install
pip install sentencepiece tokenizers

# train sentencepiece
spm_train --input=data/all_texts.txt --model_prefix=sanchari_spm --vocab_size=50000 --model_type=unigram