|
Download architecture.md from Mike369williams/Sanchari: direct link, hf CLI and curl.
- Browser
- Download file 2.9 kB
-
https://huggingface.co/Mike369williams/Sanchari/resolve/main/architecture.md
- Command line
-
hf download hf://Mike369williams/Sanchari/architecture.md
-
curl -L -o architecture.md https://huggingface.co/Mike369williams/Sanchari/resolve/main/architecture.md
2.9 kB
| # Sanchari β Technical Blueprint (Full β Investor Grade) | |
| > Version: v0.1 | |
| > Purpose: Complete technical blueprint to develop Sanchari-S β Sanchari-M β Sanchari-L | |
| > Target audience: engineers, infrastructure teams, investors | |
| --- | |
| ## Summary (one line) | |
| Build a practical, India-focused multilingual instruction-following LM family (S: ~200β350M, M: ~1β3B, L: 7B+) using modern efficient training (PyTorch + DeepSpeed/Accelerate + FlashAttention), with explicit data provenance, safety audits, and production deployment targets. | |
| --- | |
| ## 1. Design principles (what we deliver and why) | |
| 1. **Practicality first** β Sanchari-S must be cheap to train and fast to run for API/mobile use. | |
| 2. **Indian language competence** β prioritize English (Indian), Hindi, Telugu, mixed-script. | |
| 3. **Safety & governance** β every training stage includes PII scrubbing and red-team testing before any checkpoint release. | |
| 4. **Modular pipeline** β tokenizer, preprocessing, training config, adapter-based instruction tuning. | |
| 5. **Deliverables at each phase** β private checkpoints for investors (under NDA), evaluation reports, HF repo + demo. | |
| --- | |
| ## 2. Model architecture & approach | |
| ### 2.1 Base model family | |
| - **Sanchari-S:** ~200β350M parameters. Decoder-only transformer (GPT-like). Primary use: fast inference, API & mobile clients. | |
| - **Sanchari-M:** ~1β3B parameters. Better instruction following and multi-turn coherence. | |
| - **Sanchari-L:** ~7B+ parameters. Full foundation for enterprise applications. | |
| ### 2.2 Model topology (recommended) | |
| - **Type:** Decoder-only transformer. | |
| - **Layer scaling (example recipes):** | |
| - Sanchari-S: 24 layers Γ 8 heads Γ 2048 hidden (example β 300M) | |
| - Sanchari-M: 36 layers Γ 16 heads Γ 4096 hidden (example β 1.3B) | |
| - Sanchari-L: 48β80 layers Γ 32 heads Γ 6144 hidden (approx 7B) | |
| - **Attention:** FlashAttention2 compatible (for GPU memory and speed). | |
| - **Norms and embeddings:** RMSNorm or LayerNorm, rotary positional embeddings (RoPE) for stable long-range. | |
| - **Efficiency:** Provide LoRA adapters and QLoRA options for cost-effective instruction tuning. | |
| > Rationale: decoder-only is industry standard for instruction-following and easier to deploy as an API. | |
| --- | |
| ## 3. Tokenizer & text processing | |
| ### 3.1 Tokenizer strategy | |
| - **Tokenizer:** SentencePiece Unigram or BPE with ~50k vocabulary, trained on mixed Indic + English corpus. | |
| - **Features required:** | |
| - Mixed script normalization (normalize Unicode, NFKC) | |
| - Preserve whitespace tokens for code-like text | |
| - Subword segmentation for Indic scripts | |
| - **Tooling:** `sentencepiece` or `tokenizers` (Hugging Face), with training script. | |
| ### 3.2 Tokenizer commands (example) | |
| ```bash | |
| # install | |
| pip install sentencepiece tokenizers | |
| # train sentencepiece | |
| spm_train --input=data/all_texts.txt --model_prefix=sanchari_spm --vocab_size=50000 --model_type=unigram |