Create training/README.md
Browse files- training/README.md +21 -0
training/README.md
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Training skeleton (Sanchari-S)
|
| 2 |
+
|
| 3 |
+
How-to (quick):
|
| 4 |
+
1. Prepare data:
|
| 5 |
+
- Create `data/all_texts.txt` as newline-separated documents or a single large text.
|
| 6 |
+
- Ensure `tokenizer/` has your trained tokenizer files: `sanchari_spm.model`, `sanchari_spm.vocab` (or use the placeholders until you generate them).
|
| 7 |
+
|
| 8 |
+
2. Install deps (Colab / local):
|
| 9 |
+
pip install transformers datasets sentencepiece accelerate
|
| 10 |
+
|
| 11 |
+
3. Run training (example local):
|
| 12 |
+
python training/train.py --config training/config_s.json --tokenizer_dir tokenizer --data_file data/all_texts.txt --output_dir outputs/sanchari-s
|
| 13 |
+
|
| 14 |
+
4. Inspect outputs:
|
| 15 |
+
- Model + tokenizer saved in `outputs/sanchari-s`
|
| 16 |
+
- Logs printed to console
|
| 17 |
+
|
| 18 |
+
Notes:
|
| 19 |
+
- This script is intentionally minimal for seed-stage proof of concept.
|
| 20 |
+
- For larger runs use DeepSpeed/Accelerate and multi-GPU clusters.
|
| 21 |
+
- Before any investor checkpoint release, run safety & PII checks.
|