Mazgu_Small-T_130M
Mazgu_Small-T_130M is a bilingual Large Language Model (LLM) developed by Saints Studios, specifically optimized for Tumbuka (Tumbuka) and English.
Model Details
- Architecture: Llama-based architecture
- Parameters: ~130 Million
- Vocabulary Size: 32,000 tokens
- Context Length: 512 tokens
- Training Data: A curated blend of Tumbuka translated datasets (Gutenberg, TinyStories) and English Light Novels.
Training Progress
The model was pretrained for 2,000 steps using a Tesla T4 GPU.
| Step | Training Loss | Validation Loss |
|---|---|---|
| 500 | 3.0454 | 3.4716 |
| 1000 | 2.9579 | 3.3821 |
| 1500 | 2.8874 | 3.3160 |
| 2000 | 2.8695 | 3.2677 |
Inference Examples
Tumbuka
Prompt: Mumutendeko
Output: Mumutendeko, ndipo imwe
atatu. Nchakusuzga yayi kuti mu
e na maghanoghano ghaheni chomene chifukwa cha inewakuyowoya unenesko; kweni pakuchita nthena pera nkhugomezgaNkhuwonga Chiuta wane ku
a wakukondwa, nangauli nkhawona nga ni umo viku
ira na chitima chikurukutemwa kwinuKweni icho chikachitika
Prompt: Kasi iyo wakati
Output: Kasi iyo wakati mu mtima wakhe?Ndipo wakafumba kuti, "Kasi nchivichi icho chikachitika pa nyengo iyo?Mukwenera kuluta ku malo ghanyake ndipo muzamukwera lu
iro nthena yayi.Mu mphindi zichoko waka pambere dazi lindambe, ndipo pamanyuma pakuti
afika pa mphaka ya nyumba
English
Prompt: Once upon a time
Output: Once upon a time with the sky. It seemed to be someone who'd decided to do much in such an enormous, but it didn't have anything more about this. As far as I was thinking of what, I could tell everything that I
Prompt: The people were
Output: The people were the same. Even if their owner had to give up, they wouldn't be able to gather everything in his heart as well. He felt a bit undead and looked at him with excitement."You did think we could do that?" the
How to Use
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("SaintsStudios/Mazgu_Small-T_130M")
model = AutoModelForCausalLM.from_pretrained("SaintsStudios/Mazgu_Small-T_130M")
prompt = "Mumutendeko"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0]))
Limitations
As a small-scale model (130M parameters), it may exhibit hallucination or BPE spacing artifacts. It is intended as a base for further fine-tuning on specific Tumbuka tasks.
- Downloads last month
- 396