Safetensors
Tumbuka
English
llama
Mixture of Experts
tumbuka
bilingual
causal-lm

Mazgu_Small-T_130M

Mazgu_Small-T_130M is a bilingual Large Language Model (LLM) developed by Saints Studios, specifically optimized for Tumbuka (Tumbuka) and English.

Model Details

  • Architecture: Llama-based architecture
  • Parameters: ~130 Million
  • Vocabulary Size: 32,000 tokens
  • Context Length: 512 tokens
  • Training Data: A curated blend of Tumbuka translated datasets (Gutenberg, TinyStories) and English Light Novels.

Training Progress

The model was pretrained for 2,000 steps using a Tesla T4 GPU.

Step Training Loss Validation Loss
500 3.0454 3.4716
1000 2.9579 3.3821
1500 2.8874 3.3160
2000 2.8695 3.2677

Inference Examples

Tumbuka

Prompt: Mumutendeko
Output: Mumutendeko, ndipo imwe …atatu. Nchakusuzga yayi kuti mu…e na maghanoghano ghaheni chomene chifukwa cha inewakuyowoya unenesko; kweni pakuchita nthena pera nkhugomezgaNkhuwonga Chiuta wane ku…a wakukondwa, nangauli nkhawona nga ni umo viku…ira na chitima chikurukutemwa kwinuKweni icho chikachitika

Prompt: Kasi iyo wakati
Output: Kasi iyo wakati mu mtima wakhe?Ndipo wakafumba kuti, "Kasi nchivichi icho chikachitika pa nyengo iyo?Mukwenera kuluta ku malo ghanyake ndipo muzamukwera lu…iro nthena yayi.Mu mphindi zichoko waka pambere dazi lindambe, ndipo pamanyuma pakuti …afika pa mphaka ya nyumba

English

Prompt: Once upon a time
Output: Once upon a time with the sky. It seemed to be someone who'd decided to do much in such an enormous, but it didn't have anything more about this. As far as I was thinking of what, I could tell everything that I

Prompt: The people were
Output: The people were the same. Even if their owner had to give up, they wouldn't be able to gather everything in his heart as well. He felt a bit undead and looked at him with excitement."You did think we could do that?" the

How to Use

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("SaintsStudios/Mazgu_Small-T_130M")
model = AutoModelForCausalLM.from_pretrained("SaintsStudios/Mazgu_Small-T_130M")

prompt = "Mumutendeko"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0]))

Limitations

As a small-scale model (130M parameters), it may exhibit hallucination or BPE spacing artifacts. It is intended as a base for further fine-tuning on specific Tumbuka tasks.

Downloads last month
396
Safetensors
Model size
52.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train SaintsStudios/Mazgu_Small-T_130M