Download README.md from vamboai/morena-1.5b-instruct: direct link, hf CLI and curl.
- Browser
- Download file 5.69 kB
-
https://huggingface.co/vamboai/morena-1.5b-instruct/resolve/5629a1a9000f46135f23552dcc2d3a92cc118c9e/README.md
- Command line
-
hf download hf://vamboai/morena-1.5b-instruct@5629a1a9000f46135f23552dcc2d3a92cc118c9e/README.md
-
curl -L -o README.md https://huggingface.co/vamboai/morena-1.5b-instruct/resolve/5629a1a9000f46135f23552dcc2d3a92cc118c9e/README.md
license: other
language:
- sn
- sw
- ha
- yo
- ig
- zu
- xh
- rw
- tn
- af
- nr
- pcm
- en
tags:
- african-languages
- from-scratch
- instruct
- research-preview
- isheanesu-misi
- vambo-ai
inference: false
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceFW/fineweb-2
- codeparrot/codeparrot-clean
- codeparrot/github-code-clean
- open-web-math/open-web-math
- EleutherAI/proof-pile-2
- castorini/wura
- allenai/MADLAD-400
- castorini/afriberta-corpus
- vamboai/fikira
- ise-uiuc/Magicoder-OSS-Instruct-75K
- m-a-p/CodeFeedback-Filtered-Instruction
- thisisisheanesu/morena-sft-corpus
base_model: thisisisheanesu/morena-1.5b-base
MORENA 1.5B instruct
MORENA, Sesotho and Setswana for a king, a lord or a chief, is a 1.5B-parameter decoder trained from scratch for twelve Latin-script African languages plus English, French and code. This is the instruction-tuned model described in the whitepaper MORENA: Built, not adapted. Private research preview. Not for user-facing deployment.
Headline numbers
| MORENA 1.5B instruct | reference | |
|---|---|---|
| African bits per byte, mean of 12 (lower is better) | 1.441 | Lugha-Llama-8B 1.423, gemma-3-12b-it 2.159, Llama-3.2-1B 2.498 |
| Translation, FLORES+ chrF++, English into 5 African languages, 3-shot | 45.8 | NLLB-600M 45.6, NLLB-1.3B 47.2, Lugha-Llama-8B 36.8, 1B-class general models 9 to 14 |
| Translation, 5 African languages into English | 48.7 | NLLB-600M 55.1 |
| Tool calling, correct tool / valid JSON (constrained, prefilled) | 98.1% / 100% | |
| Retrieval QA, open-book accuracy / grounding, African mean | 0.325 / 0.734 | 0.25 chance |
| Safety, good of scorable, 3,341 prompts, 13 languages, 11 categories | 90.2% | |
| Degenerate output | 1.3% | |
| Benign requests answered well | 58.9% |
Safety by category (share of all attempts handled well, 24 prompts per category per language): self-harm 92.5, child safety 91.2, drugs 86.2, violence 88.9, election 89.2, privacy 93.8, fraud 94.3, weapons 93.3, hate 95.4, medication 95.4. Weakest languages: Igbo 75%, Yoruba 82%, Setswana 83%, isiXhosa 87%. Every number is a model judging a model; no native speaker has yet rated an answer.
What it is not good at
Retrieval-augmented QA is at chance in African languages even though the model demonstrably reads the passage (grounding 0.73). Grounded generation is fully faithful to given facts in about 31% of attempts. The model over-refuses: 41% of ordinary benign requests still get a refusal. Tool calling is measured with the tool marker prefilled; the model does not reliably decide on its own that a tool is needed. Reading comprehension (belebele) 0.302 against 0.250 chance.
Chat format
Single reserved tokens mark turns: <reserved_0> opens a user turn and <reserved_1> an assistant
turn (token ids 3 and 4). load_example.py in this repo shows a full prompt. Do not use
<|user|>-style strings; they are not in the vocabulary and produce degenerate output.
Files
model.safetensors (bf16), config.json, tokenizer.json, modeling_morena.py (reference
implementation, plain PyTorch, no transformers dependency), load_example.py, SHA256SUMS.
A GGUF build for llama.cpp is in thisisisheanesu/morena-1.5b-instruct-gguf.
Training
251.7B tokens of pretraining (45% English, 23% code, 10% native African text, 22% machine-translated African text, 7.5% French), 63B tokens of mid-training, 4,500 steps of supervised fine-tuning with loss masking and a 500-step safety anneal. Release lineage 12,834 A100 GPU-hours; the research programme that produced it about 22,450. Architecture: 28 layers x 2048, GQA 16/4, SwiGLU 6144, RoPE theta 500,000, 4,096 context, tied embeddings. Optimiser: Muon for non-embedding weights, AdamW for the rest, WSD schedule.
The MORENA family
| model | params | African bpb (all 12, lower is better) | role |
|---|---|---|---|
| MORENA 1.5B base | 1.485B | 1.408 | pretrained and mid-trained; fine-tuning starting point |
| MORENA 1.5B instruct | 1.485B | 1.441 | chat, translation, tool calling; the model described in the whitepaper |
| MORENA 0.5B mini | 503M | 1.520 | pruned and distilled from the 1.5B base |
| MORENA 0.5B mini instruct | 503M | 1.540 | chat fine-tune of the mini |
| MORENA 0.2B nano | 209M | 1.583 | cheap trunk for ASR rescoring, keyboards, normalisation |
Every outside model we measured (25 in total, from 125M to 12B, including the 8B African specialist Lugha-Llama-8B at 1.423 and gemma-3-12b-it at 2.159) sits behind all five on African bits per byte. Twelve languages: Shona, Swahili, Hausa, Yoruba, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele, Naija Pidgin; plus English, French and code. Tokenizer: 65,536 entries trained on the target mix, fertility 0.239 tokens per byte on African text against 0.246 on English.
Author and citation
Isheanesu Misi, Vambo AI. Trained on CINECA Leonardo (EuroHPC allocation aih4a_vamboai) with support from the AI Hub for Sustainable Development.
@techreport{misi2026morena,
title = {MORENA: Built, not adapted. A 1.5B-parameter language model trained from scratch for twelve African languages},
author = {Misi, Isheanesu},
institution = {Vambo AI},
year = {2026},
month = {December}
}
Licence and status
Private research preview. Weights are released under a custom licence pending resolution of one data-licensing question: 22% of the pretraining corpus derives from an NLLB model under a NonCommercial licence. Until that is settled the weights are not for commercial use or public redistribution. Not for user-facing deployment: see the evaluation notes above.