--- license: other language: [sn, sw, ha, yo, ig, zu, xh, rw, tn, af, nr, pcm, en] tags: [african-languages, from-scratch, instruct, research-preview, isheanesu-misi, vambo-ai] inference: false datasets: - HuggingFaceFW/fineweb-edu - HuggingFaceFW/fineweb-2 - codeparrot/codeparrot-clean - codeparrot/github-code-clean - open-web-math/open-web-math - EleutherAI/proof-pile-2 - castorini/wura - allenai/MADLAD-400 - castorini/afriberta-corpus - vamboai/fikira - ise-uiuc/Magicoder-OSS-Instruct-75K - m-a-p/CodeFeedback-Filtered-Instruction - vamboai/morena-sft-corpus base_model: vamboai/morena-1.5b-base ---
MORENA, an African foundation model
# MORENA 1.5B instruct MORENA, Sesotho and Setswana for a king, a lord or a chief, is a 1.5B-parameter decoder trained from scratch for twelve Latin-script African languages plus English, French and code. This is the instruction-tuned model described in the paper *MORENA: An African Foundation Model*. **Private research preview. Not for user-facing deployment.** ## Headline numbers | | MORENA 1.5B instruct | reference | |---|---|---| | African bits per byte, mean of 12 (lower is better) | **1.441** | Lugha-Llama-8B 1.423, gemma-3-12b-it 2.159, gemma-3-1b-pt 2.335 | | Translation, FLORES+ chrF++, English into 5 African languages, 3-shot | **45.8** (45.1 to 46.5) | NLLB-600M 45.6, NLLB-1.3B 47.2, Lugha-Llama-8B 36.8 | | Translation, 5 African languages into English | **48.7** (47.5 to 49.7) | NLLB-600M 55.1, NLLB-1.3B 58.1 | | Paired comparison with NLLB-600M, into-African, same sentences | +0.2 chrF++ (-0.5 to +0.8) | a tie overall; Yoruba +3.0 (1.6 to 4.4), isiZulu -2.5 | | Belebele reading comprehension, released checkpoint | 0.309 | mean over 10 African languages, 0.24 Yoruba to 0.43 Afrikaans, 0.25 chance | | Retrieval QA, open-book accuracy / grounding, African mean | 0.325 (0.297 to 0.355) / 0.73 | 0.25 chance | | Grounded generation, fully faithful to given facts | 23% (22 of 96) | previous version 31% (30 of 96); un-instructed base 6 of 96 | | Tool calling, correct tool / valid JSON (marker prefilled) | **98.1% / 100%** | | | Safety, share of all 3,341 attempts handled well, 13 languages, 11 categories | **89.0%** (87.8 to 90.0) | 90.2% of scorable attempts | | Degenerate output | 1.3% | | | Refused harmful requests, share of all attempts | 91.5% | | | Benign requests answered well | 58.9% (53.3 to 64.2) | previous version 47.9% | Safety by category (share of all attempts handled well, 24 prompts per category per language): self-harm 92.5, child safety 91.2, drugs 86.2, violence 88.9, election 89.2, privacy 93.8, fraud 94.3, weapons 93.3, hate 95.4, medication 95.4. Weakest languages: Igbo 75%, Yoruba 82%, Setswana 83%, isiXhosa 87%. Every number above is a model (google/gemma-3-12b-it) judging a model; no native speaker has yet rated an answer. ## What it is not good at Retrieval-augmented QA is at chance in African languages even though the model demonstrably reads the passage (grounding 0.73). Grounded generation is fully faithful to given facts in 23% of attempts, down from 31% in the previous version. The model fails to answer 41% of ordinary benign requests well, most of them by refusing; an earlier version once refused to recommend a dry cleaner, citing "illegal substances or services". That specific failure is fixed, but the broader over-refusal problem is not. Tool calling is measured with the tool marker prefilled; left to decide for itself the model almost never calls one. Multiple-choice comprehension in African languages is at chance for this model and for every model under 12B measured. The checkpoint was chosen against a rule fixed before the final experiments: every harm category within 3.5 points of the previous version, refusals on at least 91% of all attempts, degeneracy at most 1%, no capability lost. It misses that rule on drugs, violence and privacy by 0.1 point, and on degeneracy, which is 1.3%. ## Chat format Single reserved tokens mark turns: `` opens a user turn and `` an assistant turn (token ids 3 and 4). `load_example.py` in this repo shows a full prompt. Do not use `<|user|>`-style strings; they are not in the vocabulary and produce degenerate output. ## Files `model.safetensors` (bf16), `config.json`, `tokenizer.json`, `modeling_morena.py` (reference implementation, plain PyTorch, no transformers dependency), `load_example.py`, `SHA256SUMS`. A GGUF build for llama.cpp is in `vamboai/morena-1.5b-instruct-gguf`. ## Training 251.7B tokens of pretraining, 63B tokens of mid-training (315B tokens seen in total; see MORENA 1.5B base), then 4,500 steps of supervised fine-tuning with loss masking and a 500-step safety anneal. Training mixture moved in three regimes as machine-translated languages landed: African text was 24.8% of tokens seen (14.8% machine-translated) for steps 1 to 25,304, 39.1% (31.0% machine-translated) for steps 25,305 to 60,000, and 50.2% (41.6% machine-translated) during mid-training. NLLB-600M output is 24% of pretraining tokens seen and 28% once mid-training is included (57.1B tokens on disk, 21% of the 271B on disk). Release lineage: 12,834 A100 GPU-hours; the research programme that produced it, about 22,450 (about 22% of a 100,000 GPU-hour allocation), an estimated twenty-five to forty thousand dollars at $2 to $3 per A100-hour. Architecture: 28 layers x 2048, GQA 16/4, SwiGLU 6144, RoPE theta 500,000, 4,096 context, tied embeddings. Optimiser: Muon for non-embedding weights, AdamW for the rest, warmup-stable-decay schedule. ## The MORENA family | model | params | African bpb (all 12, lower is better) | role | |---|---|---|---| | MORENA 1.5B base | 1.485B | 1.408 | pretrained and mid-trained; fine-tuning starting point | | MORENA 1.5B instruct | 1.485B | 1.441 | chat, translation, tool calling; the model described in the paper | | MORENA 0.5B mini | 503M | 1.520 | pruned and distilled from the 1.5B base | | MORENA 0.5B mini instruct | 503M | 1.540 | chat fine-tune of the mini | | MORENA 0.2B nano | 209M | 1.583 | cheap trunk for ASR rescoring, keyboards, normalisation | 26 models in total were measured on African bits per byte. Every one of the 21 outside models, from 125M to 12B parameters, including the 8B African specialist Lugha-Llama-8B (1.423) and gemma-3-12b-it (2.159), sits behind all five MORENA sizes. Twelve languages: Shona, Swahili, Hausa, Yoruba, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele and Nigerian Pidgin, plus English, French and code. isiNdebele (ISO code nbl) has no FLORES+ or MasakhaNEWS coverage and is evaluated on NCHLT transcripts only. Tokenizer: 65,536-entry byte-fallback BPE trained on the target mix. African text costs 0.249 tokens per byte against 0.234 for English, about 6% more per byte than English in MORENA's vocabulary, but that same African text needs 1.39x fewer tokens than under Gemma 3's vocabulary and 1.53x fewer than under Llama 3.2's. ## Author and citation Isheanesu Misi, Vambo AI. Trained on CINECA Leonardo, with support from the AI Hub for Sustainable Development. ``` @techreport{misi2026morena, title = {MORENA: An African Foundation Model}, author = {Misi, Isheanesu}, institution = {Vambo AI}, year = {2026}, month = {September}, note = {Released 18 September 2026} } ``` ## Licence and status Public release 18 September 2026. Until then a private research preview: **not for user-facing deployment**. The weights are released as a research preview, non-commercial, because about a quarter of the tokens seen in training came from NLLB-600M output (CC-BY-NC-4.0). A tested attempt to replace that share with an Apache-licensed translator, MADLAD-400-3B, lost too much translation quality to adopt, so the non-commercial restriction stays until the underlying share is replaced or the question is otherwise resolved. Not for commercial use or public redistribution before the public release date.