Update model card to final paper numbers and house style
Browse files
README.md
CHANGED
|
@@ -20,38 +20,55 @@ datasets:
|
|
| 20 |
base_model: vamboai/morena-1.5b-base
|
| 21 |
---
|
| 22 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
# MORENA 1.5B instruct
|
| 24 |
|
| 25 |
MORENA, Sesotho and Setswana for a king, a lord or a chief, is a 1.5B-parameter decoder trained
|
| 26 |
from scratch for twelve Latin-script African languages plus English, French and code. This is the
|
| 27 |
-
instruction-tuned model described in the
|
| 28 |
-
preview. Not for user-facing deployment.**
|
| 29 |
|
| 30 |
## Headline numbers
|
| 31 |
|
| 32 |
| | MORENA 1.5B instruct | reference |
|
| 33 |
|---|---|---|
|
| 34 |
-
| African bits per byte, mean of 12 (lower is better) | **1.441** | Lugha-Llama-8B 1.423, gemma-3-12b-it 2.159,
|
| 35 |
-
| Translation, FLORES+ chrF++, English into 5 African languages, 3-shot | **45.8** | NLLB-600M 45.6, NLLB-1.3B 47.2, Lugha-Llama-8B 36.8
|
| 36 |
-
| Translation, 5 African languages into English | **48.7** | NLLB-600M 55.1 |
|
| 37 |
-
|
|
| 38 |
-
|
|
| 39 |
-
|
|
|
|
|
|
|
|
|
|
|
| 40 |
| Degenerate output | 1.3% | |
|
| 41 |
-
|
|
|
|
|
| 42 |
|
| 43 |
Safety by category (share of all attempts handled well, 24 prompts per category per language):
|
| 44 |
self-harm 92.5, child safety 91.2, drugs 86.2, violence 88.9, election 89.2, privacy 93.8, fraud 94.3,
|
| 45 |
weapons 93.3, hate 95.4, medication 95.4. Weakest languages: Igbo 75%, Yoruba 82%, Setswana 83%,
|
| 46 |
-
isiXhosa 87%. Every number is a model judging a model; no native
|
|
|
|
| 47 |
|
| 48 |
## What it is not good at
|
| 49 |
|
| 50 |
Retrieval-augmented QA is at chance in African languages even though the model demonstrably reads
|
| 51 |
-
the passage (grounding 0.73). Grounded generation is fully faithful to given facts in
|
| 52 |
-
attempts. The model
|
| 53 |
-
|
| 54 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
## Chat format
|
| 57 |
|
|
@@ -67,33 +84,46 @@ A GGUF build for llama.cpp is in `vamboai/morena-1.5b-instruct-gguf`.
|
|
| 67 |
|
| 68 |
## Training
|
| 69 |
|
| 70 |
-
251.7B tokens of pretraining
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
|
| 77 |
## The MORENA family
|
| 78 |
|
| 79 |
| model | params | African bpb (all 12, lower is better) | role |
|
| 80 |
|---|---|---|---|
|
| 81 |
| MORENA 1.5B base | 1.485B | 1.408 | pretrained and mid-trained; fine-tuning starting point |
|
| 82 |
-
| MORENA 1.5B instruct | 1.485B | 1.441 | chat, translation, tool calling; the model described in the
|
| 83 |
| MORENA 0.5B mini | 503M | 1.520 | pruned and distilled from the 1.5B base |
|
| 84 |
| MORENA 0.5B mini instruct | 503M | 1.540 | chat fine-tune of the mini |
|
| 85 |
| MORENA 0.2B nano | 209M | 1.583 | cheap trunk for ASR rescoring, keyboards, normalisation |
|
| 86 |
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
|
| 93 |
## Author and citation
|
| 94 |
|
| 95 |
-
Isheanesu Misi, Vambo AI. Trained on CINECA Leonardo with support
|
| 96 |
-
|
| 97 |
|
| 98 |
```
|
| 99 |
@techreport{misi2026morena,
|
|
@@ -108,7 +138,10 @@ from the AI Hub for Sustainable Development.
|
|
| 108 |
|
| 109 |
## Licence and status
|
| 110 |
|
| 111 |
-
Public release 18 September 2026. Until then a private research preview
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
base_model: vamboai/morena-1.5b-base
|
| 21 |
---
|
| 22 |
|
| 23 |
+
<div align="center">
|
| 24 |
+
<img src="poster.png" alt="MORENA, an African foundation model" width="600">
|
| 25 |
+
</div>
|
| 26 |
+
|
| 27 |
# MORENA 1.5B instruct
|
| 28 |
|
| 29 |
MORENA, Sesotho and Setswana for a king, a lord or a chief, is a 1.5B-parameter decoder trained
|
| 30 |
from scratch for twelve Latin-script African languages plus English, French and code. This is the
|
| 31 |
+
instruction-tuned model described in the paper *MORENA: An African Foundation Model*. **Private
|
| 32 |
+
research preview. Not for user-facing deployment.**
|
| 33 |
|
| 34 |
## Headline numbers
|
| 35 |
|
| 36 |
| | MORENA 1.5B instruct | reference |
|
| 37 |
|---|---|---|
|
| 38 |
+
| African bits per byte, mean of 12 (lower is better) | **1.441** | Lugha-Llama-8B 1.423, gemma-3-12b-it 2.159, gemma-3-1b-pt 2.335 |
|
| 39 |
+
| Translation, FLORES+ chrF++, English into 5 African languages, 3-shot | **45.8** (45.1 to 46.5) | NLLB-600M 45.6, NLLB-1.3B 47.2, Lugha-Llama-8B 36.8 |
|
| 40 |
+
| Translation, 5 African languages into English | **48.7** (47.5 to 49.7) | NLLB-600M 55.1, NLLB-1.3B 58.1 |
|
| 41 |
+
| Paired comparison with NLLB-600M, into-African, same sentences | +0.2 chrF++ (-0.5 to +0.8) | a tie overall; Yoruba +3.0 (1.6 to 4.4), isiZulu -2.5 |
|
| 42 |
+
| Belebele reading comprehension, released checkpoint | 0.309 | mean over 10 African languages, 0.24 Yoruba to 0.43 Afrikaans, 0.25 chance |
|
| 43 |
+
| Retrieval QA, open-book accuracy / grounding, African mean | 0.325 (0.297 to 0.355) / 0.73 | 0.25 chance |
|
| 44 |
+
| Grounded generation, fully faithful to given facts | 23% (22 of 96) | previous version 31% (30 of 96); un-instructed base 6 of 96 |
|
| 45 |
+
| Tool calling, correct tool / valid JSON (marker prefilled) | **98.1% / 100%** | |
|
| 46 |
+
| Safety, share of all 3,341 attempts handled well, 13 languages, 11 categories | **89.0%** (87.8 to 90.0) | 90.2% of scorable attempts |
|
| 47 |
| Degenerate output | 1.3% | |
|
| 48 |
+
| Refused harmful requests, share of all attempts | 91.5% | |
|
| 49 |
+
| Benign requests answered well | 58.9% (53.3 to 64.2) | previous version 47.9% |
|
| 50 |
|
| 51 |
Safety by category (share of all attempts handled well, 24 prompts per category per language):
|
| 52 |
self-harm 92.5, child safety 91.2, drugs 86.2, violence 88.9, election 89.2, privacy 93.8, fraud 94.3,
|
| 53 |
weapons 93.3, hate 95.4, medication 95.4. Weakest languages: Igbo 75%, Yoruba 82%, Setswana 83%,
|
| 54 |
+
isiXhosa 87%. Every number above is a model (google/gemma-3-12b-it) judging a model; no native
|
| 55 |
+
speaker has yet rated an answer.
|
| 56 |
|
| 57 |
## What it is not good at
|
| 58 |
|
| 59 |
Retrieval-augmented QA is at chance in African languages even though the model demonstrably reads
|
| 60 |
+
the passage (grounding 0.73). Grounded generation is fully faithful to given facts in 23% of
|
| 61 |
+
attempts, down from 31% in the previous version. The model fails to answer 41% of ordinary benign
|
| 62 |
+
requests well, most of them by refusing; an earlier version once refused to recommend a dry
|
| 63 |
+
cleaner, citing "illegal substances or services". That specific failure is fixed, but the broader
|
| 64 |
+
over-refusal problem is not. Tool calling is measured with the tool marker prefilled; left to decide
|
| 65 |
+
for itself the model almost never calls one. Multiple-choice comprehension in African languages is
|
| 66 |
+
at chance for this model and for every model under 12B measured.
|
| 67 |
+
|
| 68 |
+
The checkpoint was chosen against a rule fixed before the final experiments: every harm category
|
| 69 |
+
within 3.5 points of the previous version, refusals on at least 91% of all attempts, degeneracy at
|
| 70 |
+
most 1%, no capability lost. It misses that rule on drugs, violence and privacy by 0.1 point, and on
|
| 71 |
+
degeneracy, which is 1.3%.
|
| 72 |
|
| 73 |
## Chat format
|
| 74 |
|
|
|
|
| 84 |
|
| 85 |
## Training
|
| 86 |
|
| 87 |
+
251.7B tokens of pretraining, 63B tokens of mid-training (315B tokens seen in total; see MORENA
|
| 88 |
+
1.5B base), then 4,500 steps of supervised fine-tuning with loss masking and a 500-step safety
|
| 89 |
+
anneal. Training mixture moved in three regimes as machine-translated languages landed: African
|
| 90 |
+
text was 24.8% of tokens seen (14.8% machine-translated) for steps 1 to 25,304, 39.1% (31.0%
|
| 91 |
+
machine-translated) for steps 25,305 to 60,000, and 50.2% (41.6% machine-translated) during
|
| 92 |
+
mid-training. NLLB-600M output is 24% of pretraining tokens seen and 28% once mid-training is
|
| 93 |
+
included (57.1B tokens on disk, 21% of the 271B on disk).
|
| 94 |
+
|
| 95 |
+
Release lineage: 12,834 A100 GPU-hours; the research programme that produced it, about 22,450
|
| 96 |
+
(about 22% of a 100,000 GPU-hour allocation), an estimated twenty-five to forty thousand dollars
|
| 97 |
+
at $2 to $3 per A100-hour. Architecture: 28 layers x 2048, GQA 16/4, SwiGLU 6144, RoPE theta
|
| 98 |
+
500,000, 4,096 context, tied embeddings. Optimiser: Muon for non-embedding weights, AdamW for
|
| 99 |
+
the rest, warmup-stable-decay schedule.
|
| 100 |
|
| 101 |
## The MORENA family
|
| 102 |
|
| 103 |
| model | params | African bpb (all 12, lower is better) | role |
|
| 104 |
|---|---|---|---|
|
| 105 |
| MORENA 1.5B base | 1.485B | 1.408 | pretrained and mid-trained; fine-tuning starting point |
|
| 106 |
+
| MORENA 1.5B instruct | 1.485B | 1.441 | chat, translation, tool calling; the model described in the paper |
|
| 107 |
| MORENA 0.5B mini | 503M | 1.520 | pruned and distilled from the 1.5B base |
|
| 108 |
| MORENA 0.5B mini instruct | 503M | 1.540 | chat fine-tune of the mini |
|
| 109 |
| MORENA 0.2B nano | 209M | 1.583 | cheap trunk for ASR rescoring, keyboards, normalisation |
|
| 110 |
|
| 111 |
+
26 models in total were measured on African bits per byte. Every one of the 21 outside models,
|
| 112 |
+
from 125M to 12B parameters, including the 8B African specialist Lugha-Llama-8B (1.423) and
|
| 113 |
+
gemma-3-12b-it (2.159), sits behind all five MORENA sizes. Twelve languages: Shona, Swahili,
|
| 114 |
+
Hausa, Yoruba, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele and
|
| 115 |
+
Nigerian Pidgin, plus English, French and code. isiNdebele (ISO code nbl) has no FLORES+ or
|
| 116 |
+
MasakhaNEWS coverage and is evaluated on NCHLT transcripts only.
|
| 117 |
+
|
| 118 |
+
Tokenizer: 65,536-entry byte-fallback BPE trained on the target mix. African text costs 0.249
|
| 119 |
+
tokens per byte against 0.234 for English, about 6% more per byte than English in MORENA's
|
| 120 |
+
vocabulary, but that same African text needs 1.39x fewer tokens than under Gemma 3's
|
| 121 |
+
vocabulary and 1.53x fewer than under Llama 3.2's.
|
| 122 |
|
| 123 |
## Author and citation
|
| 124 |
|
| 125 |
+
Isheanesu Misi, Vambo AI. Trained on CINECA Leonardo, with support from the AI Hub for
|
| 126 |
+
Sustainable Development.
|
| 127 |
|
| 128 |
```
|
| 129 |
@techreport{misi2026morena,
|
|
|
|
| 138 |
|
| 139 |
## Licence and status
|
| 140 |
|
| 141 |
+
Public release 18 September 2026. Until then a private research preview: **not for user-facing
|
| 142 |
+
deployment**. The weights are released as a research preview, non-commercial, because about a
|
| 143 |
+
quarter of the tokens seen in training came from NLLB-600M output (CC-BY-NC-4.0). A tested
|
| 144 |
+
attempt to replace that share with an Apache-licensed translator, MADLAD-400-3B, lost too much
|
| 145 |
+
translation quality to adopt, so the non-commercial restriction stays until the underlying share is
|
| 146 |
+
replaced or the question is otherwise resolved. Not for commercial use or public redistribution
|
| 147 |
+
before the public release date.
|