thisisisheanesu commited on
Commit
c45f648
·
verified ·
1 Parent(s): 0d15d4f

Update model card to final paper numbers and house style

Browse files
Files changed (1) hide show
  1. README.md +65 -32
README.md CHANGED
@@ -20,38 +20,55 @@ datasets:
20
  base_model: vamboai/morena-1.5b-base
21
  ---
22
 
 
 
 
 
23
  # MORENA 1.5B instruct
24
 
25
  MORENA, Sesotho and Setswana for a king, a lord or a chief, is a 1.5B-parameter decoder trained
26
  from scratch for twelve Latin-script African languages plus English, French and code. This is the
27
- instruction-tuned model described in the whitepaper *MORENA: An African Foundation Model*. **Private research
28
- preview. Not for user-facing deployment.**
29
 
30
  ## Headline numbers
31
 
32
  | | MORENA 1.5B instruct | reference |
33
  |---|---|---|
34
- | African bits per byte, mean of 12 (lower is better) | **1.441** | Lugha-Llama-8B 1.423, gemma-3-12b-it 2.159, Llama-3.2-1B 2.498 |
35
- | Translation, FLORES+ chrF++, English into 5 African languages, 3-shot | **45.8** | NLLB-600M 45.6, NLLB-1.3B 47.2, Lugha-Llama-8B 36.8, 1B-class general models 9 to 14 |
36
- | Translation, 5 African languages into English | **48.7** | NLLB-600M 55.1 |
37
- | Tool calling, correct tool / valid JSON (constrained, prefilled) | **98.1% / 100%** | |
38
- | Retrieval QA, open-book accuracy / grounding, African mean | 0.325 / 0.734 | 0.25 chance |
39
- | Safety, good of scorable, 3,341 prompts, 13 languages, 11 categories | **90.2%** | |
 
 
 
40
  | Degenerate output | 1.3% | |
41
- | Benign requests answered well | 58.9% | |
 
42
 
43
  Safety by category (share of all attempts handled well, 24 prompts per category per language):
44
  self-harm 92.5, child safety 91.2, drugs 86.2, violence 88.9, election 89.2, privacy 93.8, fraud 94.3,
45
  weapons 93.3, hate 95.4, medication 95.4. Weakest languages: Igbo 75%, Yoruba 82%, Setswana 83%,
46
- isiXhosa 87%. Every number is a model judging a model; no native speaker has yet rated an answer.
 
47
 
48
  ## What it is not good at
49
 
50
  Retrieval-augmented QA is at chance in African languages even though the model demonstrably reads
51
- the passage (grounding 0.73). Grounded generation is fully faithful to given facts in about 31% of
52
- attempts. The model over-refuses: 41% of ordinary benign requests still get a refusal. Tool calling
53
- is measured with the tool marker prefilled; the model does not reliably decide on its own that a tool
54
- is needed. Reading comprehension (belebele) 0.302 against 0.250 chance.
 
 
 
 
 
 
 
 
55
 
56
  ## Chat format
57
 
@@ -67,33 +84,46 @@ A GGUF build for llama.cpp is in `vamboai/morena-1.5b-instruct-gguf`.
67
 
68
  ## Training
69
 
70
- 251.7B tokens of pretraining (45% English, 23% code, 10% native African text, 22% machine-translated
71
- African text, 7.5% French), 63B tokens of mid-training, 4,500 steps of supervised fine-tuning with
72
- loss masking and a 500-step safety anneal. Release lineage 12,834 A100 GPU-hours; the research
73
- programme that produced it about 22,450. Architecture: 28 layers x 2048, GQA 16/4, SwiGLU 6144,
74
- RoPE theta 500,000, 4,096 context, tied embeddings. Optimiser: Muon for non-embedding weights,
75
- AdamW for the rest, WSD schedule.
 
 
 
 
 
 
 
76
 
77
  ## The MORENA family
78
 
79
  | model | params | African bpb (all 12, lower is better) | role |
80
  |---|---|---|---|
81
  | MORENA 1.5B base | 1.485B | 1.408 | pretrained and mid-trained; fine-tuning starting point |
82
- | MORENA 1.5B instruct | 1.485B | 1.441 | chat, translation, tool calling; the model described in the whitepaper |
83
  | MORENA 0.5B mini | 503M | 1.520 | pruned and distilled from the 1.5B base |
84
  | MORENA 0.5B mini instruct | 503M | 1.540 | chat fine-tune of the mini |
85
  | MORENA 0.2B nano | 209M | 1.583 | cheap trunk for ASR rescoring, keyboards, normalisation |
86
 
87
- Every outside model we measured (25 in total, from 125M to 12B, including the 8B African specialist
88
- Lugha-Llama-8B at 1.423 and gemma-3-12b-it at 2.159) sits behind all five on African bits per byte.
89
- Twelve languages: Shona, Swahili, Hausa, Yoruba, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana,
90
- Afrikaans, isiNdebele, Naija Pidgin; plus English, French and code. Tokenizer: 65,536 entries trained
91
- on the target mix, fertility 0.239 tokens per byte on African text against 0.246 on English.
 
 
 
 
 
 
92
 
93
  ## Author and citation
94
 
95
- Isheanesu Misi, Vambo AI. Trained on CINECA Leonardo with support
96
- from the AI Hub for Sustainable Development.
97
 
98
  ```
99
  @techreport{misi2026morena,
@@ -108,7 +138,10 @@ from the AI Hub for Sustainable Development.
108
 
109
  ## Licence and status
110
 
111
- Public release 18 September 2026. Until then a private research preview. Weights are released under a custom licence pending resolution of one
112
- data-licensing question: 22% of the pretraining corpus derives from an NLLB model under a
113
- NonCommercial licence. Until that is settled the weights are not for commercial use or public
114
- redistribution. **Not for user-facing deployment**: see the evaluation notes above.
 
 
 
 
20
  base_model: vamboai/morena-1.5b-base
21
  ---
22
 
23
+ <div align="center">
24
+ <img src="poster.png" alt="MORENA, an African foundation model" width="600">
25
+ </div>
26
+
27
  # MORENA 1.5B instruct
28
 
29
  MORENA, Sesotho and Setswana for a king, a lord or a chief, is a 1.5B-parameter decoder trained
30
  from scratch for twelve Latin-script African languages plus English, French and code. This is the
31
+ instruction-tuned model described in the paper *MORENA: An African Foundation Model*. **Private
32
+ research preview. Not for user-facing deployment.**
33
 
34
  ## Headline numbers
35
 
36
  | | MORENA 1.5B instruct | reference |
37
  |---|---|---|
38
+ | African bits per byte, mean of 12 (lower is better) | **1.441** | Lugha-Llama-8B 1.423, gemma-3-12b-it 2.159, gemma-3-1b-pt 2.335 |
39
+ | Translation, FLORES+ chrF++, English into 5 African languages, 3-shot | **45.8** (45.1 to 46.5) | NLLB-600M 45.6, NLLB-1.3B 47.2, Lugha-Llama-8B 36.8 |
40
+ | Translation, 5 African languages into English | **48.7** (47.5 to 49.7) | NLLB-600M 55.1, NLLB-1.3B 58.1 |
41
+ | Paired comparison with NLLB-600M, into-African, same sentences | +0.2 chrF++ (-0.5 to +0.8) | a tie overall; Yoruba +3.0 (1.6 to 4.4), isiZulu -2.5 |
42
+ | Belebele reading comprehension, released checkpoint | 0.309 | mean over 10 African languages, 0.24 Yoruba to 0.43 Afrikaans, 0.25 chance |
43
+ | Retrieval QA, open-book accuracy / grounding, African mean | 0.325 (0.297 to 0.355) / 0.73 | 0.25 chance |
44
+ | Grounded generation, fully faithful to given facts | 23% (22 of 96) | previous version 31% (30 of 96); un-instructed base 6 of 96 |
45
+ | Tool calling, correct tool / valid JSON (marker prefilled) | **98.1% / 100%** | |
46
+ | Safety, share of all 3,341 attempts handled well, 13 languages, 11 categories | **89.0%** (87.8 to 90.0) | 90.2% of scorable attempts |
47
  | Degenerate output | 1.3% | |
48
+ | Refused harmful requests, share of all attempts | 91.5% | |
49
+ | Benign requests answered well | 58.9% (53.3 to 64.2) | previous version 47.9% |
50
 
51
  Safety by category (share of all attempts handled well, 24 prompts per category per language):
52
  self-harm 92.5, child safety 91.2, drugs 86.2, violence 88.9, election 89.2, privacy 93.8, fraud 94.3,
53
  weapons 93.3, hate 95.4, medication 95.4. Weakest languages: Igbo 75%, Yoruba 82%, Setswana 83%,
54
+ isiXhosa 87%. Every number above is a model (google/gemma-3-12b-it) judging a model; no native
55
+ speaker has yet rated an answer.
56
 
57
  ## What it is not good at
58
 
59
  Retrieval-augmented QA is at chance in African languages even though the model demonstrably reads
60
+ the passage (grounding 0.73). Grounded generation is fully faithful to given facts in 23% of
61
+ attempts, down from 31% in the previous version. The model fails to answer 41% of ordinary benign
62
+ requests well, most of them by refusing; an earlier version once refused to recommend a dry
63
+ cleaner, citing "illegal substances or services". That specific failure is fixed, but the broader
64
+ over-refusal problem is not. Tool calling is measured with the tool marker prefilled; left to decide
65
+ for itself the model almost never calls one. Multiple-choice comprehension in African languages is
66
+ at chance for this model and for every model under 12B measured.
67
+
68
+ The checkpoint was chosen against a rule fixed before the final experiments: every harm category
69
+ within 3.5 points of the previous version, refusals on at least 91% of all attempts, degeneracy at
70
+ most 1%, no capability lost. It misses that rule on drugs, violence and privacy by 0.1 point, and on
71
+ degeneracy, which is 1.3%.
72
 
73
  ## Chat format
74
 
 
84
 
85
  ## Training
86
 
87
+ 251.7B tokens of pretraining, 63B tokens of mid-training (315B tokens seen in total; see MORENA
88
+ 1.5B base), then 4,500 steps of supervised fine-tuning with loss masking and a 500-step safety
89
+ anneal. Training mixture moved in three regimes as machine-translated languages landed: African
90
+ text was 24.8% of tokens seen (14.8% machine-translated) for steps 1 to 25,304, 39.1% (31.0%
91
+ machine-translated) for steps 25,305 to 60,000, and 50.2% (41.6% machine-translated) during
92
+ mid-training. NLLB-600M output is 24% of pretraining tokens seen and 28% once mid-training is
93
+ included (57.1B tokens on disk, 21% of the 271B on disk).
94
+
95
+ Release lineage: 12,834 A100 GPU-hours; the research programme that produced it, about 22,450
96
+ (about 22% of a 100,000 GPU-hour allocation), an estimated twenty-five to forty thousand dollars
97
+ at $2 to $3 per A100-hour. Architecture: 28 layers x 2048, GQA 16/4, SwiGLU 6144, RoPE theta
98
+ 500,000, 4,096 context, tied embeddings. Optimiser: Muon for non-embedding weights, AdamW for
99
+ the rest, warmup-stable-decay schedule.
100
 
101
  ## The MORENA family
102
 
103
  | model | params | African bpb (all 12, lower is better) | role |
104
  |---|---|---|---|
105
  | MORENA 1.5B base | 1.485B | 1.408 | pretrained and mid-trained; fine-tuning starting point |
106
+ | MORENA 1.5B instruct | 1.485B | 1.441 | chat, translation, tool calling; the model described in the paper |
107
  | MORENA 0.5B mini | 503M | 1.520 | pruned and distilled from the 1.5B base |
108
  | MORENA 0.5B mini instruct | 503M | 1.540 | chat fine-tune of the mini |
109
  | MORENA 0.2B nano | 209M | 1.583 | cheap trunk for ASR rescoring, keyboards, normalisation |
110
 
111
+ 26 models in total were measured on African bits per byte. Every one of the 21 outside models,
112
+ from 125M to 12B parameters, including the 8B African specialist Lugha-Llama-8B (1.423) and
113
+ gemma-3-12b-it (2.159), sits behind all five MORENA sizes. Twelve languages: Shona, Swahili,
114
+ Hausa, Yoruba, Igbo, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, isiNdebele and
115
+ Nigerian Pidgin, plus English, French and code. isiNdebele (ISO code nbl) has no FLORES+ or
116
+ MasakhaNEWS coverage and is evaluated on NCHLT transcripts only.
117
+
118
+ Tokenizer: 65,536-entry byte-fallback BPE trained on the target mix. African text costs 0.249
119
+ tokens per byte against 0.234 for English, about 6% more per byte than English in MORENA's
120
+ vocabulary, but that same African text needs 1.39x fewer tokens than under Gemma 3's
121
+ vocabulary and 1.53x fewer than under Llama 3.2's.
122
 
123
  ## Author and citation
124
 
125
+ Isheanesu Misi, Vambo AI. Trained on CINECA Leonardo, with support from the AI Hub for
126
+ Sustainable Development.
127
 
128
  ```
129
  @techreport{misi2026morena,
 
138
 
139
  ## Licence and status
140
 
141
+ Public release 18 September 2026. Until then a private research preview: **not for user-facing
142
+ deployment**. The weights are released as a research preview, non-commercial, because about a
143
+ quarter of the tokens seen in training came from NLLB-600M output (CC-BY-NC-4.0). A tested
144
+ attempt to replace that share with an Apache-licensed translator, MADLAD-400-3B, lost too much
145
+ translation quality to adopt, so the non-commercial restriction stays until the underlying share is
146
+ replaced or the question is otherwise resolved. Not for commercial use or public redistribution
147
+ before the public release date.