NLLB-200 Mooré LoRA (3.3B)

LoRA adapter for facebook/nllb-200-3.3B fine-tuned on ~205k cleaned Mooré↔French/English sentence pairs, all four directions (eng_Latn↔mos_Latn, fra_Latn↔mos_Latn).

Part of Mooré-Voice — open translation and speech recognition for Mooré (Mòoré / Mossi, ISO 639-3 mos), spoken by ~8 million people in and around Burkina Faso.

Evaluation

Model Direction BLEU chrF++
zero-shot base eng_Latn→mos_Latn 3.72 23.77
zero-shot base fra_Latn→mos_Latn 3.18 22.82
zero-shot base mos_Latn→eng_Latn 10.77 32.01
zero-shot base mos_Latn→fra_Latn 9.14 29.86
fine-tuned eng_Latn→mos_Latn 3.88 23.41
fine-tuned fra_Latn→mos_Latn 3.5 23.32
fine-tuned mos_Latn→eng_Latn 11.86 33.06
fine-tuned mos_Latn→fra_Latn 11.24 32.66

Training data

Curated corpus v0.1 (see repo data/CORPORA.md): MT560 (Bible-register, ~89%), community instruction pairs, NLLB-mined bitext (LASER ≥ 1.15), translatewiki. Detokenised, LID-gated, FLORES-decontaminated. FLORES-200 devtest held out for eval.

Limitations

  • Register skew: mostly religious text → weaker on administrative/technical register.
  • Mooré orthography follows the 1976/2003 standard as used by the source corpora; diacritic usage varies upstream.
  • Not human-evaluated yet; BLEU/chrF++ on FLORES only.

License note

Released CC-BY-NC-4.0 because a large share of the training text derives from sources whose redistribution terms are research-use-only or undeclared (see the repo's data/CORPORA.md / data/AUDIO_CORPORA.md). A fully permissive release is planned once the corpus is rebuilt on cleared sources (Common Voice mos + translatewiki + NLLB-mined).

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rekin226/nllb-3.3B-moore-lora-v0

Adapter
(21)
this model