MMS-1B-all — Lingala adapter (MALOBA)
Language adapter only. This repository holds the fine-tuned lin_maloba adapter for facebook/mms-1b-all: about 2.3 M parameters, 0.23 % of the model. The base model is not redistributed here: the adapter loads on top of it, at the pinned revision below.
Non-commercial — see Licence. Part of the MALOBA project, Congo Digital Services (CDS SARL) with UNDP Republic of Congo.
Summary
| Task | Speech-to-text transcription of spoken Lingala |
| Base model | facebook/mms-1b-all @ 3d33597e, frozen |
| Method | Language adapter and output layer only; alphabet extended from 81 to 88 symbols (adds É”, É›) |
| Training data | audios-lingala-annotatees-v3.2, harmonised transcriptions: 15,148 segments, 47.79 h |
| Reference result | WER 44.12 %, CER 16.26 % on the leak-free v3.2 test set (reading C) |
| Training cost | 1 h 50 on one NVIDIA A40 |
| Licence | CC-BY-NC 4.0, inherited from the base model: no commercial use |
What this adapter changes
The stock lin adapter of mms-1b-all has an 81-symbol alphabet that contains neither É” (U+0254) nor É› (U+025B). Those are letters of the Lingala alphabet, not diacritics. A model that cannot form them loses a whole word every time one is expected, which is why its WER was far worse than its CER suggested.
This adapter extends the alphabet to 88 symbols, keeping the original 81 in their original order and appending ɔ ɛ ê ô û ù ɲ. The output-layer rows of the first 81 are copied unchanged; ɔ is initialised from o and ɛ from e; the other five from the mean of the existing rows.
Usage
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from transformers import Wav2Vec2ForCTC, Wav2Vec2Processor
REPO = "Congo-digital-service/mms-1b-all-lingala-maloba-adapter"
BASE, REV = "facebook/mms-1b-all", "3d33597edbdaaba14a8e858e2c8caa76e3cec0cd"
processor = Wav2Vec2Processor.from_pretrained(REPO)
model = Wav2Vec2ForCTC.from_pretrained(
BASE, revision=REV, target_lang="lin",
ignore_mismatched_sizes=True)
state = load_file(hf_hub_download(REPO, "adapter.lin_maloba.safetensors"))
model.lm_head = torch.nn.Linear(model.lm_head.in_features,
state["lm_head.weight"].shape[0])
model.config.vocab_size = state["lm_head.weight"].shape[0]
model.load_state_dict(state, strict=False) # 0 unexpected keys
model.eval()
Audio must be 16 kHz mono. Decoding is plain CTC argmax followed by processor.batch_decode, with no language model.
Note on load_adapter. model.load_adapter("lin_maloba") does not work here: it looks for the adapter inside the base model's repository, where lin_maloba does not exist. The output layer also has to be resized from 81 to 88 rows before loading, which load_adapter does not do. Hence the explicit code above, verified by reloading from the Hub and reproducing this run's transcriptions on 5 validation segments: 5/5 identical.
Training data
Congo-digital-service/audios-lingala-annotatees-v3.2 @ fed93c0e49fe (gated access, NOODL-1.0): programmes of Radio Rurale of the Republic of Congo, segmented and transcribed by MALOBA annotators, with harmonised transcriptions.
| Split | Segments kept | Hours |
|---|---|---|
| train | 15,148 / 15,175 | 47.787 |
| validation | 1,781 / 1,795 | 5.327 |
Target text is the transcription_harmonisee column, passed through texte_ctc: NFC, lowercase, non-lexical tags ([...], (...)) removed, punctuation and typographic quotes removed, whitespace collapsed. The straight apostrophe is kept: it marks elision in Lingala, and removing it would glue two words together. No digit-to-word conversion: there is no agreed reading of numbers in Lingala in this project, and inventing one would teach the model a pronunciation nobody validated.
Lines containing a character outside the extended alphabet were dropped and counted: 27 in train, 14 in validation. ∫ (U+222B) reached the inclusion threshold with 33 occurrences and was excluded deliberately: it is a mathematical sign, most likely a typo for ʃ, and adding it would teach the model to write an integral sign.
Training
Only the language adapter and the output layer are trained; the base model is frozen. init_adapter_layers() is deliberately NOT called: the published recipe reinitialises the adapter at random, and we keep what lin already knew.
- trainable parameters: 2,263,896 of 964,761,304 (adapters 2,151,168; output layer 112,728);
- learning rate 1e-3, linear schedule, 100 warm-up steps, AdamW, weight decay 0, all dropouts 0;
- effective batch 32 (8 × 4 accumulation), bf16, gradient checkpointing, length-grouped batches, seed 12345;
- 1,422 steps (3 epochs), best checkpoint at step 1,000, selected on validation WER;
- validation WER 57.55 % → 41.75 %, CER 21.70 % → 14.47 % (normalised);
- 1 h 50 on one NVIDIA A40, 11.19 GiB peak GPU memory.
A copy-check runs before training: on 20 validation segments the extended model must produce exactly what the original lin adapter produced, once É”/É› are folded back to o/e. It passed with 0 discrepancies. This is the only check that catches a shifted index in the output layer: the tensors keep the right shape and the loss falls normally either way.
Results
MALOBA v3.2 test set (reference)
Test split of the same dataset, 1,797 segments, 19 source recordings, 4.83 h, all marked eval_propre. Reading C: reference and hypothesis both in the harmonised convention. Normalised WER/CER in %.
| Model | WER | CER |
|---|---|---|
| this adapter | 44.12 | 16.26 |
| Whisper Small + MALOBA QLoRA | 44.17 | 18.95 |
mms-1b-all, stock lin adapter |
58.21 | 22.87 |
openai/whisper-small, base |
197.46 | 104.32 |
95 % intervals come from a bootstrap by source recording (1,000 resamples, seed 12345), never by segment: the 1,797 segments come from 19 recordings and are not independent.
- Fine-tuning gain: −14.02 pt, CI [−15.30, −12.83], no sign flip: established.
- Against the fine-tuned Whisper: −0.09 pt, CI [−2.12, +1.87], sign flipping on 48.3 % of resamples: the two models are indistinguishable at the word level. At the character level, this adapter makes fewer errors.
FLEURS Lingala (external, read speech)
google/fleurs, ln_cd, test split, 478 sentences, normalised:
| Model | WER | CER |
|---|---|---|
mms-1b-all, stock lin adapter |
14.66 | 4.19 |
| this adapter | 22.48 | 6.84 |
| Whisper Small + MALOBA QLoRA | 39.39 | 15.36 |
The ranking reverses on FLEURS, as expected. MMS was pre-trained largely on read speech; this adapter specialises it on radio speech and on the MALOBA spelling convention. It gains 14 WER points on the target domain and loses 7.8 on read speech: domain specialisation, not a regression. FLEURS never writes É” or É›; neutralising the spelling convention on both sides gives back 1.75 points to this adapter and only 0.14 to the stock adapter, which cannot produce these letters: direct evidence that the alphabet extension was learned.
Human evaluation
A campaign on the v3.2 test set is under way: the same 25 segments are judged for this adapter and for the fine-tuned Whisper, by three Lingala-speaking validators. Results will be added to this card.
Limitations
- 19 test recordings. Intervals are about ±3 points wide; a difference smaller than 3–4 points cannot be established on this test set.
éis removed from French loanwords by the corpus harmonisation, so the adapter does not write it in that position.- Radio speech. The adapter is specialised for broadcast and field audio; on read studio speech it does worse than the stock adapter (see FLEURS).
- One training run, one seed. Stability across seeds is not measured.
- Non-commercial licence, inherited from the base model.
Licence
CC-BY-NC 4.0, the licence of facebook/mms-1b-all. This adapter is a derivative work and carries the same terms: no commercial use, whatever the licence tier of the training data.
Project
MALOBA — Lingala digitisation, Congo Digital Services (CDS SARL) with UNDP Republic of Congo and the Ministry of Posts, Telecommunications and the Digital Economy. Contact: contact@congo-digital.com. Card updated on 30 September 2026.
Model tree for Congo-digital-service/mms-1b-all-lingala-maloba-adapter
Base model
facebook/mms-1b-allDatasets used to train Congo-digital-service/mms-1b-all-lingala-maloba-adapter
Congo-digital-service/audios-lingala-annotatees-v3.2
Evaluation results
- WER (normalised, reading C, %) on MALOBA Lingala speech v3.2, leak-free test split (1,797 segments), harmonised reading Ctest set self-reported44.120
- CER (normalised, reading C, %) on MALOBA Lingala speech v3.2, leak-free test split (1,797 segments), harmonised reading Ctest set self-reported16.260
- WER (normalised, %) on FLEURS Lingala (ln_cd), test split (478 sentences)test set self-reported22.480
- CER (normalised, %) on FLEURS Lingala (ln_cd), test split (478 sentences)test set self-reported6.840