Kusaal ASR (MMS adapter)
Author: Prince Nasamu Alhassan
Overview
A per-language MMS adapter over facebook/mms-1b-all.
This is now the best Kusaal recogniser in the project. The previous champion, KASA-42, scores 41.0 / 19.2 and cannot be fine-tuned from — it ships raw .pt files and an inference-only int8 ONNX export with no architectures or model_type. This model beats it by 10.6 WER points and can be trained further.
Use it
This repo holds only the adapter (about 9 MB), not a whole model. Load the base and apply it — and take the head width from the adapter, not from the tokenizer: they differ by the added special tokens, and letting the tokenizer decide raises a size mismatch.
AutoConfig.from_pretrainedon this repo fails with "Unrecognized model … should have amodel_typekey", andAutoModelForCTCfails too. That is expected, not a broken upload: there is no model here, only an adapter and its tokenizer.AutoProcessor.from_pretraineddoes work. Aconfig.jsonis deliberately not shipped — the base's would advertise a vocabulary of 154 where this head is 58–62 wide, and would invite loading weights that are not in the repo.
import torch, soundfile as sf
from huggingface_hub import hf_hub_download, list_repo_files
from safetensors.torch import load_file
from transformers import AutoProcessor, Wav2Vec2ForCTC
repo = "PrinceAlhassanNasamu/tekyerema-asr-mms-kus"
adapter = [f for f in list_repo_files(repo)
if f.startswith("adapter.") and f.endswith(".safetensors")][0]
sd = load_file(hf_hub_download(repo, adapter))
proc = AutoProcessor.from_pretrained(repo)
model = Wav2Vec2ForCTC.from_pretrained(
"facebook/mms-1b-all",
vocab_size=sd["lm_head.weight"].shape[0], # the ADAPTER decides this
ignore_mismatched_sizes=True).eval()
missing, unexpected = model.load_state_dict(sd, strict=False)
assert not unexpected, unexpected # never transcribe with a half-loaded model
wav, sr = sf.read("clip.wav", dtype="float32")
inp = proc(wav, sampling_rate=16_000, return_tensors="pt")
with torch.no_grad():
logits = model(**inp).logits
print(proc.batch_decode(logits.argmax(-1))[0])
Training data
Trained on the Ghana Speech dataset and related Ghanaian corpora, licensed CC BY-NC 4.0.
Measured
On kusaal_scripture, same items and same scorer as the baseline:
| model | WER / CER |
|---|---|
| baseline it was fine-tuned from | 59.34 / 22.10 |
| this model | 30.44 / 13.52 |
A WER alone is not informative — compared against another language it means nothing. Compared against the model it started from, it means everything.
Intended use & license
Non-commercial use only (CC BY-NC 4.0). This is inherited from the training data and required by the terms under which the compute was granted: models trained in that window are non-commercial by condition of access, not by inference.
Limitations, stated plainly
- Dagbani had no recogniser of its own for this whole project, and the
reason given for that was wrong. Every card here said "one fine-tuning
session on 74 validation rows would not change that". Those 74 rows are
the eng-dag machine-translation validation split. The Dagbani
speech data in this same account is
waxal_dag: 13,228 training rows, 1,750 validation rows, ~71 hours, 1,041 speakers with the largest at 1% — more data and better speaker diversity than Ewe, which produced a working 42.19 WER recogniser. A number was carried across from a translation table into a speech claim, and then repeated on every model card on the account. It is training now, on 2026-08-31. Until it is scored, the honest statement is that Dagbani's best available recogniser scores 86.6 WER and nobody had tried fine-tuning on the data already in hand. - Evaluation is on read and machine-translated text. No recordings of people speaking agent commands in these languages exist. Numbers measured this way are optimistic about phrasing and pessimistic about code-switching, and should not be read as field performance.
- Research work from a hackathon entry, not a supported product.
The rest of the family
Recognisers
whisper-large-v3-turbo-tekyerema-eng-foundation— Ghanaian English ASR — course 1 (foundation)kusaal-whisper-small-lora— Kusaal ASR (Whisper-small LoRA, superseded)kasa42-asr— KASA-42 (Kusaal, third-party export)tekyerema-asr-ctc— Twi ASR (w2v-BERT CTC)tekyerema-asr-mms-ewe— Ewe ASR (MMS adapter)tekyerema-asr-mms-dag— Dagbani ASR (MMS adapter)tekyerema-asr-mms-hau— Hausa ASR (MMS adapter)tekyerema-asr-mms-kus— Kusaal ASR (MMS adapter)whisper-large-v3-turbo-tekyerema-eng— Ghanaian English ASR (Whisper large-v3-turbo)
Voices
tekyerema-tts-twi— Twi TTS (VITS)tekyerema-tts-kus— Kusaal TTS (VITS)tekyerema-tts-ewe— Ewe TTS (VITS)tekyerema-tts-hau— Hausa TTS (VITS)tekyerema-tts-eng— Ghanaian English TTS (VITS)
Agent models
tekyerema-1-reply— Tɛkyerɛma-1 reply adapter (arm ①)tekyerema-1-native-reply— Tɛkyerɛma-1 reply adapter (arm ②)tekyerema-1-tool— Tɛkyerɛma-1 tool adapter (arm 1)tekyerema-audio-native— Tɛkyerɛma-1 audio-native (arm 3)tekyerema-audio-native-4k— Tɛkyerɛma-1 audio-native, 4,000 clips (arm 3 v2)tekyerema-1-native-tool— Tɛkyerɛma-1 tool adapter (arm 2)
Translation
tekyerema-nllb600m-v1— Tɛkyerɛma MT v1 (NLLB-600M)kusaal-nllb-600M— Kusaal MT specialist (NLLB-600M)
Routing
tekyerema-intent-afroxlmr— Intent classifier (AfroXLMR)
Acknowledgements
Compute resources provided by AI Skills and Compute Africa (AISCA).
Trained on the Ghana NLP H200 GPU. Please keep derivatives non-commercial
and share improvements back with the Ghana NLP community
(ghananlpcommunity).
Model tree for PrinceAlhassanNasamu/tekyerema-asr-mms-kus
Base model
facebook/mms-1b-allEvaluation results
- wer on kusaal_scriptureself-reported30.440
- cer on kusaal_scriptureself-reported13.520