Automatic Speech Recognition
Transformers
Safetensors
Tigrinya
wav2vec2
african-languages
waxal
waxalnet
Instructions to use waxal-benchmarking/mms-300m-waxal-tir with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use waxal-benchmarking/mms-300m-waxal-tir with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="waxal-benchmarking/mms-300m-waxal-tir")# Load model directly from transformers import AutoProcessor, AutoModelForCTC processor = AutoProcessor.from_pretrained("waxal-benchmarking/mms-300m-waxal-tir") model = AutoModelForCTC.from_pretrained("waxal-benchmarking/mms-300m-waxal-tir", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| library_name: transformers | |
| license: cc-by-nc-4.0 | |
| base_model: facebook/mms-300m | |
| language: | |
| - tir | |
| tags: | |
| - automatic-speech-recognition | |
| - african-languages | |
| - waxal | |
| - waxalnet | |
| - tir | |
| datasets: | |
| - waxal-benchmarking/waxal | |
| metrics: | |
| - wer | |
| - cer | |
| # MMS-300M fine-tuned on WAXAL — Tigrinya | |
| This model is part of **[WAXALNet](https://huggingface.co/waxal-benchmarking)**, a suite of ASR models fine-tuned on the [WAXAL corpus](https://huggingface.co/waxal-benchmarking) across 19 African languages, developed as part of the WAXAL ASR Benchmark study. | |
| ## Model Details | |
| | | | | |
| |---|---| | |
| | **Language** | Tigrinya (`tir`) | | |
| | **Language Family** | Afro-Asiatic | | |
| | **Architecture** | MMS-300M (300M parameters) | | |
| | **Base Model** | [facebook/mms-300m](https://huggingface.co/facebook/mms-300m) | | |
| | **Training Data** | WAXAL corpus (conversational spontaneous speech) | | |
| | **Test WER** | 57.1% | | |
| | **Test CER** | 37.2% | | |
| | **License** | cc-by-nc-4.0 | | |
| ## Intended Use | |
| This model is intended for automatic speech recognition of **Tigrinya** conversational speech. It was evaluated on the WAXAL test set (spontaneous, image-prompted speech) and partially on FLEURS (read speech). It is suitable for research and low-resource ASR applications. It is not recommended for high-stakes production use without further validation. | |
| ## Training Data | |
| Fine-tuned on the [WAXAL corpus](https://huggingface.co/waxal-benchmarking), a large-scale dataset of transcribed, image-prompted spontaneous speech across 19 African languages recorded in participants' natural environments. The Tigrinya training split contains conversational speech across diverse speakers. Data is released under CC-BY 4.0. | |
| ## Usage | |
| ```python | |
| from transformers import pipeline | |
| asr = pipeline("automatic-speech-recognition", | |
| model="waxal-benchmarking/mms-300m-waxal-tir") | |
| result = asr("audio.wav") | |
| print(result["text"]) | |
| ``` | |
| ## Test Set Performance (WAXAL Benchmark) | |
| Evaluated on the filtered WAXAL test set (duration >= 1.5s, speech rate >= 4 WPS). | |
| | Metric | Score | | |
| |---|---| | |
| | **WER** | 57.1% | | |
| | **CER** | 37.2% | | |
| Full benchmark results across all 19 languages and 6 models are reported in the [WAXAL ASR Benchmark paper](https://arxiv.org/abs/2606.02375) (citation below). | |
| ## Training procedure | |
| ### Training hyperparameters | |
| The following hyperparameters were used during training: | |
| - learning_rate: 0.0001 | |
| - train_batch_size: 4 | |
| - eval_batch_size: 8 | |
| - seed: 42 | |
| - gradient_accumulation_steps: 8 | |
| - total_train_batch_size: 32 | |
| - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments | |
| - lr_scheduler_type: linear | |
| - lr_scheduler_warmup_steps: 500 | |
| - num_epochs: 30 | |
| - mixed_precision_training: Native AMP | |
| ### Training results | |
| | Training Loss | Epoch | Step | Validation Loss | Wer | Cer | | |
| |:-------------:|:------:|:----:|:---------------:|:------:|:------:| | |
| | 32.5284 | 0.3932 | 500 | 4.3067 | 1.0 | 1.0 | | |
| | 31.3983 | 0.7865 | 1000 | 4.1331 | 0.9973 | 0.9949 | | |
| | 9.5147 | 1.1793 | 1500 | 1.6329 | 0.6428 | 0.2630 | | |
| | 6.6968 | 1.5726 | 2000 | 1.4294 | 0.5261 | 0.2077 | | |
| | 6.7439 | 1.9658 | 2500 | 1.3102 | 0.4827 | 0.1897 | | |
| | 5.5428 | 2.3586 | 3000 | 1.3659 | 0.4639 | 0.1819 | | |
| | 5.3838 | 2.7519 | 3500 | 1.3011 | 0.4520 | 0.1759 | | |
| | 5.6637 | 3.1447 | 4000 | 1.2025 | 0.4413 | 0.1731 | | |
| | 5.2161 | 3.5379 | 4500 | 1.3206 | 0.4345 | 0.1691 | | |
| | 5.3430 | 3.9312 | 5000 | 1.2392 | 0.4305 | 0.1670 | | |
| | 5.3286 | 4.3240 | 5500 | 1.1944 | 0.4309 | 0.1689 | | |
| | 4.6537 | 4.7173 | 6000 | 1.3149 | 0.4183 | 0.1633 | | |
| | 4.4528 | 5.1101 | 6500 | 1.3301 | 0.4082 | 0.1598 | | |
| | 5.0576 | 5.5033 | 7000 | 1.4483 | 0.4084 | 0.1599 | | |
| ### Framework versions | |
| - Transformers 5.0.0 | |
| - Pytorch 2.10.0+cu128 | |
| - Datasets 4.0.0 | |
| - Tokenizers 0.22.2 | |
| ## Citation | |
| ```bibtex | |
| @article{waxalnet2026, | |
| title = {The WAXAL ASR Benchmark: Fine-Tuned Edge Models Across 19 African Languages}, | |
| author = {Olufemi, Victor Tolulope and Babatunde, Oreoluwa and Njema, Ramsey and | |
| Gbotemi, Bolarinwa and Yen, Wanchi Lucia and Uzodinma, John and | |
| Ajayi, Sunday and Williams, Oluwademilade and Moshood, Kausar and | |
| Anyaele, Innocent Elendu and Arefaine, Akebert Tesfahunegn and | |
| Hunzwi, Candace and Daniel, Wongel Dawit and Namuganga, Emmilly Immaculate and | |
| Kadima, Cleophas and Bahizire, Athanase Biluge and Ranaivoson, Onitsiky and | |
| Aaron, Emmanuel and Ladislaus, Nicholaus Dismas and Muhammed, Idris and | |
| Simenya, Jonathan Enoch and Koome, Martin and Endaylalu, Matewos Tegete and | |
| Adeyemo, Peter Ifeoluwa and Birindwa, Hondi Prisca and Eze-Mbey, Ukachi Agnes and | |
| Oduro-Yeboah, Yacoba and Aremu, Toluwani and Adjovi, Pericles and | |
| Ngueajio, Mikel K and Mitra, Prasenjit}, | |
| year = {2026}, | |
| note = {arXiv preprint arXiv:2606.02375} | |
| } | |
| ``` | |
| ## Authors | |
| Victor Tolulope Olufemi · Oreoluwa Babatunde · Ramsey Njema · Bolarinwa Gbotemi · Wanchi Lucia Yen · John Uzodinma · Sunday Ajayi · Oluwademilade Williams · Kausar Moshood · Innocent Elendu Anyaele · Akebert Tesfahunegn Arefaine · Candace Hunzwi · Wongel Dawit Daniel · Emmilly Immaculate Namuganga · Cleophas Kadima · Athanase Biluge Bahizire · Onitsiky Ranaivoson · Emmanuel Aaron · Nicholaus Dismas Ladislaus · Idris Muhammed · Jonathan Enoch Simenya · Martin Koome · Matewos Tegete Endaylalu · Peter Ifeoluwa Adeyemo · Hondi Prisca Birindwa · Ukachi Agnes Eze-Mbey · Yacoba Oduro-Yeboah · Toluwani Aremu · Pericles Adjovi · Mikel K Ngueajio · Prasenjit Mitra | |
| ## Acknowledgements | |
| We thank the following contributors for their language expertise and native-speaker evaluation support: | |
| Ajara Oyinloye, Abubakari Sadic Mohammed, Hafiz Adjei, Aliga Norah Lele, Marie-Louise B. Ndamuso, and Odong Diana. | |
| This work was supported by **[Lynguallabs](https://lynguallabs.org/)** (compute, researchers & storage), | |
| **[Open Token](https://opentoken.global/)** (compute resources), and | |
| **[CMU Africa](https://www.africa.engineering.cmu.edu/)** (researchers & native speakers). | |