--- language: - ny tags: - automatic-speech-recognition - whisper - chichewa - nyanja - fleurs license: apache-2.0 pipeline_tag: automatic-speech-recognition base_model: openai/whisper-small --- # Fleurs Synthetic Whisper Small Fleurs Synthetic Hours0P75 Checkpoint ## Summary This repository contains a Whisper checkpoint for Chichewa/Nyanja automatic speech recognition, fine-tuned from `openai/whisper-small`. - Experiment type: `fleurs-synthetic` - Base model: `openai/whisper-small` - Training condition: `fleurs_synthetic_hours0p75` - Release artifact: full fine-tuned checkpoint selected from the best training checkpoint ## Intended use This checkpoint is intended for research and evaluation on Chichewa/Nyanja ASR. It is not a production-ready speech system and should be validated carefully before downstream use. ## How to use This repository contains a full fine-tuned checkpoint. It can be loaded directly with Transformers. ```python from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor model_id = "ai4good-labyrinth/fleurs-synthetic-hours0p75-whisper-small-no-language" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForSpeechSeq2Seq.from_pretrained(model_id) ``` ## Training data - Training source: `FLEURS train + Prepared dataset` - Evaluation source during training: `FLEURS dev` - Train examples before duration filtering: `8243` - Train examples after duration filtering: `8202` - Dev examples before duration filtering: `311` - Dev examples after duration filtering: `305` - Duration filter used during training: `min_duration_seconds=0.0`, `max_duration_seconds=30.0` ## Training procedure - Fine-tuning script: `experiments/whisper_finetune/finetune_whisper.py` - Base model: `openai/whisper-small` - Task: `transcribe` - Language hint during training/evaluation: none, corresponding to `--language auto` in standalone evaluation - Mixed precision: `no` - Gradient checkpointing: `False` - Selected checkpoint step: `6600` - Selected checkpoint epoch: `6.43` ## Training-time dev selection The best checkpoint was selected using trainer-side dev evaluation on the duration-filtered FLEURS dev split. - Dev WER: `0.5012` - Dev CER: `0.1509` - Dev loss: `0.7995` These values come from the training pipeline and may differ slightly from standalone post-hoc evaluation because the decoding path is not perfectly identical. ## Evaluation protocol Standalone evaluation is recommended for the final release. Filtered and unfiltered results should be reported separately. - Filtered evaluation: `min_duration_seconds=0`, `max_duration_seconds=30` - Unfiltered evaluation: no duration constraint - Decoding task: `transcribe` - Language hint: `auto` ## Evaluation summary | Dataset | Split | Setting | Num examples | WER | CER | Notes | |---|---|---|---:|---:|---:|---| | FLEURS | dev | filtered | 305 | 0.5063 | 0.1414 | Filtered to 30 seconds | | FLEURS | dev | unfiltered | TBD | TBD | TBD | Standalone eval pending | | FLEURS | test | filtered | 745 | 0.4965 | 0.1460 | Filtered to 30 seconds | | FLEURS | test | unfiltered | TBD | TBD | TBD | Standalone eval pending | | Zambezi | dev | filtered | TBD | TBD | TBD | Standalone eval pending | | Zambezi | dev | unfiltered | TBD | TBD | TBD | Standalone eval pending | | Zambezi | test | filtered | 427 | 0.6231 | 0.1545 | Filtered to 30 seconds | | Zambezi | test | unfiltered | 428 | 0.6203 | 0.1536 | No duration filter | ## Files in this repository - Model weights and config: repository root - Processor/tokenizer files: repository root - Evaluation JSON files: `eval/...` ## Known limitations - Whisper does not provide an official Nyanja/Chichewa language token. - Users must also comply with the upstream dataset licenses and any upstream model license obligations. - Standalone evaluation and trainer-side evaluation can differ slightly even on the same split and duration filter. - Cross-dataset results should be interpreted carefully because transcription conventions may differ across corpora. ## Citation If you use this checkpoint, please cite: - the Whisper paper - the FLEURS dataset - this repository ```bibtex @misc{fleurs_synthetic_hours0p75_whisper_small_no_language_2026, title = {Fleurs Synthetic Whisper Small Fleurs Synthetic Hours0P75 Checkpoint}, author = {AI4Good Labyrinth Team}, year = {2026}, howpublished = {\url{https://huggingface.co/ai4good-labyrinth/fleurs-synthetic-hours0p75-whisper-small-no-language}}, note = {Whisper fine-tuning for Chichewa/Nyanja ASR} } ```