FLEURS-Trigram Whisper Medium No-Language LoRA Adapter

Summary

This repository contains a Whisper checkpoint for Chichewa/Nyanja automatic speech recognition, fine-tuned from openai/whisper-medium.

  • Experiment type: fleurs-trigram
  • Base model: openai/whisper-medium
  • Training condition: no_language
  • Release artifact: LoRA adapter checkpoint selected from the best training checkpoint

Intended use

This adapter is intended for research and evaluation on Chichewa/Nyanja ASR. It must be used together with the base Whisper model. It is not a production-ready speech system and should be validated carefully before downstream use.

How to use

This repository contains an adapter, not a fully merged standalone model. Load the base model first, then attach the adapter.

from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
from peft import PeftModel

base_model_id = "openai/whisper-medium"
adapter_repo_id = "ai4good-labyrinth/fleurs-trigram-hours0p75-whisper-medium-no-language-lora-adapter"

processor = AutoProcessor.from_pretrained(base_model_id)
base_model = AutoModelForSpeechSeq2Seq.from_pretrained(base_model_id)
model = PeftModel.from_pretrained(base_model, adapter_repo_id)

The local evaluation script in this repository can also load the adapter directly because it reads adapter_config.json and automatically fetches the base model.

Training data

  • Training source: FLEURS train + Chichewa Trigrams train
  • Evaluation source during training: FLEURS dev
  • Train examples before duration filtering: 22118
  • Train examples after duration filtering: 22077
  • Dev examples before duration filtering: 311
  • Dev examples after duration filtering: 305
  • Duration filter used during training: min_duration_seconds=0.0, max_duration_seconds=30.0

Training procedure

  • Fine-tuning script: experiments/whisper_finetune/finetune_whisper.py
  • Base model: openai/whisper-medium
  • Task: transcribe
  • Language hint during training/evaluation: none, corresponding to --language auto in standalone evaluation
  • LoRA: yes
  • LoRA rank: 32
  • LoRA alpha: 64
  • LoRA dropout: 0.05
  • LoRA target modules: q_proj,fc1,out_proj,fc2,k_proj,v_proj
  • Extra trainable modules: embed_tokens,proj_out
  • Mixed precision: fp16
  • Gradient checkpointing: True
  • Selected checkpoint step: TBD
  • Selected checkpoint epoch: TBD

Training-time dev selection

The best checkpoint was selected using trainer-side dev evaluation on the duration-filtered FLEURS dev split.

  • Dev WER: 0.4721
  • Dev CER: 0.1239
  • Dev loss: 0.6224

These values come from the training pipeline and may differ slightly from standalone post-hoc evaluation because the decoding path is not perfectly identical.

Evaluation protocol

Standalone evaluation is recommended for the final release. Filtered and unfiltered results should be reported separately.

  • Filtered evaluation: min_duration_seconds=0, max_duration_seconds=30
  • Unfiltered evaluation: no duration constraint
  • Decoding task: transcribe
  • Language hint: auto

Evaluation summary

Dataset Split Setting Num examples WER CER Notes
FLEURS dev filtered 305 0.5140 0.1261 Filtered to 30 seconds
FLEURS dev unfiltered TBD TBD TBD Standalone eval pending
FLEURS test filtered 745 0.4668 0.1226 Filtered to 30 seconds
FLEURS test unfiltered TBD TBD TBD Standalone eval pending
Zambezi dev filtered 613 0.8179 0.2829 Filtered to 30 seconds
Zambezi dev unfiltered TBD TBD TBD Standalone eval pending
Zambezi test filtered 427 0.8890 0.3105 Filtered to 30 seconds
Zambezi test unfiltered 428 0.8843 0.3083 No duration filter

Files in this repository

  • Adapter weights and config: repository root
  • Processor/tokenizer files: repository root
  • Evaluation JSON files: eval/...

Known limitations

  • Whisper does not provide an official Nyanja/Chichewa language token.
  • This repository contains an adapter only, so users must also comply with the upstream base model license and dataset licenses.
  • Standalone evaluation and trainer-side evaluation can differ slightly even on the same split and duration filter.
  • Cross-dataset results should be interpreted carefully because transcription conventions may differ across corpora.

Citation

If you use this checkpoint, please cite:

  • the Whisper paper
  • the FLEURS dataset
  • this repository
@misc{fleurs_trigram_hours0p75_whisper_medium_no_language_lora_adapter_2026,
  title        = {FLEURS-Trigram Whisper Medium No-Language LoRA Adapter},
  author       = {AI4Good Labyrinth Team},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/ai4good-labyrinth/fleurs-trigram-hours0p75-whisper-medium-no-language-lora-adapter}},
  note         = {Whisper LoRA adapter fine-tuning for Chichewa/Nyanja ASR}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai4good-labyrinth/fleurs-trigram-hours0p75-whisper-medium-no-language-lora-adapter

Finetuned
(946)
this model