Synthetic Only Whisper Medium Synthetic Only Medium Lora Constwarmup Best7000 LoRA Adapter

Summary

This repository contains a Whisper checkpoint for Chichewa/Nyanja automatic speech recognition, fine-tuned from openai/whisper-medium.

  • Experiment type: synthetic-only
  • Base model: openai/whisper-medium
  • Training condition: synthetic_only_medium_lora_constwarmup_best7000
  • Release artifact: LoRA adapter checkpoint selected from the best training checkpoint

Intended use

This adapter is intended for research and evaluation on Chichewa/Nyanja ASR. It must be used together with the base Whisper model. It is not a production-ready speech system and should be validated carefully before downstream use.

How to use

This repository contains an adapter, not a fully merged standalone model. Load the base model first, then attach the adapter.

from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
from peft import PeftModel

base_model_id = "openai/whisper-medium"
adapter_repo_id = "ai4good-labyrinth/synthetic-only-whisper-medium-lora-r32-no-language-v2"

processor = AutoProcessor.from_pretrained(base_model_id)
base_model = AutoModelForSpeechSeq2Seq.from_pretrained(base_model_id)
model = PeftModel.from_pretrained(base_model, adapter_repo_id)

The local evaluation script in this repository can also load the adapter directly because it reads adapter_config.json and automatically fetches the base model.

Training data

  • Training source: Prepared dataset
  • Evaluation source during training: FLEURS dev
  • Train examples before duration filtering: 374612
  • Train examples after duration filtering: 374576
  • Dev examples before duration filtering: 311
  • Dev examples after duration filtering: 305
  • Duration filter used during training: min_duration_seconds=0.0, max_duration_seconds=30.0

Training procedure

  • Fine-tuning script: experiments/whisper_finetune/finetune_whisper.py
  • Base model: openai/whisper-medium
  • Task: transcribe
  • Language hint during training/evaluation: none, corresponding to --language auto in standalone evaluation
  • LoRA: yes
  • LoRA rank: 32
  • LoRA alpha: 64
  • LoRA dropout: 0.05
  • LoRA target modules: k_proj,fc2,out_proj,v_proj,fc1,q_proj
  • Extra trainable modules: embed_tokens,proj_out
  • Mixed precision: fp16
  • Gradient checkpointing: True
  • Selected checkpoint step: 7000
  • Selected checkpoint epoch: 0.15

Training-time dev selection

The best checkpoint was selected using trainer-side dev evaluation on the duration-filtered FLEURS dev split.

  • Dev WER: 0.7020
  • Dev CER: 0.2136
  • Dev loss: 1.5151

These values come from the training pipeline and may differ slightly from standalone post-hoc evaluation because the decoding path is not perfectly identical.

Evaluation protocol

Standalone evaluation is recommended for the final release. Filtered and unfiltered results should be reported separately.

  • Filtered evaluation: min_duration_seconds=0, max_duration_seconds=30
  • Unfiltered evaluation: no duration constraint
  • Decoding task: transcribe
  • Language hint: auto

Evaluation summary

Dataset Split Setting Num examples WER CER Notes
FLEURS dev filtered 305 0.6977 0.2321 Filtered to 30 seconds
FLEURS dev unfiltered TBD TBD TBD Standalone eval pending
FLEURS test filtered 745 0.7192 0.2479 Filtered to 30 seconds
FLEURS test unfiltered TBD TBD TBD Standalone eval pending
Zambezi dev filtered 613 0.7392 0.2207 Filtered to 30 seconds
Zambezi dev unfiltered TBD TBD TBD Standalone eval pending
Zambezi test filtered 427 0.7382 0.1981 Filtered to 30 seconds
Zambezi test unfiltered TBD TBD TBD Standalone eval pending

Files in this repository

  • Adapter weights and config: repository root
  • Processor/tokenizer files: repository root
  • Evaluation JSON files: eval/...

Known limitations

  • Whisper does not provide an official Nyanja/Chichewa language token.
  • This repository contains an adapter only, so users must also comply with the upstream base model license and dataset licenses.
  • Standalone evaluation and trainer-side evaluation can differ slightly even on the same split and duration filter.
  • Cross-dataset results should be interpreted carefully because transcription conventions may differ across corpora.

Citation

If you use this checkpoint, please cite:

  • the Whisper paper
  • the FLEURS dataset
  • this repository
@misc{synthetic_only_whisper_medium_lora_r32_no_language_v2_2026,
  title        = {Synthetic Only Whisper Medium Synthetic Only Medium Lora Constwarmup Best7000 LoRA Adapter},
  author       = {AI4Good Labyrinth Team},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/ai4good-labyrinth/synthetic-only-whisper-medium-lora-r32-no-language-v2}},
  note         = {Whisper LoRA adapter fine-tuning for Chichewa/Nyanja ASR}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ai4good-labyrinth/synthetic-only-whisper-medium-no-language-lora-adapter-v2

Finetuned
(946)
this model