Instructions to use yehoshua01/waxal-sunbird51-sna-pl2-spk with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yehoshua01/waxal-sunbird51-sna-pl2-spk with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="yehoshua01/waxal-sunbird51-sna-pl2-spk")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("yehoshua01/waxal-sunbird51-sna-pl2-spk") model = AutoModelForSpeechSeq2Seq.from_pretrained("yehoshua01/waxal-sunbird51-sna-pl2-spk", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Shona ASR — Sunbird-51 + pseudo-labels + per-speaker micro-fine-tune
Component of the 4th-place solution to the Google WAXAL ASR Challenge (Zindi, phase 2):
892 unseen clips, two African languages, no language metadata, scored 1 - (WER + CER) / 2 on
raw text. Private leaderboard 0.771848284.
Code, full method and one-command verification: yehoshua0/waxal-asr-phase2 The repository reproduces the submitted CSV byte for byte on a laptop in about a minute, and re-decodes every input from the audio on rented GPUs in about four hours.
Role in the system
Primary — the Shona half of the shipped submission. Sunbird's 51-language whisper, domain-adapted with pseudo-labels on the test audio and a per-speaker micro-fine-tune at LR 1e-6.
What it measured
Replacing this half costs -0.017824 (mms-1b-all fine-tuned on Shona) to -0.029706 (omniASR CTC-7B). The per-speaker step itself is near a no-op — 10 clips of 445 — the asset is the pseudo-label domain adaptation.
Usage
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import torch, soundfile as sf, torchaudio.functional as AF
proc = WhisperProcessor.from_pretrained("yehoshua01/waxal-sunbird51-sna-pl2-spk")
model = WhisperForConditionalGeneration.from_pretrained("yehoshua01/waxal-sunbird51-sna-pl2-spk").eval().cuda()
tok = proc.tokenizer
# the language slot is a LEARNED decoder state -- use the token this checkpoint trained under
forced = [(1, tok.convert_tokens_to_ids("50410 # Sunbird card map: sna")),
(2, tok.convert_tokens_to_ids("<|transcribe|>")),
(3, tok.convert_tokens_to_ids("<|notimestamps|>"))]
wav, sr = sf.read("clip.wav", dtype="float32")
wav = AF.resample(torch.from_numpy(wav), sr, 16_000).numpy()[:30 * 16_000]
f = proc.feature_extractor(wav, sampling_rate=16_000, return_tensors="pt").input_features
out = model.generate(f.to("cuda", model.dtype), forced_decoder_ids=forced, max_new_tokens=220)
print(proc.batch_decode(out, skip_special_tokens=True)[0])
The rest of the system
| code, method, verification | yehoshua0/waxal-asr-phase2 |
| cached decodes and chain inputs | yehoshua01/waxal-phase2-chain-inputs |
| all checkpoints | yehoshua01 on the Hub |
Sibling checkpoints (primaries, voters and ablations of the same system): waxal-mms-1b-lin-pl2-spk · waxal-whisper-turbo-lin-r1 · waxal-whisper-turbo-lin-r2 · waxal-qlora-largev3-lin · waxal-omni-ctc1b-lin · waxal-omni-ctc1b-sna · waxal-sunbird51-lin-ft-r2 · waxal-sunbird51-lin-ft-light · waxal-mms-1b-lin-full · waxal-mms-1b-lin-fullmeta · waxal-ssa-hubert-lin
Licence and intended use
cc-by-sa-4.0. Training data is google/WaxalNLP
(CC-BY-SA-4.0, share-alike), so derivatives carry that too.
These weights are not a general-purpose ASR model. Pseudo-labels were computed on the phase-2 test audio (transductive self-training, permitted for phase-2 training by the host), so the checkpoint is partly adapted to that specific set.
- Downloads last month
- 10
Model tree for yehoshua01/waxal-sunbird51-sna-pl2-spk
Base model
huwenjie333/whisper-v3-ft-af51-0903