Instructions to use yehoshua01/waxal-qlora-largev3-lin with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yehoshua01/waxal-qlora-largev3-lin with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="yehoshua01/waxal-qlora-largev3-lin")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("yehoshua01/waxal-qlora-largev3-lin", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Lingala ASR — whisper-large-v3 + QLoRA adapter
Component of the 4th-place solution to the Google WAXAL ASR Challenge (Zindi, phase 2):
892 unseen clips, two African languages, no language metadata, scored 1 - (WER + CER) / 2 on
raw text. Private leaderboard 0.771848284.
Code, full method and one-command verification: yehoshua0/waxal-asr-phase2 The repository reproduces the submitted CSV byte for byte on a laptop in about a minute, and re-decodes every input from the audio on rented GPUs in about four hours.
Role in the system
Witness (voter). Load onto the base model with PEFT and keep the regime it was selected under: 4-bit NF4 base, adapter NOT merged, bf16 autocast. Merging into a full-precision base changes the arithmetic it was measured in.
What it measured
Drives the final arbitration stage (qlora_voter, +0.000216). As a solo Lingala half: 0.742006.
Usage
from peft import PeftModel
from transformers import (BitsAndBytesConfig, WhisperForConditionalGeneration,
WhisperProcessor)
import torch
# keep the regime the adapter was selected under: 4-bit NF4 base, NOT merged, bf16 autocast
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16)
base = WhisperForConditionalGeneration.from_pretrained(
"openai/whisper-large-v3", quantization_config=bnb, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "yehoshua01/waxal-qlora-largev3-lin").eval()
proc = WhisperProcessor.from_pretrained("openai/whisper-large-v3")
The rest of the system
| code, method, verification | yehoshua0/waxal-asr-phase2 |
| cached decodes and chain inputs | yehoshua01/waxal-phase2-chain-inputs |
| all checkpoints | yehoshua01 on the Hub |
Sibling checkpoints (primaries, voters and ablations of the same system): waxal-mms-1b-lin-pl2-spk · waxal-sunbird51-sna-pl2-spk · waxal-whisper-turbo-lin-r1 · waxal-whisper-turbo-lin-r2 · waxal-omni-ctc1b-lin · waxal-omni-ctc1b-sna · waxal-sunbird51-lin-ft-r2 · waxal-sunbird51-lin-ft-light · waxal-mms-1b-lin-full · waxal-mms-1b-lin-fullmeta · waxal-ssa-hubert-lin
Licence and intended use
apache-2.0. Training data is google/WaxalNLP
(CC-BY-SA-4.0, share-alike), so derivatives carry that too.
These weights are not a general-purpose ASR model. Pseudo-labels were computed on the phase-2 test audio (transductive self-training, permitted for phase-2 training by the host), so the checkpoint is partly adapted to that specific set.
Model tree for yehoshua01/waxal-qlora-largev3-lin
Base model
openai/whisper-large-v3