Lingala ASR — whisper-large-v3 + QLoRA adapter

Component of the 4th-place solution to the Google WAXAL ASR Challenge (Zindi, phase 2): 892 unseen clips, two African languages, no language metadata, scored 1 - (WER + CER) / 2 on raw text. Private leaderboard 0.771848284.

Code, full method and one-command verification: yehoshua0/waxal-asr-phase2 The repository reproduces the submitted CSV byte for byte on a laptop in about a minute, and re-decodes every input from the audio on rented GPUs in about four hours.

Role in the system

Witness (voter). Load onto the base model with PEFT and keep the regime it was selected under: 4-bit NF4 base, adapter NOT merged, bf16 autocast. Merging into a full-precision base changes the arithmetic it was measured in.

What it measured

Drives the final arbitration stage (qlora_voter, +0.000216). As a solo Lingala half: 0.742006.

Usage

from peft import PeftModel
from transformers import (BitsAndBytesConfig, WhisperForConditionalGeneration,
                          WhisperProcessor)
import torch

# keep the regime the adapter was selected under: 4-bit NF4 base, NOT merged, bf16 autocast
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_use_double_quant=True,
                         bnb_4bit_compute_dtype=torch.bfloat16)
base = WhisperForConditionalGeneration.from_pretrained(
    "openai/whisper-large-v3", quantization_config=bnb, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "yehoshua01/waxal-qlora-largev3-lin").eval()
proc = WhisperProcessor.from_pretrained("openai/whisper-large-v3")

The rest of the system

code, method, verification yehoshua0/waxal-asr-phase2
cached decodes and chain inputs yehoshua01/waxal-phase2-chain-inputs
all checkpoints yehoshua01 on the Hub

Sibling checkpoints (primaries, voters and ablations of the same system): waxal-mms-1b-lin-pl2-spk · waxal-sunbird51-sna-pl2-spk · waxal-whisper-turbo-lin-r1 · waxal-whisper-turbo-lin-r2 · waxal-omni-ctc1b-lin · waxal-omni-ctc1b-sna · waxal-sunbird51-lin-ft-r2 · waxal-sunbird51-lin-ft-light · waxal-mms-1b-lin-full · waxal-mms-1b-lin-fullmeta · waxal-ssa-hubert-lin

Licence and intended use

apache-2.0. Training data is google/WaxalNLP (CC-BY-SA-4.0, share-alike), so derivatives carry that too.

These weights are not a general-purpose ASR model. Pseudo-labels were computed on the phase-2 test audio (transductive self-training, permitted for phase-2 training by the host), so the checkpoint is partly adapted to that specific set.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yehoshua01/waxal-qlora-largev3-lin

Finetuned
(1085)
this model