yehoshua01 commited on
Commit
6e5ee5d
·
verified ·
1 Parent(s): 25c8531

Link the card to the GitHub repository and its siblings

Browse files
Files changed (1) hide show
  1. README.md +63 -8
README.md CHANGED
@@ -1,14 +1,69 @@
1
  ---
2
- license: cc-by-nc-4.0
3
- language: [ln, sn]
4
- tags: [automatic-speech-recognition, waxal, zindi]
 
 
 
 
 
 
 
 
5
  ---
6
 
7
- # waxal-qlora-largev3-lin
8
 
9
- lin QLoRA adapter over whisper-large-v3 (PEFT, load onto the base model)
 
 
10
 
11
- Trained for the Google WAXAL ASR Challenge (Zindi), phase 2. Full method:
12
- `docs/METHOD.md` in the solution bundle. This card is a stub written at upload time and is
13
- replaced by the documented card in the same commit series.
14
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: apache-2.0
3
+ language:
4
+ - ln
5
+ base_model: openai/whisper-large-v3
6
+ library_name: transformers
7
+ pipeline_tag: automatic-speech-recognition
8
+ tags:
9
+ - automatic-speech-recognition
10
+ - waxal
11
+ - zindi
12
+ - african-languages
13
  ---
14
 
15
+ # Lingala ASR — whisper-large-v3 + QLoRA adapter
16
 
17
+ Component of the **4th-place solution** to the [Google WAXAL ASR Challenge](https://github.com/yehoshua0/waxal-asr-phase2) (Zindi, phase 2):
18
+ 892 unseen clips, two African languages, **no language metadata**, scored `1 - (WER + CER) / 2` on
19
+ raw text. Private leaderboard **0.771848284**.
20
 
21
+ > **Code, full method and one-command verification: [yehoshua0/waxal-asr-phase2](https://github.com/yehoshua0/waxal-asr-phase2)**
22
+ > The repository reproduces the submitted CSV byte for byte on a laptop in about a minute, and
23
+ > re-decodes every input from the audio on rented GPUs in about four hours.
24
 
25
+ ## Role in the system
26
+
27
+ **Witness (voter).** Load onto the base model with PEFT and keep the regime it was selected under: 4-bit NF4 base, adapter NOT merged, bf16 autocast. Merging into a full-precision base changes the arithmetic it was measured in.
28
+
29
+ ## What it measured
30
+
31
+ Drives the final arbitration stage (`qlora_voter`, +0.000216). As a solo Lingala half: 0.742006.
32
+
33
+ ## Usage
34
+
35
+ ```python
36
+ from peft import PeftModel
37
+ from transformers import (BitsAndBytesConfig, WhisperForConditionalGeneration,
38
+ WhisperProcessor)
39
+ import torch
40
+
41
+ # keep the regime the adapter was selected under: 4-bit NF4 base, NOT merged, bf16 autocast
42
+ bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
43
+ bnb_4bit_use_double_quant=True,
44
+ bnb_4bit_compute_dtype=torch.bfloat16)
45
+ base = WhisperForConditionalGeneration.from_pretrained(
46
+ "openai/whisper-large-v3", quantization_config=bnb, dtype=torch.bfloat16)
47
+ model = PeftModel.from_pretrained(base, "yehoshua01/waxal-qlora-largev3-lin").eval()
48
+ proc = WhisperProcessor.from_pretrained("openai/whisper-large-v3")
49
+ ```
50
+
51
+ ## The rest of the system
52
+
53
+ | | |
54
+ |---|---|
55
+ | code, method, verification | [yehoshua0/waxal-asr-phase2](https://github.com/yehoshua0/waxal-asr-phase2) |
56
+ | cached decodes and chain inputs | [`yehoshua01/waxal-phase2-chain-inputs`](https://huggingface.co/datasets/yehoshua01/waxal-phase2-chain-inputs) |
57
+ | all checkpoints | [`yehoshua01` on the Hub](https://huggingface.co/yehoshua01) |
58
+
59
+ **Sibling checkpoints** (primaries, voters and ablations of the same system): [`waxal-mms-1b-lin-pl2-spk`](https://huggingface.co/yehoshua01/waxal-mms-1b-lin-pl2-spk) · [`waxal-sunbird51-sna-pl2-spk`](https://huggingface.co/yehoshua01/waxal-sunbird51-sna-pl2-spk) · [`waxal-whisper-turbo-lin-r1`](https://huggingface.co/yehoshua01/waxal-whisper-turbo-lin-r1) · [`waxal-whisper-turbo-lin-r2`](https://huggingface.co/yehoshua01/waxal-whisper-turbo-lin-r2) · [`waxal-omni-ctc1b-lin`](https://huggingface.co/yehoshua01/waxal-omni-ctc1b-lin) · [`waxal-omni-ctc1b-sna`](https://huggingface.co/yehoshua01/waxal-omni-ctc1b-sna) · [`waxal-sunbird51-lin-ft-r2`](https://huggingface.co/yehoshua01/waxal-sunbird51-lin-ft-r2) · [`waxal-sunbird51-lin-ft-light`](https://huggingface.co/yehoshua01/waxal-sunbird51-lin-ft-light) · [`waxal-mms-1b-lin-full`](https://huggingface.co/yehoshua01/waxal-mms-1b-lin-full) · [`waxal-mms-1b-lin-fullmeta`](https://huggingface.co/yehoshua01/waxal-mms-1b-lin-fullmeta) · [`waxal-ssa-hubert-lin`](https://huggingface.co/yehoshua01/waxal-ssa-hubert-lin)
60
+
61
+
62
+ ## Licence and intended use
63
+
64
+ `apache-2.0`. Training data is [`google/WaxalNLP`](https://huggingface.co/datasets/google/WaxalNLP)
65
+ (CC-BY-SA-4.0, share-alike), so derivatives carry that too.
66
+
67
+ **These weights are not a general-purpose ASR model.** Pseudo-labels were computed on the phase-2
68
+ test audio (transductive self-training, permitted for phase-2 training by the host), so the
69
+ checkpoint is partly adapted to that specific set.