--- language: bo license: apache-2.0 library_name: peft tags: - automatic-speech-recognition - tibetan - low-resource - whisper - lora - peft - 8bit base_model: openai/whisper-tiny datasets: - billingsmoore/nict-tib1 metrics: - cer - wer model-index: - name: tibetan-asr-nict-tib1-whisper-tiny-lora-8bit results: - task: type: automatic-speech-recognition name: Automatic Speech Recognition dataset: type: billingsmoore/nict-tib1 name: NICT-Tib1 (Lhasa Tibetan) split: test metrics: - type: cer value: 0.7555 name: CER - type: wer value: 1.0461 name: BoTok-SWER (Segmented WER) --- # Whisper Tiny — Tibetan ASR (QLoRA, 8-bit) Fine-tuned **`openai/whisper-tiny`** for automatic speech recognition (ASR) on **Lhasa Tibetan**, released alongside: > J. Moore, S. Li and P. Lauren, "Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics," in *IEEE Access*, vol. 14, pp. 101790-101805, 2026, doi: 10.1109/ACCESS.2026.3709206. This repo is one of 14 model checkpoints released with the paper, benchmarking 5 ASR architectures (Whisper Tiny/Base/Small, Wav2Vec 2.0 Base, HuBERT Base) under full fine-tuning, LoRA, and QLoRA (8-bit/4-bit) adaptation. See [billingsmoore/tibetan-asr-nict-tib1-*](https://huggingface.co/billingsmoore) for the full set. ## Model description Whisper is an encoder–decoder transformer, originally pre-trained by OpenAI on 680,000 hours of weakly-labeled multilingual speech-text data, fine-tuned here for sequence-to-sequence Tibetan transcription. This checkpoint uses **QLoRA fine-tuning (LoRA adapters trained on top of an 8-bit quantized base model)**. **LoRA configuration:** r=8, alpha=16, dropout=0.1, targeting the attention projections (`k_proj`, `q_proj`, `v_proj`, `out_proj`) and both feed-forward layers (`fc1`, `fc2`) of every Whisper transformer block. Base model weights are loaded in 8-bit precision (symmetric rounding, double quantization) via `bitsandbytes` before LoRA adapters are applied. ## Training data Fine-tuned on [NICT-Tib1](https://huggingface.co/datasets/billingsmoore/nict-tib1), a corpus of transcribed Lhasa Tibetan speech from 20 speakers (Soky, Gong & Li, 2022). The data was split by speaker (85/15, not by utterance) so that no speaker appears in both splits: 15,099 training utterances from 17 speakers, 1,547 test utterances from the remaining 3 speakers. ## Training procedure - Base model: `openai/whisper-tiny` - Learning rate: 3.75e-5 - Max steps: 4000 (warmup 500) - Effective batch size: 16 (per-device batch size 2, gradient accumulation 8) - Mixed precision: fp16, with gradient checkpointing - Generation max length: 225 ## Evaluation results Evaluated on the 1,547-utterance NICT-Tib1 test set under five metrics: Character Error Rate (CER), Syllable Error Rate (SER), and three Segmented Word Error Rate (SWER) variants using different automatic word segmenters (BoTok-SWER, BERT-SWER, and Gem-SWER, the last computed on a 500-utterance subset for API cost reasons). All values are micro-averaged. See the paper for full bootstrap confidence intervals. | Metric | Value | |---|---| | CER (micro) | 0.7555 | | SER (micro) | 0.9908 | | BoTok-SWER (micro) | 1.0461 | | BERT-SWER (micro, full 1,547-utt. test set) | 0.558 | | Gem-SWER (micro, 500-utt. subset) | 1.2158 | ### Full model family comparison (standard fine-tuning) | Model | Repo | CER | SER | BoTok-SWER | BERT-SWER | Gem-SWER | |---|---|---|---|---|---|---| | HuBERT Base | [tibetan-asr-nict-tib1-hubert-base](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-hubert-base) | 0.1352 | 0.3690 | 0.4477 | 0.161 | 0.9653 | | Wav2Vec 2.0 Base | [tibetan-asr-nict-tib1-wav2vec2-base](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-wav2vec2-base) | 0.0745 | 0.2152 | 0.2747 | 0.097 | 0.7447 | | Whisper Tiny | [tibetan-asr-nict-tib1-whisper-tiny](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-tiny) | 0.1560 | 0.2351 | 0.2975 | 0.118 | 0.6759 | | Whisper Base | [tibetan-asr-nict-tib1-whisper-base](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-base) | 0.1417 | 0.2083 | 0.2600 | 0.105 | 0.6314 | | **Whisper Small** | [tibetan-asr-nict-tib1-whisper-small](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-small) | **0.1185** | **0.1692** | **0.2042** | **0.086** | **0.5337** | > [!WARNING] > **This is a LoRA/QLoRA fine-tuned checkpoint. In the paper's benchmark, all LoRA and QLoRA configurations showed catastrophic word-level degradation** relative to full fine-tuning of the same base model, despite retaining partial character-level accuracy. This repo is published for reproducibility of that (negative) result, not as a recommended deployment artifact. If you need a usable Tibetan ASR model, use the [standard fine-tuned checkpoint](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-tiny) instead. ## How to use ```python from transformers import WhisperForConditionalGeneration, WhisperProcessor from transformers import BitsAndBytesConfig from peft import PeftModel bnb_config = BitsAndBytesConfig(load_in_8bit=True, llm_int8_skip_modules=None) base_model = WhisperForConditionalGeneration.from_pretrained( "openai/whisper-tiny", quantization_config=bnb_config, device_map="auto" ) model = PeftModel.from_pretrained(base_model, "billingsmoore/tibetan-asr-nict-tib1-whisper-tiny-lora-8bit") processor = WhisperProcessor.from_pretrained("openai/whisper-tiny", language="bo", task="transcribe") # generate as usual, e.g.: # inputs = processor(audio, sampling_rate=16000, return_tensors="pt") # predicted_ids = model.generate(**inputs) # transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True) ``` ## Limitations - Trained and evaluated only on read-speech, modern **Lhasa (Central) Tibetan** news recordings; performance on Amdo or Kham dialects, Classical Tibetan, or spontaneous/conversational speech is unknown. - The test set contains only 3 speakers (by design, to prevent speaker leakage), which limits how confidently results generalize across speaker, age, and prosodic variation. - CER alone understates errors that matter at the word level for Tibetan; see the paper for why SER and SWER are necessary complements when judging output quality. - LoRA/QLoRA adapters showed severe sequence-level collapse in this study (SWER often exceeding 1.0, i.e. worse than a random-length hypothesis) despite retaining moderate CER — treat this checkpoint as a research artifact documenting that failure mode, not as a usable transcription model. ## Citation If you use this model, please cite the paper and the source NICT-Tib1 corpus: ```bibtex @ARTICLE{11592371, author={Moore, Jacob and Li, Sheng and Lauren, Paula}, journal={IEEE Access}, title={Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics}, year={2026}, volume={14}, number={}, pages={101790-101805}, keywords={Modeling;Automatic speech recognition;Error analysis;LoRa;Measurement;Ranking (statistics);Quantization (signal);Bit error rate;Standards;Training;Tibetan;automatic speech recognition;word error rate;low-resource language}, doi={10.1109/ACCESS.2026.3709206} } @inproceedings{soky2022nict, title={Nict-tib1: A public speech corpus of lhasa dialect for benchmarking tibetan language speech recognition systems}, author={Soky, Kak and Gong, Zhuo and Li, Sheng}, booktitle={2022 25th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA)}, pages={1--5}, year={2022}, organization={IEEE} } ``` ## License Released under Apache License 2.0, inherited from the base model `openai/whisper-tiny`.