Automatic Speech Recognition
Transformers
Safetensors
Tibetan
whisper
tibetan
low-resource
Eval Results (legacy)
Instructions to use billingsmoore/tibetan-asr-nict-tib1-whisper-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use billingsmoore/tibetan-asr-nict-tib1-whisper-small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="billingsmoore/tibetan-asr-nict-tib1-whisper-small")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("billingsmoore/tibetan-asr-nict-tib1-whisper-small") model = AutoModelForSpeechSeq2Seq.from_pretrained("billingsmoore/tibetan-asr-nict-tib1-whisper-small", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: bo | |
| license: apache-2.0 | |
| library_name: transformers | |
| tags: | |
| - automatic-speech-recognition | |
| - tibetan | |
| - low-resource | |
| - whisper | |
| base_model: openai/whisper-small | |
| datasets: | |
| - billingsmoore/nict-tib1 | |
| metrics: | |
| - cer | |
| - wer | |
| model-index: | |
| - name: tibetan-asr-nict-tib1-whisper-small | |
| results: | |
| - task: | |
| type: automatic-speech-recognition | |
| name: Automatic Speech Recognition | |
| dataset: | |
| type: billingsmoore/nict-tib1 | |
| name: NICT-Tib1 (Lhasa Tibetan) | |
| split: test | |
| metrics: | |
| - type: cer | |
| value: 0.1185 | |
| name: CER | |
| - type: wer | |
| value: 0.2042 | |
| name: BoTok-SWER (Segmented WER) | |
| # Whisper Small — Tibetan ASR (full fine-tune) | |
| Fine-tuned **`openai/whisper-small`** for automatic speech recognition (ASR) on **Lhasa Tibetan**, released alongside: | |
| > J. Moore, S. Li and P. Lauren, "Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics," in *IEEE Access*, vol. 14, pp. 101790-101805, 2026, doi: 10.1109/ACCESS.2026.3709206. | |
| This repo is one of 14 model checkpoints released with the paper, benchmarking 5 ASR architectures (Whisper Tiny/Base/Small, Wav2Vec 2.0 Base, HuBERT Base) under full fine-tuning, LoRA, and QLoRA (8-bit/4-bit) adaptation. See [billingsmoore/tibetan-asr-nict-tib1-*](https://huggingface.co/billingsmoore) for the full set. | |
| ## Model description | |
| Whisper is an encoder–decoder transformer, originally pre-trained by OpenAI on 680,000 hours of weakly-labeled multilingual speech-text data, fine-tuned here for sequence-to-sequence Tibetan transcription. This checkpoint uses **full fine-tuning (all parameters updated, no quantization or adapters)**. | |
| ## Training data | |
| Fine-tuned on [NICT-Tib1](https://huggingface.co/datasets/billingsmoore/nict-tib1), a corpus of transcribed Lhasa Tibetan speech from 20 speakers (Soky, Gong & Li, 2022). The data was split by speaker (85/15, not by utterance) so that no speaker appears in both splits: 15,099 training utterances from 17 speakers, 1,547 test utterances from the remaining 3 speakers. | |
| ## Training procedure | |
| - Base model: `openai/whisper-small` | |
| - Learning rate: 1.25e-5 | |
| - Max steps: 4000 (warmup 500) | |
| - Effective batch size: 16 (per-device batch size 2, gradient accumulation 8) | |
| - Mixed precision: fp16, with gradient checkpointing | |
| - Generation max length: 225 | |
| ## Evaluation results | |
| Evaluated on the 1,547-utterance NICT-Tib1 test set under five metrics: Character Error Rate (CER), Syllable Error Rate (SER), and three Segmented Word Error Rate (SWER) variants using different automatic word segmenters (BoTok-SWER, BERT-SWER, and Gem-SWER, the last computed on a 500-utterance subset for API cost reasons). All values are micro-averaged. See the paper for full bootstrap confidence intervals. | |
| | Metric | Value | | |
| |---|---| | |
| | CER (micro) | 0.1185 | | |
| | SER (micro) | 0.1692 | | |
| | BoTok-SWER (micro) | 0.2042 | | |
| | BERT-SWER (micro, full 1,547-utt. test set) | 0.086 | | |
| | Gem-SWER (micro, 500-utt. subset) | 0.5337 | | |
| This configuration achieved the best score on every metric (CER, SER, and all three SWER variants) among all 14 models benchmarked in the paper. | |
| ### Full model family comparison (standard fine-tuning) | |
| | Model | Repo | CER | SER | BoTok-SWER | BERT-SWER | Gem-SWER | | |
| |---|---|---|---|---|---|---| | |
| | HuBERT Base | [tibetan-asr-nict-tib1-hubert-base](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-hubert-base) | 0.1352 | 0.3690 | 0.4477 | 0.161 | 0.9653 | | |
| | Wav2Vec 2.0 Base | [tibetan-asr-nict-tib1-wav2vec2-base](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-wav2vec2-base) | 0.0745 | 0.2152 | 0.2747 | 0.097 | 0.7447 | | |
| | Whisper Tiny | [tibetan-asr-nict-tib1-whisper-tiny](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-tiny) | 0.1560 | 0.2351 | 0.2975 | 0.118 | 0.6759 | | |
| | Whisper Base | [tibetan-asr-nict-tib1-whisper-base](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-base) | 0.1417 | 0.2083 | 0.2600 | 0.105 | 0.6314 | | |
| | **Whisper Small** | [tibetan-asr-nict-tib1-whisper-small](https://huggingface.co/billingsmoore/tibetan-asr-nict-tib1-whisper-small) | **0.1185** | **0.1692** | **0.2042** | **0.086** | **0.5337** | | |
| ## How to use | |
| ```python | |
| from transformers import pipeline | |
| pipe = pipeline("automatic-speech-recognition", model="billingsmoore/tibetan-asr-nict-tib1-whisper-small") | |
| result = pipe("path/to/audio.wav") | |
| print(result["text"]) | |
| ``` | |
| ## Limitations | |
| - Trained and evaluated only on read-speech, modern **Lhasa (Central) Tibetan** news recordings; performance on Amdo or Kham dialects, Classical Tibetan, or spontaneous/conversational speech is unknown. | |
| - The test set contains only 3 speakers (by design, to prevent speaker leakage), which limits how confidently results generalize across speaker, age, and prosodic variation. | |
| - CER alone understates errors that matter at the word level for Tibetan; see the paper for why SER and SWER are necessary complements when judging output quality. | |
| ## Citation | |
| If you use this model, please cite the paper and the source NICT-Tib1 corpus: | |
| ```bibtex | |
| @ARTICLE{11592371, | |
| author={Moore, Jacob and Li, Sheng and Lauren, Paula}, | |
| journal={IEEE Access}, | |
| title={Evaluating Tibetan ASR With Segmented Word Error Rate: Beyond Character-Level Metrics}, | |
| year={2026}, | |
| volume={14}, | |
| number={}, | |
| pages={101790-101805}, | |
| keywords={Modeling;Automatic speech recognition;Error analysis;LoRa;Measurement;Ranking (statistics);Quantization (signal);Bit error rate;Standards;Training;Tibetan;automatic speech recognition;word error rate;low-resource language}, | |
| doi={10.1109/ACCESS.2026.3709206} | |
| } | |
| @inproceedings{soky2022nict, | |
| title={Nict-tib1: A public speech corpus of lhasa dialect for benchmarking tibetan language speech recognition systems}, | |
| author={Soky, Kak and Gong, Zhuo and Li, Sheng}, | |
| booktitle={2022 25th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA)}, | |
| pages={1--5}, | |
| year={2022}, | |
| organization={IEEE} | |
| } | |
| ``` | |
| ## License | |
| Released under Apache License 2.0, inherited from the base model `openai/whisper-small`. | |