taphuynh's picture
CT2 float16 conversion of taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16
a952ae9 verified
|
Raw History Blame Contribute Delete
3.62 kB
---
base_model: openai/whisper-large-v3-turbo
datasets:
- taphuynh/arcaai-medical-malayalam-english
- thennal/indic_tts_ml
- thennal/ulca_ml
- thennal/GMaSC
- vrclc/imasc_slr
- smcproject/MSC
- taphuynh/MayoClinic_00001
language:
- ml
- en
library_name: transformers
license: apache-2.0
metrics:
- wer
- cer
pipeline_tag: automatic-speech-recognition
tags:
- whisper-turbo
- malayalam-english
- code-switch
- pruned
- vocab-prune
- full-ft
- vocabulary-pruned
- full-fine-tune
- whisper
- automatic-speech-recognition
- generated_from_trainer
- arca-tuner-lite
model-index:
- name: whisper-turbo-ml-en-codeswitch-fullft-2607.29.1
results:
- task:
type: automatic-speech-recognition
name: Automatic Speech Recognition
dataset:
name: taphuynh/arcaai-medical-malayalam-english
type: taphuynh/arcaai-medical-malayalam-english
metrics:
- type: wer
value: 15.4686
name: WER
- type: cer
value: 13.5835
name: CER
---
> **CTranslate2 conversion.** This repo is [`taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16`](https://huggingface.co/taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16) converted to CTranslate2 (`float16`) for use with faster-whisper.
>
> ```python
> from faster_whisper import WhisperModel
> model = WhisperModel("taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16-ct2", compute_type="float16")
> ```
>
> Converted with `ct2-transformers-converter`. Original model card below.
# taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1
Fine-tuned from [`openai/whisper-large-v3-turbo`](https://huggingface.co/openai/whisper-large-v3-turbo) with
`arca-tuner-lite` (`prune_finetune_ml_en_codeswitch`).
- **Base model:** `openai/whisper-large-v3-turbo`
- **Recipe:** full fine-tune
- **Language(s):** ml, en
- **Run tags:** `whisper-turbo`, `malayalam-english`, `code-switch`, `pruned`, `vocab-prune`, `full-ft`
- **Run group:** `ml-en-cs-fullft`
## Evaluation
Metrics on the held-out eval split, on the best checkpoint (the one this repo
contains — training used early stopping / `load_best_model_at_end`):
| Metric | Value |
| --- | --- |
| WER | 15.4686 |
| CER | 13.5835 |
| loss | 0.0292 |
| wer_ml | 10.6873 |
| cer_ml | 9.3094 |
| n_ml | 604.0000 |
| wer_en | 12.7083 |
| cer_en | 9.6942 |
| n_en | 157.0000 |
| wer_mixed | 19.8012 |
| cer_mixed | 17.2802 |
| n_mixed | 439.0000 |
| script_drop_rate | 2.1667 |
| hyp_ml_word_share | 90.8188 |
| cs_score | 12.6363 |
| epoch | 0.3637 |
## Usage
```python
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1")
print(asr("audio.wav")["text"])
```
## Training data
- `taphuynh/arcaai-medical-malayalam-english`
- `thennal/indic_tts_ml`
- `thennal/ulca_ml`
- `thennal/GMaSC`
- `vrclc/imasc_slr`
- `smcproject/MSC`
- `taphuynh/MayoClinic_00001`
## Training procedure
| Hyperparameter | Value |
| --- | --- |
| learning rate | 1e-05 |
| effective batch size | 8 (8 × 1 grad-accum) |
| max steps | 34000 |
| warmup steps | 1000 |
| lr scheduler | cosine |
| precision | bf16 |
| early stopping patience | 6 |
| metric for best model | cs_score |
| seed | 42 |
The exact resolved configuration and environment are in `run_card.json` in this
repo.
## Notes & limitations
- Fine-tuned on domain-specific speech; expect the usual Whisper failure modes
(hallucination on silence/noise, degradation far out of domain).
- Full weights are included; load directly with `transformers`.
- Not a medical device and not for clinical decision-making.