taphuynh's picture
CT2 float16 conversion of taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16
a952ae9 verified
|
Raw History Blame Contribute Delete
3.62 kB
metadata
base_model: openai/whisper-large-v3-turbo
datasets:
  - taphuynh/arcaai-medical-malayalam-english
  - thennal/indic_tts_ml
  - thennal/ulca_ml
  - thennal/GMaSC
  - vrclc/imasc_slr
  - smcproject/MSC
  - taphuynh/MayoClinic_00001
language:
  - ml
  - en
library_name: transformers
license: apache-2.0
metrics:
  - wer
  - cer
pipeline_tag: automatic-speech-recognition
tags:
  - whisper-turbo
  - malayalam-english
  - code-switch
  - pruned
  - vocab-prune
  - full-ft
  - vocabulary-pruned
  - full-fine-tune
  - whisper
  - automatic-speech-recognition
  - generated_from_trainer
  - arca-tuner-lite
model-index:
  - name: whisper-turbo-ml-en-codeswitch-fullft-2607.29.1
    results:
      - task:
          type: automatic-speech-recognition
          name: Automatic Speech Recognition
        dataset:
          name: taphuynh/arcaai-medical-malayalam-english
          type: taphuynh/arcaai-medical-malayalam-english
        metrics:
          - type: wer
            value: 15.4686
            name: WER
          - type: cer
            value: 13.5835
            name: CER

CTranslate2 conversion. This repo is taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16 converted to CTranslate2 (float16) for use with faster-whisper.

from faster_whisper import WhisperModel
model = WhisperModel("taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1-fp16-ct2", compute_type="float16")

Converted with ct2-transformers-converter. Original model card below.

taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1

Fine-tuned from openai/whisper-large-v3-turbo with arca-tuner-lite (prune_finetune_ml_en_codeswitch).

  • Base model: openai/whisper-large-v3-turbo
  • Recipe: full fine-tune
  • Language(s): ml, en
  • Run tags: whisper-turbo, malayalam-english, code-switch, pruned, vocab-prune, full-ft
  • Run group: ml-en-cs-fullft

Evaluation

Metrics on the held-out eval split, on the best checkpoint (the one this repo contains — training used early stopping / load_best_model_at_end):

Metric Value
WER 15.4686
CER 13.5835
loss 0.0292
wer_ml 10.6873
cer_ml 9.3094
n_ml 604.0000
wer_en 12.7083
cer_en 9.6942
n_en 157.0000
wer_mixed 19.8012
cer_mixed 17.2802
n_mixed 439.0000
script_drop_rate 2.1667
hyp_ml_word_share 90.8188
cs_score 12.6363
epoch 0.3637

Usage

from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="taphuynh/whisper-turbo-ml-en-codeswitch-fullft-2607.29.1")
print(asr("audio.wav")["text"])

Training data

  • taphuynh/arcaai-medical-malayalam-english
  • thennal/indic_tts_ml
  • thennal/ulca_ml
  • thennal/GMaSC
  • vrclc/imasc_slr
  • smcproject/MSC
  • taphuynh/MayoClinic_00001

Training procedure

Hyperparameter Value
learning rate 1e-05
effective batch size 8 (8 × 1 grad-accum)
max steps 34000
warmup steps 1000
lr scheduler cosine
precision bf16
early stopping patience 6
metric for best model cs_score
seed 42

The exact resolved configuration and environment are in run_card.json in this repo.

Notes & limitations

  • Fine-tuned on domain-specific speech; expect the usual Whisper failure modes (hallucination on silence/noise, degradation far out of domain).
  • Full weights are included; load directly with transformers.
  • Not a medical device and not for clinical decision-making.