Automatic Speech Recognition
Transformers
PyTorch
TensorFlow
JAX
Safetensors
whisper
audio
hf-asr-leaderboard
Eval Results (legacy)
Eval Results
Instructions to use openai/whisper-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use openai/whisper-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="openai/whisper-large")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("openai/whisper-large") model = AutoModelForSpeechSeq2Seq.from_pretrained("openai/whisper-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Commit ·
b65535b
1
Parent(s): baca495
Add pad token to tokenizer
Browse filesWhisper large is missing the pad token, which is otherwise added in the tiny-medium models (e.g. https://huggingface.co/openai/whisper-medium/blob/main/special_tokens_map.json#L125). This PR adds the pad token for the large checkpoint.
- special_tokens_map.json +1 -0
special_tokens_map.json
CHANGED
|
@@ -122,6 +122,7 @@
|
|
| 122 |
"rstrip": false,
|
| 123 |
"single_word": false
|
| 124 |
},
|
|
|
|
| 125 |
"unk_token": {
|
| 126 |
"content": "",
|
| 127 |
"lstrip": false,
|
|
|
|
| 122 |
"rstrip": false,
|
| 123 |
"single_word": false
|
| 124 |
},
|
| 125 |
+
"pad_token": "<|endoftext|>",
|
| 126 |
"unk_token": {
|
| 127 |
"content": "",
|
| 128 |
"lstrip": false,
|