Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 1.75 kB xet | 1eaf991e | |
| README.md | 30.6 kB xet | dda799ed | |
| added_tokens.json | 34.6 kB xet | 647dbabd | |
| config.json | 1.25 kB xet | 8f00e8c9 | |
| generation_config.json | 3.76 kB xet | 06b00493 | |
| merges.txt | 494 kB xet | b3236b79 | |
| model.safetensors | 3.24 GB xet | 5cbacece | |
| normalizer.json | 52.7 kB xet | 1301c121 | |
| preprocessor_config.json | 340 Bytes xet | 0bf6e84e | |
| special_tokens_map.json | 2.19 kB xet | 99ebb6fd | |
| test_sound_en.wav | 1.3 MB xet | d395c6cd | |
| test_sound_nonspeech.wav | 320 kB xet | 5a6f2809 | |
| test_sound_ru.wav | 959 kB xet | 6b9a1723 | |
| test_sound_ru_longform.wav | 9.53 MB xet | 03691d09 | |
| tokenizer.json | 2.71 MB xet | 928f1b82 | |
| tokenizer_config.json | 283 kB xet | 7e2f0c8f | |
| vocab.json | 1.04 MB xet | 70a15bbc |
Whisper-Podlodka-Turbo
Whisper-Podlodka-Turbo is a new fine-tuned version of a Whisper large-v3-turbo. The main goal of the fine-tuning is to improve the quality of speech recognition and speech translation for Russian and English, as well as reduce the occurrence of hallucinations when processing non-speech audio signals.
Model Description
Whisper-Podlodka-Turbo is a new fine-tuned version of Whisper-Large-V3-Turbo, optimized for high-quality Russian speech recognition with proper punctuation + capitalization and enhanced with noise resistance capability.
Key Benefits
- π― Improved Russian speech recognition quality compared to the base Whisper-Large-V3-Turbo model
- βοΈ Correct Russian punctuation and capitalization
- π§ Enhanced background noise resistance
- π« Reduced number of hallucinations, especially in non-speech segments
Supported Tasks
- Automatic Speech Recognition (ASR):
- π·πΊ Russian (primary focus)
- π¬π§ English
- Speech Translation:
- Russian βοΈ English
- Speech Language Detection (including non-speech detection)
Uses
Installation
Whisper-Podlodka-Turbo is supported in Hugging Face π€ Transformers. To run the model, first install the Transformers library. For this example, we'll also install π€ Datasets to load toy audio dataset from the Hugging Face Hub, and π€ Accelerate to reduce the model loading time:
pip install --upgrade pip
pip install --upgrade transformers datasets[audio] accelerate
Also, I recommend using whisper-lid for initial spoken language detection. Therefore, this library is also worth installing:
pip install --upgrade whisper-lid
Usages Cases
Speech recognition
The model can be used with the pipeline class to transcribe audios of arbitrary language:
import librosa # for loading sound from local file
from transformers import pipeline # for working with Whisper-Podlodka-Turbo
import wget # for downloading demo sound from its URL
from whisper_lid.whisper_lid import detect_language_in_speech # for spoken language detection
model_id = "bond005/whisper-podlodka-turbo" # the best Whisper model :-)
target_sampling_rate = 16_000 # Hz
asr = pipeline(model=model_id, device_map='auto', torch_dtype='auto')
# An example of speech recognition in Russian, spoken by a native speaker of this language
sound_ru_url = 'https://huggingface.co/bond005/whisper-podlodka-turbo/resolve/main/test_sound_ru.wav'
sound_ru_name = wget.download(sound_ru_url)
sound_ru = librosa.load(sound_ru_name, sr=target_sampling_rate, mono=True)[0]
print('Duration of sound with Russian speech = {0:.3f} seconds.'.format(
sound_ru.shape[0] / target_sampling_rate
))
detected_languages = detect_language_in_speech(
sound_ru,
asr.feature_extractor,
asr.tokenizer,
asr.model
)
print('Top-3 languages:')
lang_text_width = max([len(it[0]) for it in detected_languages])
for it in detected_languages[0:3]:
print(' {0:>{1}} {2:.4f}'.format(it[0], lang_text_width, it[1]))
recognition_result = asr(
sound_ru,
generate_kwargs={'task': 'transcribe', 'language': detected_languages[0][0]},
return_timestamps=False
)
print(recognition_result['text'] + '\n')
# An example of speech recognition in English, pronounced by a non-native speaker of that language with an accent
sound_en_url = 'https://huggingface.co/bond005/whisper-podlodka-turbo/resolve/main/test_sound_en.wav'
sound_en_name = wget.download(sound_en_url)
sound_en = librosa.load(sound_en_name, sr=target_sampling_rate, mono=True)[0]
print('Duration of sound with English speech = {0:.3f} seconds.'.format(
sound_en.shape[0] / target_sampling_rate
))
detected_languages = detect_language_in_speech(
sound_en,
asr.feature_extractor,
asr.tokenizer,
asr.model
)
print('Top-3 languages:')
lang_text_width = max([len(it[0]) for it in detected_languages])
for it in detected_languages[0:3]:
print(' {0:>{1}} {2:.4f}'.format(it[0], lang_text_width, it[1]))
recognition_result = asr(
sound_en,
generate_kwargs={'task': 'transcribe', 'language': detected_languages[0][0]},
return_timestamps=False
)
print(recognition_result['text'] + '\n')
As a result, you can see a text output like this:
Duration of sound with Russian speech = 29.947 seconds.
Top-3 languages:
russian 0.9568
english 0.0372
ukrainian 0.0013
ΠΡ, Π²ΠΈΡΠΏΠ΅Ρ ΡΠ°ΠΌ ΠΏΠΎ ΡΠ΅Π±Π΅. Π§ΡΠΎ ΡΠ°ΠΊΠΎΠ΅ Π²ΠΈΡΠΏΠ΅Ρ? ΠΠΈΡΠΏΠ΅Ρ β ΡΡΠΎ ΡΠΆΠ΅ ΠΏΠΎΠ»Π½ΠΎΡΠ΅Π½Π½ΠΎΠ΅ end-to-end Π½Π΅ΠΉΡΠΎΡΠ΅ΡΠ΅Π²ΠΎΠ΅ ΡΠ΅ΡΠ΅Π½ΠΈΠ΅ Ρ Π°Π²ΡΠΎΡΠ΅Π³ΡΠ΅ΡΡΠΈΠΎΠ½Π½ΡΠΌ Π΄Π΅ΠΊΠΎΠ΄Π΅ΡΠΎΠΌ, ΡΠΎ Π΅ΡΡΡ ΡΡΠΎ Π½Π΅ ΡΠΈΡΡΡΠΉ ΡΠ½ΠΊΠΎΠ΄Π΅Ρ, ΠΊΠ°ΠΊ Wave2Vec, ΡΡΠΎ Π½Π΅ ΠΏΡΠΎΡΡΠΎ ΡΠ΅ΠΊΡΡΠΎΠ²ΡΠΉ ΡΠ΅ΠΊ-ΡΠΎ-ΡΠ΅ΠΊ, ΡΠ½ΠΊΠΎΠ΄Π΅Ρ-Π΄Π΅ΠΊΠΎΠ΄Π΅Ρ, ΠΊΠ°ΠΊ T5, ΡΡΠΎ ΠΏΠΎΠ»Π½ΠΎΡΠ΅Π½Π½ΡΠΉ Π°Π»Π³ΠΎΡΠΈΡΠΌ ΠΏΡΠ΅ΠΎΠ±ΡΠ°Π·ΠΎΠ²Π°Π½ΠΈΡ ΡΠ΅ΡΠΈ Π² ΡΠ΅ΠΊΡΡ, Π³Π΄Π΅ ΡΠ½ΠΊΠΎΠ΄Π΅Ρ ΡΡΠΈΡΡΠ²Π°Π΅Ρ, ΠΏΡΠ΅ΠΆΠ΄Π΅ Π²ΡΠ΅Π³ΠΎ, Π°ΠΊΡΡΡΠΈΡΠ΅ΡΠΊΠΈΠ΅ ΡΠΈΡΠΈ ΡΠ΅ΡΠΈ, Π½Ρ ΠΈ ΡΠ΅ΠΌΠ°Π½ΡΠΈΠΊΠ° ΡΠΎΠΆΠ΅ ΠΏΠΎΡΡΠ΅ΠΏΠ΅Π½Π½ΠΎ ΠΏΠΎΠ΄ΠΌΠ΅ΡΠΈΠ²Π°Π΅ΡΡΡ, Π° Π΄Π΅ΠΊΠΎΠ΄Π΅Ρ β ΡΡΠΎ ΡΠΆΠ΅ ΡΠ·ΡΠΊΠΎΠ²Π°Ρ ΠΌΠΎΠ΄Π΅Π»Ρ, ΠΊΠΎΡΠΎΡΠ°Ρ Π³Π΅Π½Π΅ΡΠΈΡΡΠ΅Ρ ΡΠΎΠΊΠ΅Π½ Π·Π° ΡΠΎΠΊΠ΅Π½ΠΎΠΌ.
Duration of sound with English speech = 20.247 seconds.
Top-3 languages:
english 0.9526
russian 0.0311
polish 0.0006
Ensembling can help us to solve a well-known bias-variance trade-off. We can decrease variance on basis of large ensemble, large ensemble of different algorithms.
Speech recognition with timestamps
In addition to the usual recognition, the model can also provide timestamps for recognized speech fragments:
recognition_result = asr(
sound_ru,
generate_kwargs={'task': 'transcribe', 'language': 'russian',
return_timestamps=True
)
print('Recognized chunks of Russian speech:')
for it in recognition_result['chunks']:
print(f' {it}')
recognition_result = asr(
sound_en,
generate_kwargs={'task': 'transcribe', 'language': 'english'},
return_timestamps=True
)
print('\nRecognized chunks of English speech:')
for it in recognition_result['chunks']:
print(f' {it}')
As a result, you can see a text output like this:
Recognized chunks of Russian speech:
{'timestamp': (0.0, 4.8), 'text': 'ΠΡ, Π²ΠΈΡΠΏΠ΅Ρ, ΡΠ°ΠΌ ΠΏΠΎ ΡΠ΅Π±Π΅, ΡΡΠΎ ΡΠ°ΠΊΠΎΠ΅ Π²ΠΈΡΠΏΠ΅Ρ. ΠΠΈΡΠΏΠ΅Ρ β ΡΡΠΎ ΡΠΆΠ΅ ΠΏΠΎΠ»Π½ΠΎΡΠ΅Π½Π½ΠΎΠ΅'}
{'timestamp': (4.8, 8.4), 'text': ' end-to-end Π½Π΅ΠΉΡΠΎΡΠ΅ΡΠ΅Π²ΠΎΠ΅ ΡΠ΅ΡΠ΅Π½ΠΈΠ΅ Ρ Π°Π²ΡΠΎΡΠ΅Π³ΡΠ΅ΡΡΠΈΠΎΠ½Π½ΡΠΌ Π΄Π΅ΠΊΠΎΠ΄Π΅ΡΠΎΠΌ.'}
{'timestamp': (8.4, 10.88), 'text': ' Π’ΠΎ Π΅ΡΡΡ, ΡΡΠΎ Π½Π΅ ΡΠΈΡΡΡΠΉ ΡΠ½ΠΊΠΎΠ΄Π΅Ρ, ΠΊΠ°ΠΊ Wave2Vec.'}
{'timestamp': (10.88, 15.6), 'text': ' ΠΡΠΎ Π½Π΅ ΠΏΡΠΎΡΡΠΎ ΡΠ΅ΠΊΡΡΠΎΠ²ΡΠΉ ΡΠ΅ΠΊ-ΡΠΎ-ΡΠ΅ΠΊ, ΡΠ½ΠΊΠΎΠ΄Π΅Ρ-Π΄Π΅ΠΊΠΎΠ΄Π΅Ρ, ΠΊΠ°ΠΊ T5.'}
{'timestamp': (15.6, 19.12), 'text': ' ΠΡΠΎ ΠΏΠΎΠ»Π½ΠΎΡΠ΅Π½Π½ΡΠΉ Π°Π»Π³ΠΎΡΠΈΡΠΌ ΠΏΡΠ΅ΠΎΠ±ΡΠ°Π·ΠΎΠ²Π°Π½ΠΈΡ ΡΠ΅ΡΠΈ Π² ΡΠ΅ΠΊΡΡ,'}
{'timestamp': (19.12, 23.54), 'text': ' Π³Π΄Π΅ ΡΠ½ΠΊΠΎΠ΄Π΅Ρ ΡΡΠΈΡΡΠ²Π°Π΅Ρ, ΠΏΡΠ΅ΠΆΠ΄Π΅ Π²ΡΠ΅Π³ΠΎ, Π°ΠΊΡΡΡΠΈΡΠ΅ΡΠΊΠΈΠ΅ ΡΠΈΡΠΈ ΡΠ΅ΡΠΈ,'}
{'timestamp': (23.54, 25.54), 'text': ' Π½Ρ ΠΈ ΡΠ΅ΠΌΠ°Π½ΡΠΈΠΊΠ° ΡΠΎΠΆΠ΅ ΠΏΠΎΡΡΠ΅ΠΏΠ΅Π½Π½ΠΎ ΠΏΠΎΠ΄ΠΌΠ΅ΡΠΈΠ²Π°Π΅ΡΡΡ,'}
{'timestamp': (25.54, 29.94), 'text': ' Π° Π΄Π΅ΠΊΠΎΠ΄Π΅Ρ β ΡΡΠΎ ΡΠΆΠ΅ ΡΠ·ΡΠΊΠΎΠ²Π°Ρ ΠΌΠΎΠ΄Π΅Π»Ρ, ΠΊΠΎΡΠΎΡΠ°Ρ Π³Π΅Π½Π΅ΡΠΈΡΡΠ΅Ρ ΡΠΎΠΊΠ΅Π½ Π·Π° ΡΠΎΠΊΠ΅Π½ΠΎΠΌ.'}
Recognized chunks of English speech:
{'timestamp': (0.0, 8.08), 'text': 'Ensembling can help us to solve a well-known bias-variance trade-off.'}
{'timestamp': (8.96, 20.08), 'text': 'We can decrease variance on basis of large ensemble, large ensemble of different algorithms.'}
Long-form speech recognition
While previous examples demonstrate accurate transcription for audio segments under thirty seconds, practical applications often require processing extensive recordings ranging from several minutes to multiple hours. This necessitates specialized techniques like the sliding window approach to overcome memory constraints and preserve contextual coherence across the entire signal. The following example showcases the model's capability to handle such long-form audio, enabling accurate transcription of lectures, interviews, and meetings.
import nltk # for splitting long text into sentences
from whisper_lid.whisper_lid import detect_language_in_long_speech # for spoken language detection in long audio
nltk.download('punkt_tab')
long_sound_ru_url = 'https://huggingface.co/bond005/whisper-podlodka-turbo/resolve/main/test_sound_ru_longform.wav'
long_sound_ru_name = wget.download(long_sound_ru_url)
long_sound_ru = librosa.load(long_sound_ru_name, sr=target_sampling_rate, mono=True)[0]
print('Duration of long sound with Russian speech = {0:.3f} seconds.'.format(
sound_ru.shape[0] / target_sampling_rate
))
detected_languages, _ = detect_language_in_long_speech(
long_sound_ru,
asr.feature_extractor,
asr.tokenizer,
asr.model
)
print('\nTop-3 languages:')
lang_text_width = max([len(it[0]) for it in detected_languages])
for it in detected_languages[0:3]:
print(' {0:>{1}} {2:.4f}'.format(it[0], lang_text_width, it[1]))
recognition_result = asr(
sound_ru_longform,
generate_kwargs={
'max_new_tokens': 410,
'num_beams': 5, # beam search width (higher values improve accuracy at the cost of increased computation)
'condition_on_prev_tokens': False,
'compression_ratio_threshold': 2.4, # used to detect and suppress repetitive loops (a common failure mode)
'temperature': (0.0, 0.2, 0.4, 0.6, 0.8, 1.0), # controls the randomness of token sampling during generation
'logprob_threshold': -1.0, # the threshold for the average log-probability of the generated tokens (provides a filter to exclude low-confidence, potentially erroneous transcriptions)
'no_speech_threshold': 0.6, # threshold for the probability of the `<|nospeech|>` token (segments with a probability above this threshold are considered silent and skipped)
'task': 'transcribe',
'language': detected_languages[0][0]
},
return_timestamps=True
)
print('\nRecognized text in the long audio, split into sentences:')
for it in map(lambda sent: sent.strip(), nltk.sent_tokenize(recognition_result['text'])):
print(f' {it}')
As a result, you can see a text output like this:
Duration of long sound with Russian speech = 148.845 seconds.
Top-3 languages:
russian 0.9787
english 0.0186
ukrainian 0.0006
Recognized text in the long audio, split into sentences:
ΠΠ΄ΡΠ°Π²ΡΡΠ²ΡΠΉΡΠ΅, Π΄ΡΡΠ·ΡΡ!
ΠΠ΄ΡΠ°Π²ΡΡΠ²ΡΠΉΡΠ΅!
Π― ΠΎΡΠ΅Π½Ρ ΡΠ°Π΄ Π²Π°Ρ Π²ΡΠ΅Ρ
Π²ΠΈΠ΄Π΅ΡΡ Π·Π΄Π΅ΡΡ, Π² ΡΡΠΎΠΌ Π·Π°Π»Π΅ Π½Π° Π₯Π°ΠΉΠ»ΠΎΠΎΠ΄Π΅.
Π― ΠΠ²Π°Π½, ΠΊΠ°ΠΊ ΠΌΠ΅Π½Ρ ΡΠΆΠ΅ ΠΏΡΠ΅Π΄ΡΡΠ°Π²ΠΈΠ»ΠΈ, ΠΈ Ρ Π»ΡΠ±Π»Ρ ΠΌΠ°ΡΠΈΠ½Π½ΠΎΠ΅ ΠΎΠ±ΡΡΠ΅Π½ΠΈΠ΅.
Π― Π»ΡΠ±Π»Ρ ΠΌΠ°ΡΠΈΠ½Π½ΠΎΠ΅ ΠΎΠ±ΡΡΠ΅Π½ΠΈΠ΅ Ρ 2005 Π³ΠΎΠ΄Π°, ΠΊΠΎΠ³Π΄Π° ΠΎΠ½ΠΎ ΠΏΡΠΎΠ½ΠΈΠΊΠ»ΠΎ Π² ΠΌΠΎΡ ΡΠ΅ΡΠ΄ΡΠ΅ Π΅ΡΡ, ΠΊΠΎΠ³Π΄Π° Ρ Π±ΡΠ» ΡΡΡΠ΄Π΅Π½ΡΠΎΠΌ.
Π‘ 2006 Π³ΠΎΠ΄Π° Ρ ΡΠ°Π±ΠΎΡΠ°Π» Π² Π°ΠΊΠ°Π΄Π΅ΠΌΠΈΡΠ΅ΡΠΊΠΎΠΉ ΡΡΠ΅ΡΠ΅, ΠΏΡΠ΅ΠΏΠΎΠ΄Π°Π²Π°Π» Π½Π΅ΠΉΡΠΎΠ½Π½ΡΠ΅ ΡΠ΅ΡΠΈ, ΠΈΡΠΊΡΡΡΡΠ²Π΅Π½Π½ΡΠΉ ΠΈΠ½ΡΠ΅Π»Π»Π΅ΠΊΡ, ΠΌΠ°ΡΠΈΠ½Π½ΠΎΠ΅ ΠΎΠ±ΡΡΠ΅Π½ΠΈΠ΅ ΡΠ²ΠΎΠΈΠΌ ΡΡΡΠ΄Π΅Π½ΡΠ°ΠΌ.
Π‘ ΡΡΠΈΠ½Π°Π΄ΡΠ°ΡΠΎΠ³ΠΎ Π³ΠΎΠ΄Π° Ρ ΠΏΠ΅ΡΠ΅ΡΡΠ» Π² IT-ΠΈΠ½Π΄ΡΡΡΡΠΈΡ, ΡΠ°Π±ΠΎΡΠ°Π» Π² ΡΠ°Π·Π½ΡΡ
ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΡΡ
, Π·Π°Π½ΠΈΠΌΠ°ΡΡΡ ΠΏΡΠΈΠΌΠ΅ΡΠ½ΠΎ Π²ΡΡ ΡΠ΅ΠΌ ΠΆΠ΅ ΠΌΠ°ΡΠΈΠ½Π½ΡΠΌ ΠΎΠ±ΡΡΠ΅Π½ΠΈΠ΅ΠΌ ΠΈ ΠΈΡΠΊΡΡΡΡΠ²ΠΎΠΌ ΠΈΠ½ΡΠ΅Π»Π»Π΅ΠΊΡΠΎΠΌ.
Π, Π½Π°ΠΊΠΎΠ½Π΅Ρ, Π² Π΄Π²Π°Π΄ΡΠ°ΡΡ Π²ΡΠΎΡΠΎΠΌ Π³ΠΎΠ΄Ρ ΠΌΠ½Π΅ Π²ΡΡ ΡΡΠΎ Π½Π°Π΄ΠΎΠ΅Π»ΠΎ, Ρ ΡΠ΅ΡΠΈΠ» ΠΎΠ±ΡΠ°ΡΠ½ΠΎ ΠΈΠ· IT-ΠΈΠ½Π΄ΡΡΡΡΠΈΠΈ ΠΏΠ΅ΡΠ΅ΠΉΡΠΈ Π² Π°ΠΊΠ°Π΄Π΅ΠΌΠΈΡΠ΅ΡΠΊΡΡ ΡΡΠ΅ΡΡ.
Π ΡΠ΅ΠΉΡΠ°Ρ Ρ ΡΠ°Π±ΠΎΡΠ°Ρ Π² ΠΠΎΠ²ΠΎΡΠΈΠ±ΠΈΡΡΠΊΠΎΠΌ Π³ΠΎΡΡΠ΄Π°ΡΡΡΠ²Π΅Π½Π½ΠΎΠΌ ΡΠ½ΠΈΠ²Π΅ΡΡΠΈΡΠ΅ΡΠ΅, Π·Π°Π½ΠΈΠΌΠ°ΡΡΡ Π½Π°ΡΡΠ½ΡΠΌΠΈ ΠΈΡΡΠ»Π΅Π΄ΠΎΠ²Π°Π½ΠΈΡΠΌΠΈ, ΡΡΡ ΡΡΡΠ΄Π΅Π½ΡΠΎΠ², Π΄Π΅Π»Π°Π΅ΠΌ Π²ΡΡΠΊΠΈΠ΅ ΠΈΠ½ΡΠ΅ΡΠ΅ΡΠ½ΡΠ΅ ΡΡΡΠΊΠΈ.
ΠΡ, Π° Π² ΠΊΠΎΠ½ΡΠ΅ ΠΏΡΠΎΡΠ»ΠΎΠ³ΠΎ Π³ΠΎΠ΄Π° Ρ ΠΈ ΠΌΠΎΠΈ ΡΡΠ΅Π½ΠΈΠΊΠΈ ΡΠ΅ΡΠΈΠ»ΠΈ Π²ΡΡ-ΡΠ°ΠΊΠΈ Π½Π΅ ΡΠΎΠ»ΡΠΊΠΎ ΡΡΠ½Π΄Π°ΠΌΠ΅Π½ΡΠ°Π»ΡΠ½ΡΠΌΠΈ ΠΈΡΡΠ»Π΅Π΄ΠΎΠ²Π°Π½ΠΈΡΠΌΠΈ Π·Π°Π½ΠΈΠΌΠ°ΡΡΡΡ, ΠΈ Π½Π°ΡΠΊΠ° Π΄ΠΎΠ»ΠΆΠ½Π° ΠΏΡΠΎΠ½ΠΎΡΠΈΡΡ ΠΏΠΎΠ»ΡΠ·Ρ Π»ΡΠ΄ΡΠΌ, ΠΈ ΠΌΡ ΡΠ΄Π΅Π»Π°Π»ΠΈ ΠΌΠ°Π»Π΅Π½ΡΠΊΠΈΠΉ ΡΡΠ°ΡΡΠ°ΠΏ ΠΏΠΎΠ΄ Π½Π°Π·Π²Π°Π½ΠΈΠ΅ΠΌ Β«Π‘ΠΈΠ±ΠΈΡΡΠΊΠΈΠ΅ Π½Π΅ΠΉΡΠΎΡΠ΅ΡΠΈΒ».
Π’ΠΎ Π΅ΡΡΡ, ΡΠΈΠ±ΠΈΡΡΠΊΠΎΠ΅ Π·Π΄ΠΎΡΠΎΠ²ΡΠ΅ ΠΈ ΡΠΈΠ±ΠΈΡΡΠΊΠΈΠ΅ Π½Π΅ΠΉΡΠΎΡΠ΅ΡΠΈ ΡΠ΅ΠΏΠ΅ΡΡ.
ΠΡΡ, ΡΡΠΎ Ρ Π΄Π΅Π»Π°Π», ΠΎΠ½ΠΎ ΡΠ²ΡΠ·Π°Π½ΠΎ ΠΎΠ΄Π½ΠΈΠΌ, ΠΌΠΎΠ΅ΠΉ Π»ΡΠ±ΠΎΠ²ΡΡ ΠΊ ΠΌΠ°ΡΠΈΠ½Π½ΠΎΠΌΡ ΠΎΠ±ΡΡΠ΅Π½ΠΈΡ.
ΠΠ½Π΅ ΡΡΠΎ ΠΎΡΠ΅Π½Ρ ΠΈΠ½ΡΠ΅ΡΠ΅ΡΠ½ΠΎ Π±ΡΠ»ΠΎ Π²ΡΠ΅Π³Π΄Π°.
ΠΡ, Π° ΠΏΠΎΠΌΠΈΠΌΠΎ ΠΌΠ°ΡΠΈΠ½Π½ΠΎΠ³ΠΎ ΠΎΠ±ΡΡΠ΅Π½ΠΈΡ, ΠΌΠ½Π΅ ΠΈΠ½ΡΠ΅ΡΠ΅ΡΠ½ΠΎ ΡΡΠ°ΡΡΠ²ΠΎΠ²Π°ΡΡ Π² ΡΠΎΡΠ΅Π²Π½ΠΎΠ²Π°Π½ΠΈΡΡ
.
ΠΠ°ΡΡΠ½ΡΠ΅ ΡΠΎΡΠ΅Π²Π½ΠΎΠ²Π°Π½ΠΈΡ β ΡΡΠΎ Π½Π΅ ΠΏΡΠΎΡΡΠΎ ΡΠΏΠΎΡΠΎΠ± ΡΠ°Π·Π²Π»Π΅ΡΡΡΡ, ΡΡΠΎ ΡΠΏΠΎΡΠΎΠ± ΠΎΡΠ΅Π½ΠΈΡΡ ΠΏΠΎ Π³Π°ΠΌΠ±ΡΡΠ³ΡΠΊΠΎΠΌΡ ΡΡΠ΅ΡΡ, Π½Π°ΡΠΊΠΎΠ»ΡΠΊΠΎ ΡΠ²ΠΎΠΉ Π°Π»Π³ΠΎΡΠΈΡΠΌ Ρ
ΠΎΡΠΎΡ ΠΎΠ±ΡΠ΅ΠΊΡΠΈΠ²Π½ΠΎ Π² ΡΡΠ°Π²Π½Π΅Π½ΠΈΠΈ Ρ Π΄ΡΡΠ³ΠΈΠΌΠΈ.
ΠΠΎΠ³Π΄Π° ΠΌΡ ΡΠ°Π·ΡΠ°Π±Π°ΡΡΠ²Π°Π΅ΠΌ ΠΊΠ°ΠΊΡΡ-ΡΠΎ ΡΠΈΡΡΠ΅ΠΌΡ Π΄Π»Ρ Π·Π°ΠΊΠ°Π·ΡΠΈΠΊΠ°, ΠΌΡ ΠΎΡΠΈΠ΅Π½ΡΠΈΡΡΠ΅ΠΌΡΡ Π½Π° Π΅Π³ΠΎ Π΄Π°Π½Π½ΡΠ΅, ΡΡΠΈ Π΄Π°Π½Π½ΡΠ΅ Π·Π°ΠΊΡΡΡΡ, ΠΊΠ°ΠΊ ΠΏΡΠ°Π²ΠΈΠ»ΠΎ, ΠΌΡ Π΄Π΅Π»Π°Π΅ΠΌ ΠΊΠ°ΠΊΠΈΠ΅-ΡΠΎ ΠΊΠ°ΡΡΠΎΠΌΠΈΠ·Π°ΡΠΈΠΈ, ΠΊΠΎΡΠΎΡΡΠ΅ ΠΎΠ±Π΅ΡΠΏΠ΅ΡΠΈΠ²Π°ΡΡ ΠΊΠ°ΡΠ΅ΡΡΠ²ΠΎ, ΠΌΠΎΠΆΠ΅Ρ, Π΄Π°ΠΆΠ΅ ΡΠ΅ΡΠ΅ΠΏΠΈΠΊΠΈΠ½Π³ ΠΈΠ½ΠΎΠ³Π΄Π° Π΄Π΅Π»Π°Π΅ΠΌ, Ρ
ΠΎΡΡ ΡΡΠΎ ΡΡ, Π½ΠΎ ΡΠ΅ΠΌ Π½Π΅ ΠΌΠ΅Π½Π΅Π΅.
Π ΠΊΠΎΠ³Π΄Π° ΠΌΡ ΠΏΡΠ΅Π΄Π»Π°Π³Π°Π΅ΠΌ Π½Π°ΡΠΈ ΡΠ΅ΡΠ΅Π½ΠΈΡ, Π½Π°Ρ Π½Π°ΡΡΠ½ΡΠΉ ΠΌΠ΅ΡΠΎΠ΄, Π½Π°Ρ Π°Π»Π³ΠΎΡΠΈΡΠΌ Π½Π° ΡΠΎΡΠ΅Π²Π½ΠΎΠ²Π°Π½ΠΈΠΈ, ΠΌΡ Π²ΡΠ΅ Π² ΡΠ°Π²Π½ΡΡ
ΡΡΠ»ΠΎΠ²ΠΈΡΡ
, ΠΈ ΠΌΡ Π΄Π΅ΠΉΡΡΠ²ΠΈΡΠ΅Π»ΡΠ½ΠΎ ΠΌΠΎΠΆΠ΅ΠΌ ΠΎΡΠ΅Π½ΠΈΡΡ, Π½Π°ΡΠΊΠΎΠ»ΡΠΊΠΎ Ρ
ΠΎΡΠΎΡΠΎ ΡΠΎ ΠΈΠ»ΠΈ ΠΈΠ½ΠΎΠ΅ ΡΠ΅ΡΠ΅Π½ΠΈΠ΅ ΡΠ΅Π±Ρ ΠΏΠΎΠΊΠ°Π·ΡΠ²Π°Π΅Ρ.
ΠΡΠΈ ΡΡΠΎΠΌ ΡΠ°ΠΌΠΎΠ΅ Π³Π»Π°Π²Π½ΠΎΠ΅, ΡΡΠΎ ΡΠ°ΠΊΠΈΠ΅ ΡΠΎΡΠ΅Π²Π½ΠΎΠ²Π°Π½ΠΈΡ Π΄Π°ΡΡ ΠΌΠ°ΡΠ΅ΡΠΈΠ°Π»Ρ Π² Π²ΠΈΠ΄Π΅ ΠΎΡΠΊΡΡΡΡΡ
Π΄Π°ΡΠ°ΡΠ΅ΡΠΎΠ² Π΄Π»Ρ Π΄Π°Π»ΡΠ½Π΅ΠΉΡΠ΅ΠΉ Π²ΠΎΡΠΏΡΠΎΠΈΠ·Π²ΠΎΠ΄ΠΈΠΌΠΎΡΡΠΈ, Π² Π²ΠΈΠ΄Π΅ ΠΎΡΠΊΡΡΡΠΎΠ³ΠΎ ΠΊΠΎΠ΄Π°, ΠΊΠΎΡΠΎΡΡΠΉ, ΠΎΠΏΡΡΡ-ΡΠ°ΠΊΠΈ, ΠΌΠΎΠΆΠ½ΠΎ Π²ΠΎΡΠΏΡΠΎΠΈΠ·Π²ΠΎΠ΄ΠΈΡΡ Π² Π½Π°ΡΠΊΠ΅, ΠΏΡΠΎΠ±Π»Π΅ΠΌΠ° Π²ΠΎΡΠΏΡΠΎΠΈΠ·Π²ΠΎΠ΄ΠΈΠΌΠΎΡΡΠΈ ΡΡΠΎΠΈΡ ΠΎΡΡΡΠΎ, Π ΠΏΠΎΠ΄ΠΎΠ±Π½ΡΠ΅ Π½Π°ΡΡΠ½ΡΠ΅ ΡΠΎΡΠ΅Π²Π½ΠΎΠ²Π°Π½ΠΈΡ, Ρ ΠΈΠΌΠ΅Ρ Π² Π²ΠΈΠ΄Ρ Π½Π΅ ΠΠ΅Π³Π», Π° ΡΡΠΎ-ΡΠΎ Π±ΠΎΠ»Π΅Π΅ ΠΈΠ½ΡΠ΅ΡΠ΅ΡΠ½ΠΎΠ΅ ΠΈ Π±ΠΎΠ»Π΅Π΅ ΡΠ°ΠΊΠΎΠ΅ ΡΠΈΡΡΠ΅ΠΌΠ½ΠΎΠ΅, ΠΎΠ½ΠΈ ΠΏΠΎΠ·Π²ΠΎΠ»ΡΡΡ Π΄Π΅ΠΉΡΡΠ²ΠΈΡΠ΅Π»ΡΠ½ΠΎ ΠΊΠ°ΠΊΠΎΠ΅-ΡΠΎ Π΄Π²ΠΈΠΆΠ΅Π½ΠΈΠ΅ Π½Π°ΡΠΊΠΈ Π²ΠΏΠ΅ΡΠ΅Π΄ ΠΎΡΡΡΠ΅ΡΡΠ²Π»ΡΡΡ.
Voice activity detection (speech/non-speech)
Along with special language tokens, the model can also return the special token <|nospeech|>, if the input audio signal does not contain any speech (for details, see section 2.3 of the corresponding paper about Whisper). This skill of the model forms the basis of the speech/non-speech classification algorithm, as demonstrated in the following example:
nonspeech_sound_url = 'https://huggingface.co/bond005/whisper-podlodka-turbo/resolve/main/test_sound_nonspeech.wav'
nonspeech_sound_name = wget.download(nonspeech_sound_url)
nonspeech_sound = librosa.load(nonspeech_sound_name, sr=target_sampling_rate, mono=True)[0]
print('Duration of sound without speech = {0:.3f} seconds.'.format(
nonspeech_sound.shape[0] / target_sampling_rate
))
detected_languages = detect_language_in_speech(
nonspeech_sound,
asr.feature_extractor,
asr.tokenizer,
asr.model
)
print('Top-3 languages:')
lang_text_width = max([len(it[0]) for it in detected_languages])
for it in detected_languages[0:3]:
print(' {0:>{1}} {2:.4f}'.format(it[0], lang_text_width, it[1]))
As a result, you can see a text output like this:
Duration of sound without speech = 10.000 seconds.
Top-3 languages:
NO SPEECH 0.9957
lingala 0.0002
cantonese 0.0002
Speech translation
In addition to the transcription task, the model also performs speech translation (although it translates better from Russian into English than from English into Russian):
print(f'Speech translation from Russian to English:')
recognition_result = asr(
sound_ru,
generate_kwargs={'task': 'translate', 'language': 'english'},
return_timestamps=False
)
print(recognition_result['text'] + '\n')
print(f'Speech translation from English to Russian:')
recognition_result = asr(
sound_en,
generate_kwargs={'task': 'translate', 'language': 'russian'},
return_timestamps=False
)
print(recognition_result['text'] + '\n')
As a result, you can see a text output like this:
Speech translation from Russian to English:
Well, Visper, what is Visper? Visper is already a complete end-to-end neural network with an autoregressive decoder. That is, it's not a pure encoder like Wave2Vec, it's not just a text-to-seq encoder-decoder like T5, it's a complete algorithm for the transformation of speech into text, where the encoder considers, first of all, acoustic features of speech, well, and the semantics are also gradually moving, and the decoder is already a language model that generates token by token.
Speech translation from English to Russian:
ΠΠ½ΡΠ΅ΠΌΠ±Π»ΠΈΠ½Π³ ΠΌΠΎΠΆΠ΅Ρ ΠΏΠΎΠΌΠΎΡΡ Π½Π°ΠΌ ΠΎΡΡΡΠ΅ΡΡΠ²Π»ΡΡΡ Ρ
ΠΎΡΠΎΡΠΎ ΠΈΠ·Π²Π΅ΡΡΠ½ΡΠΉ ΡΠΎΡΠ³ΠΎΠ²ΡΠΉ Π±Π°ΠΉΠ·-Π²Π°ΡΠΈΠ°Π½Ρ. ΠΡ ΠΌΠΎΠΆΠ΅ΠΌ ΠΎΠ³ΡΠ°Π½ΠΈΡΠΈΡΡ Π²Π°ΡΠΈΠ°Π½ΡΡ Π½Π° ΠΎΡΠ½ΠΎΠ²Π΅ ΠΊΡΡΠΏΠ½ΠΎΠ³ΠΎ ΡΠ½ΡΠ΅ΠΌΠ±Π»Π°, ΠΊΡΡΠΏΠ½ΠΎΠ³ΠΎ ΡΠ½ΡΠ΅ΠΌΠ±Π»Π° ΡΠ°Π·Π½ΡΡ
Π°Π»Π³ΠΎΡΠΈΡΠΌΠΎΠ².
As you can see, in both examples the speech translation contains some errors, however in the example of translation from English to Russian these errors are more significant.
Bias, Risks, and Limitations
- While improvements are observed for English and translation tasks, statistically significant advantages are confirmed only for Russian ASR
- The model's performance on code-switching speech (where speakers alternate between Russian and English within the same utterance) has not been specifically evaluated
- Inherits basic limitations of the Whisper architecture
Training Details
Training Data
The model was fine-tuned on a composite dataset including:
- Common Voice (Ru, En)
- Podlodka Speech (Ru)
- Taiga Speech (Ru, synthetic)
- Golos Farfield and Golos Crowd (Ru)
- Sova Rudevices (Ru)
- Audioset (non-speech audio)
Training Features
1. Data Augmentation:
- Dynamic mixing of speech with background noise and music
- Gradual reduction of signal-to-noise ratio during training
2. Text Data Processing:
- Russian text punctuation and capitalization restoration using bond005/ruT5-ASR-large (for speech sub-corpora without punctuated annotations)
- Parallel Russian-English text generation using Qwen/Qwen2.5-14B-Instruct
- Multi-stage validation of generated texts to minimize hallucinations using bond005/xlm-roberta-xl-hallucination-detector
3. Training Strategy:
- Progressive increase in training example complexity
- Balanced sampling between speech and non-speech data
- Special handling of language tokens and no-speech detection (
<|nospeech|>)
Evaluation
The experimental evaluation focused on two main tasks:
- Russian speech recognition
- Speech activity detection (binary classification "speech/non-speech")
Testing was performed on publicly available Russian speech corpora. Speech recognition was conducted using the standard pipeline from the Hugging Face π€ Transformers library. Due to the limitations of this pipeline in language identification and non-speech detection (caused by a certain bug), the whisper-lid library was used for speech presence/absence detection in the signal.
Testing Data & Metrics
Testing Data
The quality of the Russian speech recognition task was tested on test sub-sets of six different datasets:
The quality of the long-form Russian speech recognition was tested on the dangrebenkin/long_audio_youtube_lectures dataset, developed by Daniel Grebenkin. This dataset contains seven long-form (20-40 minute) Russian audio recordings that were manually annotated. The audios cover a variety of topics and speaking styles; they are excerpts from Russian scientific lectures on various subjects: philology, mathematics, history, etc. All recordings were made in relatively quiet, lecture-hall-like acoustic environments. However, some natural background noises, such as the sound of chalk on a blackboard, are present.
The quality of the voice activity detection task was tested on test sub-sets of two different datasets:
- noised version of Golos Crowd as a source of speech samples
- filtered sub-set of Audioset corpus as a source of non-speech samples
Noise was added using a special augmenter capable of simulating the superposition of five different types of acoustic noise (reverberation, speech-like sounds, music, household sounds, and pet sounds) at a given signal-to-noise ratio (in this case, a signal-to-noise ratio of 2 dB was used).
The quality of the robust Russian speech recognition task was tested on test sub-set of above-mentioned noised Golos Crowd.
Metrics
1. Modified WER (Word Error Rate) for Russian speech recognition quality:
- Text normalization before WER calculation:
- Unification of numeral representations (digits/words)
- Standardization of foreign words (Cyrillic/Latin scripts)
- Accounting for valid transliteration variants
- Enables more accurate assessment of semantic recognition accuracy
- The lower the WER, the better the speech recognition quality
2. F1-score for speech activity detection:
- Binary classification "speech/non-speech"
- Evaluation of non-speech segment detection accuracy using
<|nospeech|>token - The higher the F1 score, the better the voice activity detection quality
Generation Parameters
For experiments with short audio signals (under 30 seconds), we used standard greedy decoding (num_beams=1). For long-form audio, two approaches were tested: simple 30-second chunking and the sequential long-form algorithm.
For the sequential long-form mode, the implementation followed the strategy from Section 4.5 of the paper about Whisper with two key hyperparameter differences:
Beam Search: The paper implies the use of beam search for optimal performance, while our initial experiments for this task used greedy decoding (
num_beams=1).Compression Ratio Threshold: A key deviation was the use of a more conservative
compression_ratio_thresholdof 1.35 (compared to 2.4 in the paper). This lower threshold makes the repetition-detection algorithm significantly more aggressive, triggering fallback mechanisms (e.g., temperature rescoring) sooner to suppress repetitive outputs.
The parameters for voice activity detection (no_speech_threshold=0.6) and low-confidence detection (logprob_threshold=-1.0) were kept aligned with the paper's recommendations. Context conditioning between segments (condition_on_prev_tokens) was disabled for this experimental run.
Results
Automatic Speech Recognition (ASR)
Result (WER, %):
| Dataset | bond005/whisper-podlodka-turbo | openai/whisper-large-v3-turbo |
|---|---|---|
| bond005/podlodka_speech | 8.17 | 8.33 |
| rulibrispeech | 9.76 | 10.25 |
| sberdevices_golos_farfield | 11.61 | 20.12 |
| sberdevices_golos_crowd | 11.85 | 14.55 |
| sova_rudevices | 15.35 | 17.70 |
| common_voice_11_0 | 5.22 | 6.63 |
Long-form ASR
Result (WER, %):
| Dataset | bond005/whisper-podlodka-turbo | openai/whisper-large-v3-turbo |
|---|---|---|
| the simple chunking | 11.66 | 15.98 |
| the sequential long-form algorithm | 7.84 | 9.59 |
Voice Activity Detection (VAD)
Result (F1):
| bond005/whisper-podlodka-turbo | openai/whisper-large-v3-turbo |
|---|---|
| 0.9235 | 0.8484 |
Robust ASR (SNR = 2 dB, speech-like noise, music, etc.)
Result (WER, %):
| Dataset | bond005/whisper-podlodka-turbo | openai/whisper-large-v3-turbo |
|---|---|---|
| sberdevices_golos_crowd (noised) | 46.58 | 75.20 |
Citation
If you use this model in your work, please cite it as:
@misc{whisper-podlodka-turbo,
author = {Ivan Bondarenko},
title = {Whisper-Podlodka-Turbo: Enhanced Whisper Model for Russian ASR},
year = {2025},
publisher = {Hugging Face},
journal = {Hugging Face Model Hub},
howpublished = {\url{https://huggingface.co/bond005/whisper-podlodka-turbo}}
}
- Total size
- 3.25 GB
- Files
- 17
- Last updated
- Jun 29
- Pre-warmed CDN
- US EU US EU