|
Download README.md from Pedro21613/realtime-multilingual-asr-light: direct link, hf CLI and curl.
- Browser
- Download file 1.8 kB
-
https://huggingface.co/Pedro21613/realtime-multilingual-asr-light/resolve/main/README.md
- Command line
-
hf download hf://Pedro21613/realtime-multilingual-asr-light/README.md
-
curl -L -o README.md https://huggingface.co/Pedro21613/realtime-multilingual-asr-light/resolve/main/README.md
1.8 kB
| language: | |
| - multilingual | |
| tags: | |
| - automatic-speech-recognition | |
| - realtime | |
| - whisper | |
| - faster-whisper | |
| - lightweight | |
| - streaming | |
| library_name: faster-whisper | |
| pipeline_tag: automatic-speech-recognition | |
| # Realtime Multilingual ASR — Light (tiny · int8 · CPU) | |
| Modelo **leve** de reconhecimento de fala em **tempo real**, multilíngue (99 idiomas do Whisper), otimizado para CPU. | |
| - Backbone: `faster-whisper tiny` (~39M params, ~75MB, quantizado `int8`) | |
| - VAD filter embutido (não transcreve silêncio) | |
| - Detecção automática de idioma (`pt/en/es/fr/de/...` ou `auto`) | |
| - Streaming por chunks (3s padrão, com overlap) via microfone | |
| - Uso: microfone, arquivo ou Gradio | |
| ## Uso rápido | |
| ```bash | |
| pip install -r requirements.txt | |
| python transcribe_file.py sample.wav auto | |
| python app_realtime.py | |
| python app_gradio.py | |
| ``` | |
| ```python | |
| from realtime_asr import LightMultilingualRealtimeASR | |
| import sounddevice as sd | |
| import numpy as np | |
| asr = LightMultilingualRealtimeASR(model_size="tiny", language="auto") | |
| text, lang, prob = asr.transcribe_file("audio.mp3") | |
| print(lang, prob, text) | |
| for seg in asr.stream_microphone(chunk_seconds=3.0): | |
| print(f"[{seg.language}] {seg.text}") | |
| ``` | |
| ## Trocar precisão / tamanho | |
| | `model_size` | Params | Tamanho | Latência CPU | | |
| |---|---|---|---| | |
| | `tiny` (padrão) | 39M | ~75MB | ~0.3-0.8s / 3s áudio | | |
| | `base` | 74M | ~145MB | ~0.8-1.5s | | |
| | `small` | 244M | ~490MB | mais preciso, mais pesado | | |
| ## Estrutura | |
| ``` | |
| realtime_asr/ | |
| __init__.py | |
| recognizer.py # classe principal LightMultilingualRealtimeASR | |
| app_realtime.py # demo microfone tempo real | |
| app_gradio.py # demo web | |
| transcribe_file.py | |
| requirements.txt | |
| ``` | |
| ## Limitações | |
| Modelo `tiny` é rápido mas menos preciso com sotaques/ruído. Para maior precisão use `base` sem mudar o código. | |