File size: 1,801 Bytes
bb2d6a4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
---
language:
- multilingual
tags:
- automatic-speech-recognition
- realtime
- whisper
- faster-whisper
- lightweight
- streaming
library_name: faster-whisper
pipeline_tag: automatic-speech-recognition
---

# Realtime Multilingual ASR — Light (tiny · int8 · CPU)

Modelo **leve** de reconhecimento de fala em **tempo real**, multilíngue (99 idiomas do Whisper), otimizado para CPU.

- Backbone: `faster-whisper tiny` (~39M params, ~75MB, quantizado `int8`)
- VAD filter embutido (não transcreve silêncio)
- Detecção automática de idioma (`pt/en/es/fr/de/...` ou `auto`)
- Streaming por chunks (3s padrão, com overlap) via microfone
- Uso: microfone, arquivo ou Gradio

## Uso rápido

```bash
pip install -r requirements.txt
python transcribe_file.py sample.wav auto
python app_realtime.py
python app_gradio.py
```

```python
from realtime_asr import LightMultilingualRealtimeASR
import sounddevice as sd
import numpy as np

asr = LightMultilingualRealtimeASR(model_size="tiny", language="auto")
text, lang, prob = asr.transcribe_file("audio.mp3")
print(lang, prob, text)

for seg in asr.stream_microphone(chunk_seconds=3.0):
    print(f"[{seg.language}] {seg.text}")
```

## Trocar precisão / tamanho

| `model_size` | Params | Tamanho | Latência CPU |
|---|---|---|---|
| `tiny` (padrão) | 39M | ~75MB | ~0.3-0.8s / 3s áudio |
| `base` | 74M | ~145MB | ~0.8-1.5s |
| `small` | 244M | ~490MB | mais preciso, mais pesado |

## Estrutura

```
realtime_asr/
  __init__.py
  recognizer.py   # classe principal LightMultilingualRealtimeASR
app_realtime.py   # demo microfone tempo real
app_gradio.py     # demo web
transcribe_file.py
requirements.txt
```

## Limitações

Modelo `tiny` é rápido mas menos preciso com sotaques/ruído. Para maior precisão use `base` sem mudar o código.