Audio Classification
Transformers
ONNX
Safetensors
Malay
English
end-of-turn-detection
turn-detection
semantic-vad
endpointing
voice-agent
livekit
whisper
telephony
Instructions to use Scicom-intl/semantic-vad-eot-whisper-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Scicom-intl/semantic-vad-eot-whisper-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Scicom-intl/semantic-vad-eot-whisper-base")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Scicom-intl/semantic-vad-eot-whisper-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Model card: absolute image URLs so the charts render
Browse files
README.md
CHANGED
|
@@ -112,13 +112,13 @@ it was designed for (this set has none).
|
|
| 112 |
| smart-turn v2 (`pipecat-ai/smart-turn-v2`, 95 M wav2vec2) | 74.1 % | 39.6 % | 2 500 ms | 2 000 ms | 0.62 |
|
| 113 |
| VAD baseline (silence timer) | 78.1 % | 43.6 % | 2 250 ms | 1 770 ms | – |
|
| 114 |
|
| 115 |
-

|
| 116 |
|
| 117 |
-

|
| 118 |
|
| 119 |
-

|
| 120 |
|
| 121 |
-

|
| 122 |
|
| 123 |
## Files
|
| 124 |
|
|
|
|
| 112 |
| smart-turn v2 (`pipecat-ai/smart-turn-v2`, 95 M wav2vec2) | 74.1 % | 39.6 % | 2 500 ms | 2 000 ms | 0.62 |
|
| 113 |
| VAD baseline (silence timer) | 78.1 % | 43.6 % | 2 250 ms | 1 770 ms | – |
|
| 114 |
|
| 115 |
+

|
| 116 |
|
| 117 |
+

|
| 118 |
|
| 119 |
+

|
| 120 |
|
| 121 |
+

|
| 122 |
|
| 123 |
## Files
|
| 124 |
|