Instructions to use nvidia/parakeet-tdt-0.6b-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/parakeet-tdt-0.6b-v3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="nvidia/parakeet-tdt-0.6b-v3")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nvidia/parakeet-tdt-0.6b-v3", device_map="auto") - Inference
- Notebooks
- Google Colab
- Kaggle
Streaming question
what are the best settings for the model streaming at low latency?
Left Context (seconds)
Larger values improve quality
Chunk Size (seconds)
Processing chunk size
Right Context (seconds)
Future context for better accuracy
help will be appreciated
what are the best settings for the model streaming at low latency?
Left Context (seconds)
Larger values improve qualityChunk Size (seconds)
Processing chunk sizeRight Context (seconds)
Future context for better accuracyhelp will be appreciated
10, 2, 2 by default (https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/asr/streaming_decoding/canary_chunked_and_streaming_decoding.html)
This is from Canary model NVIDIA webpage but I have tested on Parakeet also. BTW share your code if its not a problem, these values are the least of the problems in overall.
Anyway, the larger the better but less real-time, and what do you mean Streaming, which decoding method :) Interesting topic, if you looking for someone to together build something onParakeet let meknow:)
One thing worth adding to the 10/2/2 starting point. Right context is the piece that sets your latency floor, because you cannot emit a token until those future frames have arrived. Left context costs memory and compute but adds no delay. So when tuning for responsiveness, trim right context first and keep left context generous.