Nemotron 3.5 ASR Streaming 0.6B โ€” ONNX INT4 (CPU, fastest)

INT4 k-quant quantized ONNX model converted from NVIDIA Nemotron 3.5 ASR Streaming 0.6B using Microsoft Olive.

Fastest CPU inference โ€” smallest model size, lowest latency. Best for real-time streaming ASR on CPU.

Model Details

Property Value
Base model NVIDIA Nemotron 3.5 ASR Streaming 0.6B
Encoder layers 24
Hidden size 1024
Quantization INT4 k-quant (block_size=32)
Chunk size 0.56s (8,960 samples @ 16kHz)
Sample rate 16,000 Hz
Vocab size 13,088
Languages 80+ (multilingual)
VAD Silero VAD included
Total size ~757 MB

Files

File Size
encoder.onnx + .data 658 MB
decoder.onnx + .data 57 MB
joint.onnx + .data 36 MB
silero_vad.onnx 2.1 MB
tokenizer.json 0.7 MB

Available Variants

Model Precision Encoder Total Size Quality
FP32 FP32 2,380 MB 2,479 MB โญ Best
INT8 INT8 k-quant 922 MB 1,021 MB Good
INT4 INT4 k-quant 658 MB 757 MB โšก Fastest

License

Inherits cc-by-nc-4.0 from NVIDIA.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support