Instructions to use mehdi-hf/nemotron-asr-streaming-farsi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mehdi-hf/nemotron-asr-streaming-farsi with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mehdi-hf/nemotron-asr-streaming-farsi") transcriptions = asr_model.transcribe(["file.wav"]) - Transformers
How to use mehdi-hf/nemotron-asr-streaming-farsi with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="mehdi-hf/nemotron-asr-streaming-farsi")# Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("mehdi-hf/nemotron-asr-streaming-farsi") model = AutoModel.from_pretrained("mehdi-hf/nemotron-asr-streaming-farsi", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Nemotron ASR Streaming — Persian (Farsi)
A streaming speech recognition model for Persian (Farsi). It's fine-tuned from NVIDIA's nemotron-3.5-asr-streaming-0.6b on 1,181 hours of openly licensed Persian speech, on the base model's reserved fa-IR language slot.
- Live or offline: one model transcribes live audio chunk by chunk (cache-aware streaming) and handles whole files too.
- Look-ahead: 1.12 s for best accuracy, or 0.32 s for lower latency.
- Accuracy: 8.8% WER on FLEURS Persian, against 24.2% for NVIDIA's existing Persian model. On conversational speech (YouTube, films) it roughly halves the error rate: 26.0% vs 54.4%.
- Size: 600M parameters, FastConformer encoder + RNN-T decoder, with a new 1,024-piece Persian tokenizer.
- Code: github.com/mallahyari/nemotron-asr-streaming-farsi has data preparation, training, evaluation and ready-to-run inference scripts, including a local web app for live microphone transcription.
Results
Word error rate (WER) and character error rate (CER), in %. Lower is better. All numbers are from true chunk-by-chunk streaming inference with NeMo's cache-aware streaming script, greedy decoding.
| Test set | Clips | NVIDIA stt_fa_fastconformer_hybrid_large (zero-shot, CTC) |
This model, 0.32 s look-ahead | This model, 1.12 s look-ahead |
|---|---|---|---|---|
| FLEURS fa test (read speech) | 852 | 24.23 / 7.58 | 9.03 / 3.03 | 8.81 / 2.94 |
| Held-out conversational test (YouTube, films) | 9,154 | 54.35 / 30.74 | 28.04 / 16.74 | 25.95 / 15.11 |
| Held-out conversational dev | 5,791 | 57.43 / 34.12 | 31.63 / 19.68 | 29.59 / 17.94 |
| Common Voice 22 fa test | 10,661 | not comparable¹ | 19.94 / 6.09 | 19.12 / 5.80 |
- Scoring. Reference and hypothesis are normalized the same way: ZWNJ read as a space, punctuation removed, Arabic letter forms folded to Persian, numbers spelled out. That way
میرودvsمی رودis never counted as an error. WER is cross-checked against NeMo'sword_error_rate. - Streaming vs whole-utterance. Whole-utterance inference with the same attention context agrees with streaming within 0.15 WER points on every set.
- Held-out sets. The conversational dev/test speakers never appear in training. Their references are subtitle-derived and not hand-verified, so some "errors" are label typos.
- ¹
stt_fawas trained on Common Voice and reproduces most CV test sentences word for word. This model's training data contains no Common Voice and no FLEURS.
How to use
The model works with 🤗 Transformers (≥ 5.18, files in this repo's root) and with NVIDIA NeMo (nemotron-asr-streaming-farsi.nemo). Both give the same results. The Transformers weights are bit-identical to the .nemo, and on FLEURS (852 clips) it scores 8.81% WER at 1.12 s look-ahead vs 8.77% for NeMo, and 9.02% vs 9.13% at 0.32 s. In streaming, 98 of 100 transcripts were identical.
🤗 Transformers: transcribe a file
from transformers import AutoModelForRNNT, AutoProcessor
from transformers.audio_utils import load_audio
model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
processor.set_num_lookahead_tokens(13) # 1.12 s look-ahead (most accurate); 3 = 0.32 s (default), 6, 0
audio = load_audio("audio.mp3", sampling_rate=processor.feature_extractor.sampling_rate)
inputs = processor(audio, sampling_rate=processor.feature_extractor.sampling_rate, language="fa-IR")
inputs = inputs.to(model.device, dtype=model.dtype)
output = model.generate(**inputs, return_dict_in_generate=True)
print(processor.decode(output.sequences, skip_special_tokens=True)[0])
language accepts only "fa-IR", "fa" or "auto", and all three select the Persian prompt the model was trained with. The other languages of the base model aren't supported.
Fine-tuning: use NeMo. Transformers 5.18 can't compute this model's training loss. It has no loss entry for Nemotron3_5AsrForRNNT and falls back to a causal-LM loss; NVIDIA's original model has the same limitation. The processor's labels/decoder_input_ids are correct for this model, so this will work once Transformers adds the loss.
🤗 Transformers: live streaming
from threading import Thread
import numpy as np
from transformers import AutoModelForRNNT, AutoProcessor, TextIteratorStreamer
from transformers.audio_utils import load_audio
model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
processor.set_num_lookahead_tokens(13)
sr = processor.feature_extractor.sampling_rate
audio = load_audio("audio.mp3", sampling_rate=sr)
# Chunks are processed whole and the model holds back its look-ahead frames, so append two chunks of
# silence; otherwise the last ~1 s of speech is never transcribed. (A live app does this when it stops.)
audio = np.concatenate([audio, np.zeros(2 * processor.num_samples_per_audio_chunk, dtype=np.float32)])
first = processor(audio[: processor.num_samples_first_audio_chunk], sampling_rate=sr, is_streaming=True,
is_first_audio_chunk=True, language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
def chunks(): # in a live app, feed microphone audio here as it arrives
yield first.input_features[:, : processor.num_mel_frames_first_audio_chunk, :]
mel, hop, n_fft = processor.num_mel_frames_first_audio_chunk, processor.feature_extractor.hop_length, processor.feature_extractor.n_fft
start = mel * hop - n_fft // 2
while (end := start + processor.num_samples_per_audio_chunk) < audio.shape[0]:
x = processor(audio[start:end], sampling_rate=sr, is_streaming=True, is_first_audio_chunk=False,
language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
yield x.input_features
mel += processor.num_mel_frames_per_audio_chunk
start = mel * hop - n_fft // 2
# group_tokens=False is required: the tokenizer's decode merges repeated tokens by default (a CTC rule),
# which would drop letters from RNN-T output. processor.decode sets it for you; the streamer doesn't.
streamer = TextIteratorStreamer(processor.tokenizer, skip_special_tokens=True, group_tokens=False)
Thread(target=model.generate, kwargs={**first, "input_features": chunks(), "streamer": streamer}).start()
for text in streamer:
print(text, end="", flush=True)
NeMo
Install NeMo 3.0 (pip install "nemo_toolkit[asr]>=3.0"; Python ≥ 3.11). Then download the .nemo file:
from huggingface_hub import hf_hub_download
model_path = hf_hub_download("mehdi-hf/nemotron-asr-streaming-farsi", "nemotron-asr-streaming-farsi.nemo")
Set the
fa-IRprompt in NeMo. This model was fine-tuned only with thefa-IRlanguage prompt.
- In streaming, call
model.set_inference_prompt("fa-IR").- Don't use
model.transcribe()or NeMo'sspeech_to_text_eval.pyas they are. In NeMo 3.0 their dataloader picks the prompt at random per utterance (the untrainedautoprompt about half the time), which gives much worse, non-reproducible output.- The Transformers version doesn't have this problem: it always uses the Persian prompt.
NeMo streaming (Python)
import librosa, soundfile as sf, torch
from omegaconf import OmegaConf
from nemo.collections.asr.models import ASRModel
from nemo.collections.asr.parts.submodules.rnnt_decoding import RNNTDecodingConfig
from nemo.collections.asr.parts.utils.streaming_utils import CacheAwareStreamingAudioBuffer
model = ASRModel.restore_from(model_path, map_location="cpu").eval() # or "cuda" / "mps"
model.encoder.set_default_att_context_size([56, 13]) # 1.12 s look-ahead; [56, 3] for 0.32 s
model.change_decoding_strategy(OmegaConf.structured(RNNTDecodingConfig(fused_batch_size=-1)))
model.set_inference_prompt("fa-IR") # required: Persian language prompt
audio, sr = sf.read("audio.wav", dtype="float32")
if audio.ndim > 1:
audio = audio.mean(axis=1)
if sr != 16000:
audio = librosa.resample(audio, orig_sr=sr, target_sr=16000)
buffer = CacheAwareStreamingAudioBuffer(model=model)
buffer.append_audio(audio)
cache = model.encoder.get_initial_cache_state(batch_size=1)
hyps, pred_out = None, None
with torch.inference_mode():
for step, (chunk, length) in enumerate(buffer):
pred_out, texts, *cache, hyps = model.conformer_stream_step(
processed_signal=chunk, processed_signal_length=length,
cache_last_channel=cache[0], cache_last_time=cache[1], cache_last_channel_len=cache[2],
keep_all_outputs=buffer.is_buffer_empty(), previous_hypotheses=hyps, previous_pred_out=pred_out,
drop_extra_pre_encoded=0 if step == 0 else model.encoder.streaming_cfg.drop_extra_pre_encoded,
return_transcription=True)
print(texts[0].text) # transcript so far
The loop prints the transcript as it grows, chunk by chunk. On an Apple M1 Pro (mps) it runs about 5–10× faster than real time.
NeMo streaming (command line)
git clone --depth 1 --branch v3.0.0 https://github.com/NVIDIA-NeMo/Speech.git nemo-speech
python nemo-speech/examples/asr/asr_cache_aware_streaming/speech_to_text_cache_aware_streaming_infer.py \
model_path=nemotron-asr-streaming-farsi.nemo dataset_manifest=manifest.json output_path=out/ \
target_lang=fa-IR att_context_size=[56,13] compute_dtype=float32 batch_size=32
manifest.json is NeMo JSON lines ({"audio_filepath": ..., "duration": ..., "text": ...}). Cache-aware streaming runs in float32. On one H100 it streamed 8.3 h of audio in 82 s.
What the model outputs
- Persian script, no punctuation, with standard ZWNJ spelling (
میشه,کتابها). - Numbers as spoken words, e.g.
یازده و سی و پنجrather than11:35. Convert to digits afterwards if you need them. - No Latin letters. English words in the training transcripts were removed, so English speech comes out as Persian-script approximations or is dropped.
- Persian only. The tokenizer was replaced with a Persian one, so the base model's other languages are not supported.
Training
The full recipe, with exact commands, is in REPRODUCE.md.
| Base model | nvidia/nemotron-3.5-asr-streaming-0.6b (cache-aware FastConformer + RNN-T, multi-look-ahead [[56,3],[56,0],[56,6],[56,13]]) |
| Tokenizer | New Persian SentencePiece unigram, 1,024 pieces (nmt_nfkc, ZWNJ kept as a symbol). The RNN-T decoder and joint were re-initialized; the encoder and language-prompt layer keep their pretrained weights |
| Language prompt | fa-IR (prompt index 38, previously untrained), prompt_mode=langID |
| Data | 782,938 clips / 1,181.5 h at 16 kHz; clips longer than 20 s (51) skipped in training (see below) |
| Schedule | 15,000 steps (~30 passes), AdamW, LR 1e-4 cosine to 1e-6, 300 warmup steps |
| Batching | Lhotse duration bucketing (25 buckets), batch sizes from NeMo's OOMptimizer (42–894 clips per GPU) |
| Hardware | 8× NVIDIA H100 80 GB (Google Cloud spot), 8.9 h, NeMo Speech 3.0.0 (nvcr.io/nvidia/nemo-speech:26.07.00) |
| Checkpoint | weights at the final step (dev WER 32.6% with NeMo's in-training scoring) |
Data
All four sources are openly licensed:
| Source | License | Hours used (train) |
|---|---|---|
farsi-asr/farsi-asr-dataset (YouTube, manually created subtitles) |
MIT | 649.3 |
PerSets/youtube-persian-asr |
CC0 | 260.6 |
PerSets/filimo-persian-asr (films and series) |
CC0 | 207.3 |
MahtaFetrat/Mana-TTS (one narrator, hand-verified) |
CC0 | 64.3 |
The raw data was cut and filtered to 1,181 h:
- Label checks: every training clip was transcribed with NVIDIA's
stt_famodel, and clips whose subtitle disagreed badly with it were dropped (CER above 0.8, or 0.6 for the two YouTube sources). - Language filter: Whisper large-v3 language ID removed clips of English speech.
- Held-out split: 30 test and 20 dev speakers per source were held out entirely.
- Listening checks: two blind checks put usable labels at 99.7% (95% CI 99.2–100%).
- Text normalization: Arabic letter forms were folded to Persian, numbers spelled out, punctuation removed, and spelling unified, so
میرود,می رودandمیرودall becomeمیرود.
Limitations
- Conversational speech is still hard: ~26% WER on spontaneous film and YouTube speech with music and noise, against ~9% on read speech.
- Spacing variants still cost a little WER: compound words written with or without a space (
چندتا/چند تا) count as errors. - End of a stream (Transformers): Transformers 5.18's streaming accepts only full-size chunks, so the end of the audio is padded with silence. A word cut off by the very end of a recording can then be dropped; whole-file mode and NeMo's streaming keep it.
- Not evaluated on dialects: the training data is Iranian media and podcast speech. Performance on regional accents, Dari or Tajik hasn't been measured.
- Evaluation coverage: FLEURS Persian test speakers are all male.
License
This model is a derivative of nvidia/nemotron-3.5-asr-streaming-0.6b and is distributed under the OpenMDW License Agreement, version 1.1, the base model's license. A copy is in LICENSE. The training data is MIT / CC0, as listed above. See NOTICE for attributions.
- Downloads last month
- 79
Model tree for mehdi-hf/nemotron-asr-streaming-farsi
Base model
nvidia/nemotron-3.5-asr-streaming-0.6bDatasets used to train mehdi-hf/nemotron-asr-streaming-farsi
PerSets/youtube-persian-asr
PerSets/filimo-persian-asr
Space using mehdi-hf/nemotron-asr-streaming-farsi 1
Evaluation results
- WER (streaming, 1.12 s look-ahead) on FLEURS (Persian)test set self-reported8.810
- CER (streaming, 1.12 s look-ahead) on FLEURS (Persian)test set self-reported2.940
- WER (streaming, 1.12 s look-ahead) on Common Voice 22.0 (Persian)test set self-reported19.120
- CER (streaming, 1.12 s look-ahead) on Common Voice 22.0 (Persian)test set self-reported5.800

