Nemotron ASR Streaming — Persian (Farsi)

A streaming speech recognition model for Persian (Farsi). It's fine-tuned from NVIDIA's nemotron-3.5-asr-streaming-0.6b on 1,181 hours of openly licensed Persian speech, on the base model's reserved fa-IR language slot.

  • Live or offline: one model transcribes live audio chunk by chunk (cache-aware streaming) and handles whole files too.
  • Look-ahead: 1.12 s for best accuracy, or 0.32 s for lower latency.
  • Accuracy: 8.8% WER on FLEURS Persian, against 24.2% for NVIDIA's existing Persian model. On conversational speech (YouTube, films) it roughly halves the error rate: 26.0% vs 54.4%.
  • Size: 600M parameters, FastConformer encoder + RNN-T decoder, with a new 1,024-piece Persian tokenizer.
  • Code: github.com/mallahyari/nemotron-asr-streaming-farsi has data preparation, training, evaluation and ready-to-run inference scripts, including a local web app for live microphone transcription.

Results

Word error rate (WER) and character error rate (CER), in %. Lower is better. All numbers are from true chunk-by-chunk streaming inference with NeMo's cache-aware streaming script, greedy decoding.

Word error rate by test set

Test set Clips NVIDIA stt_fa_fastconformer_hybrid_large (zero-shot, CTC) This model, 0.32 s look-ahead This model, 1.12 s look-ahead
FLEURS fa test (read speech) 852 24.23 / 7.58 9.03 / 3.03 8.81 / 2.94
Held-out conversational test (YouTube, films) 9,154 54.35 / 30.74 28.04 / 16.74 25.95 / 15.11
Held-out conversational dev 5,791 57.43 / 34.12 31.63 / 19.68 29.59 / 17.94
Common Voice 22 fa test 10,661 not comparable¹ 19.94 / 6.09 19.12 / 5.80
  • Scoring. Reference and hypothesis are normalized the same way: ZWNJ read as a space, punctuation removed, Arabic letter forms folded to Persian, numbers spelled out. That way می‌رود vs می رود is never counted as an error. WER is cross-checked against NeMo's word_error_rate.
  • Streaming vs whole-utterance. Whole-utterance inference with the same attention context agrees with streaming within 0.15 WER points on every set.
  • Held-out sets. The conversational dev/test speakers never appear in training. Their references are subtitle-derived and not hand-verified, so some "errors" are label typos.
  • ¹ stt_fa was trained on Common Voice and reproduces most CV test sentences word for word. This model's training data contains no Common Voice and no FLEURS.

How to use

The model works with 🤗 Transformers (≥ 5.18, files in this repo's root) and with NVIDIA NeMo (nemotron-asr-streaming-farsi.nemo). Both give the same results. The Transformers weights are bit-identical to the .nemo, and on FLEURS (852 clips) it scores 8.81% WER at 1.12 s look-ahead vs 8.77% for NeMo, and 9.02% vs 9.13% at 0.32 s. In streaming, 98 of 100 transcripts were identical.

🤗 Transformers: transcribe a file

from transformers import AutoModelForRNNT, AutoProcessor
from transformers.audio_utils import load_audio

model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")

processor.set_num_lookahead_tokens(13)   # 1.12 s look-ahead (most accurate); 3 = 0.32 s (default), 6, 0
audio = load_audio("audio.mp3", sampling_rate=processor.feature_extractor.sampling_rate)
inputs = processor(audio, sampling_rate=processor.feature_extractor.sampling_rate, language="fa-IR")
inputs = inputs.to(model.device, dtype=model.dtype)
output = model.generate(**inputs, return_dict_in_generate=True)
print(processor.decode(output.sequences, skip_special_tokens=True)[0])

language accepts only "fa-IR", "fa" or "auto", and all three select the Persian prompt the model was trained with. The other languages of the base model aren't supported.

Fine-tuning: use NeMo. Transformers 5.18 can't compute this model's training loss. It has no loss entry for Nemotron3_5AsrForRNNT and falls back to a causal-LM loss; NVIDIA's original model has the same limitation. The processor's labels/decoder_input_ids are correct for this model, so this will work once Transformers adds the loss.

🤗 Transformers: live streaming

from threading import Thread
import numpy as np
from transformers import AutoModelForRNNT, AutoProcessor, TextIteratorStreamer
from transformers.audio_utils import load_audio

model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
processor.set_num_lookahead_tokens(13)

sr = processor.feature_extractor.sampling_rate
audio = load_audio("audio.mp3", sampling_rate=sr)
# Chunks are processed whole and the model holds back its look-ahead frames, so append two chunks of
# silence; otherwise the last ~1 s of speech is never transcribed. (A live app does this when it stops.)
audio = np.concatenate([audio, np.zeros(2 * processor.num_samples_per_audio_chunk, dtype=np.float32)])
first = processor(audio[: processor.num_samples_first_audio_chunk], sampling_rate=sr, is_streaming=True,
                  is_first_audio_chunk=True, language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)

def chunks():   # in a live app, feed microphone audio here as it arrives
    yield first.input_features[:, : processor.num_mel_frames_first_audio_chunk, :]
    mel, hop, n_fft = processor.num_mel_frames_first_audio_chunk, processor.feature_extractor.hop_length, processor.feature_extractor.n_fft
    start = mel * hop - n_fft // 2
    while (end := start + processor.num_samples_per_audio_chunk) < audio.shape[0]:
        x = processor(audio[start:end], sampling_rate=sr, is_streaming=True, is_first_audio_chunk=False,
                      language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
        yield x.input_features
        mel += processor.num_mel_frames_per_audio_chunk
        start = mel * hop - n_fft // 2

# group_tokens=False is required: the tokenizer's decode merges repeated tokens by default (a CTC rule),
# which would drop letters from RNN-T output. processor.decode sets it for you; the streamer doesn't.
streamer = TextIteratorStreamer(processor.tokenizer, skip_special_tokens=True, group_tokens=False)
Thread(target=model.generate, kwargs={**first, "input_features": chunks(), "streamer": streamer}).start()
for text in streamer:
    print(text, end="", flush=True)

NeMo

Install NeMo 3.0 (pip install "nemo_toolkit[asr]>=3.0"; Python ≥ 3.11). Then download the .nemo file:

from huggingface_hub import hf_hub_download
model_path = hf_hub_download("mehdi-hf/nemotron-asr-streaming-farsi", "nemotron-asr-streaming-farsi.nemo")

Set the fa-IR prompt in NeMo. This model was fine-tuned only with the fa-IR language prompt.

  • In streaming, call model.set_inference_prompt("fa-IR").
  • Don't use model.transcribe() or NeMo's speech_to_text_eval.py as they are. In NeMo 3.0 their dataloader picks the prompt at random per utterance (the untrained auto prompt about half the time), which gives much worse, non-reproducible output.
  • The Transformers version doesn't have this problem: it always uses the Persian prompt.

NeMo streaming (Python)

import librosa, soundfile as sf, torch
from omegaconf import OmegaConf
from nemo.collections.asr.models import ASRModel
from nemo.collections.asr.parts.submodules.rnnt_decoding import RNNTDecodingConfig
from nemo.collections.asr.parts.utils.streaming_utils import CacheAwareStreamingAudioBuffer

model = ASRModel.restore_from(model_path, map_location="cpu").eval()   # or "cuda" / "mps"
model.encoder.set_default_att_context_size([56, 13])   # 1.12 s look-ahead; [56, 3] for 0.32 s
model.change_decoding_strategy(OmegaConf.structured(RNNTDecodingConfig(fused_batch_size=-1)))
model.set_inference_prompt("fa-IR")                    # required: Persian language prompt

audio, sr = sf.read("audio.wav", dtype="float32")
if audio.ndim > 1:
    audio = audio.mean(axis=1)
if sr != 16000:
    audio = librosa.resample(audio, orig_sr=sr, target_sr=16000)

buffer = CacheAwareStreamingAudioBuffer(model=model)
buffer.append_audio(audio)
cache = model.encoder.get_initial_cache_state(batch_size=1)
hyps, pred_out = None, None
with torch.inference_mode():
    for step, (chunk, length) in enumerate(buffer):
        pred_out, texts, *cache, hyps = model.conformer_stream_step(
            processed_signal=chunk, processed_signal_length=length,
            cache_last_channel=cache[0], cache_last_time=cache[1], cache_last_channel_len=cache[2],
            keep_all_outputs=buffer.is_buffer_empty(), previous_hypotheses=hyps, previous_pred_out=pred_out,
            drop_extra_pre_encoded=0 if step == 0 else model.encoder.streaming_cfg.drop_extra_pre_encoded,
            return_transcription=True)
        print(texts[0].text)   # transcript so far

The loop prints the transcript as it grows, chunk by chunk. On an Apple M1 Pro (mps) it runs about 5–10× faster than real time.

NeMo streaming (command line)

git clone --depth 1 --branch v3.0.0 https://github.com/NVIDIA-NeMo/Speech.git nemo-speech
python nemo-speech/examples/asr/asr_cache_aware_streaming/speech_to_text_cache_aware_streaming_infer.py \
  model_path=nemotron-asr-streaming-farsi.nemo dataset_manifest=manifest.json output_path=out/ \
  target_lang=fa-IR att_context_size=[56,13] compute_dtype=float32 batch_size=32

manifest.json is NeMo JSON lines ({"audio_filepath": ..., "duration": ..., "text": ...}). Cache-aware streaming runs in float32. On one H100 it streamed 8.3 h of audio in 82 s.

What the model outputs

  • Persian script, no punctuation, with standard ZWNJ spelling (می‌شه, کتاب‌ها).
  • Numbers as spoken words, e.g. یازده و سی و پنج rather than 11:35. Convert to digits afterwards if you need them.
  • No Latin letters. English words in the training transcripts were removed, so English speech comes out as Persian-script approximations or is dropped.
  • Persian only. The tokenizer was replaced with a Persian one, so the base model's other languages are not supported.

Training

The full recipe, with exact commands, is in REPRODUCE.md.

Base model nvidia/nemotron-3.5-asr-streaming-0.6b (cache-aware FastConformer + RNN-T, multi-look-ahead [[56,3],[56,0],[56,6],[56,13]])
Tokenizer New Persian SentencePiece unigram, 1,024 pieces (nmt_nfkc, ZWNJ kept as a symbol). The RNN-T decoder and joint were re-initialized; the encoder and language-prompt layer keep their pretrained weights
Language prompt fa-IR (prompt index 38, previously untrained), prompt_mode=langID
Data 782,938 clips / 1,181.5 h at 16 kHz; clips longer than 20 s (51) skipped in training (see below)
Schedule 15,000 steps (~30 passes), AdamW, LR 1e-4 cosine to 1e-6, 300 warmup steps
Batching Lhotse duration bucketing (25 buckets), batch sizes from NeMo's OOMptimizer (42–894 clips per GPU)
Hardware 8× NVIDIA H100 80 GB (Google Cloud spot), 8.9 h, NeMo Speech 3.0.0 (nvcr.io/nvidia/nemo-speech:26.07.00)
Checkpoint weights at the final step (dev WER 32.6% with NeMo's in-training scoring)

Dev WER during training

Data

All four sources are openly licensed:

Source License Hours used (train)
farsi-asr/farsi-asr-dataset (YouTube, manually created subtitles) MIT 649.3
PerSets/youtube-persian-asr CC0 260.6
PerSets/filimo-persian-asr (films and series) CC0 207.3
MahtaFetrat/Mana-TTS (one narrator, hand-verified) CC0 64.3

The raw data was cut and filtered to 1,181 h:

  • Label checks: every training clip was transcribed with NVIDIA's stt_fa model, and clips whose subtitle disagreed badly with it were dropped (CER above 0.8, or 0.6 for the two YouTube sources).
  • Language filter: Whisper large-v3 language ID removed clips of English speech.
  • Held-out split: 30 test and 20 dev speakers per source were held out entirely.
  • Listening checks: two blind checks put usable labels at 99.7% (95% CI 99.2–100%).
  • Text normalization: Arabic letter forms were folded to Persian, numbers spelled out, punctuation removed, and spelling unified, so می‌رود, می رود and میرود all become می‌رود.

Limitations

  • Conversational speech is still hard: ~26% WER on spontaneous film and YouTube speech with music and noise, against ~9% on read speech.
  • Spacing variants still cost a little WER: compound words written with or without a space (چندتا / چند تا) count as errors.
  • End of a stream (Transformers): Transformers 5.18's streaming accepts only full-size chunks, so the end of the audio is padded with silence. A word cut off by the very end of a recording can then be dropped; whole-file mode and NeMo's streaming keep it.
  • Not evaluated on dialects: the training data is Iranian media and podcast speech. Performance on regional accents, Dari or Tajik hasn't been measured.
  • Evaluation coverage: FLEURS Persian test speakers are all male.

License

This model is a derivative of nvidia/nemotron-3.5-asr-streaming-0.6b and is distributed under the OpenMDW License Agreement, version 1.1, the base model's license. A copy is in LICENSE. The training data is MIT / CC0, as listed above. See NOTICE for attributions.

Downloads last month
79
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mehdi-hf/nemotron-asr-streaming-farsi

Finetuned
(59)
this model

Datasets used to train mehdi-hf/nemotron-asr-streaming-farsi

Space using mehdi-hf/nemotron-asr-streaming-farsi 1

Evaluation results