Automatic Speech Recognition
NeMo
Safetensors
Transformers
Persian
nemotron3_5_asr
feature-extraction
streaming-asr
cache-aware ASR
FastConformer
RNNT
NeMo
persian
farsi
Eval Results (legacy)
Instructions to use mehdi-hf/nemotron-asr-streaming-farsi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mehdi-hf/nemotron-asr-streaming-farsi with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mehdi-hf/nemotron-asr-streaming-farsi") transcriptions = asr_model.transcribe(["file.wav"]) - Transformers
How to use mehdi-hf/nemotron-asr-streaming-farsi with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="mehdi-hf/nemotron-asr-streaming-farsi")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("mehdi-hf/nemotron-asr-streaming-farsi") model = AutoModel.from_pretrained("mehdi-hf/nemotron-asr-streaming-farsi", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add 🤗 Transformers weights/processor (converted + verified against the .nemo); card: Transformers usage
Browse files- README.md +74 -4
- config.json +52 -0
- generation_config.json +8 -0
- model.safetensors +3 -0
- processor_config.json +29 -0
- tokenizer.json +0 -0
- tokenizer_config.json +9 -0
README.md
CHANGED
|
@@ -19,6 +19,7 @@ tags:
|
|
| 19 |
- FastConformer
|
| 20 |
- RNNT
|
| 21 |
- NeMo
|
|
|
|
| 22 |
- persian
|
| 23 |
- farsi
|
| 24 |
model-index:
|
|
@@ -85,18 +86,86 @@ Word error rate (WER) and character error rate (CER), in %. Lower is better. All
|
|
| 85 |
|
| 86 |
## How to use
|
| 87 |
|
| 88 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 89 |
|
| 90 |
```python
|
| 91 |
from huggingface_hub import hf_hub_download
|
| 92 |
model_path = hf_hub_download("mehdi-hf/nemotron-asr-streaming-farsi", "nemotron-asr-streaming-farsi.nemo")
|
| 93 |
```
|
| 94 |
|
| 95 |
-
> **Set the `fa-IR` prompt.** This model was fine-tuned only with the `fa-IR` language prompt.
|
| 96 |
> - In streaming, call `model.set_inference_prompt("fa-IR")`.
|
| 97 |
> - Don't use `model.transcribe()` or NeMo's `speech_to_text_eval.py` as they are. In NeMo 3.0 their dataloader picks the prompt **at random per utterance** (the untrained `auto` prompt about half the time), which gives much worse, non-reproducible output.
|
|
|
|
|
|
|
|
|
|
| 98 |
|
| 99 |
-
### Streaming (Python)
|
| 100 |
|
| 101 |
```python
|
| 102 |
import librosa, soundfile as sf, torch
|
|
@@ -133,7 +202,7 @@ with torch.inference_mode():
|
|
| 133 |
|
| 134 |
The loop prints the transcript as it grows, chunk by chunk. On an Apple M1 Pro (`mps`) it runs about 5–10× faster than real time.
|
| 135 |
|
| 136 |
-
###
|
| 137 |
|
| 138 |
```bash
|
| 139 |
git clone --depth 1 --branch v3.0.0 https://github.com/NVIDIA-NeMo/Speech.git nemo-speech
|
|
@@ -188,6 +257,7 @@ The raw data was cut and filtered to 1,181 h:
|
|
| 188 |
|
| 189 |
- **Conversational speech is still hard:** ~26% WER on spontaneous film and YouTube speech with music and noise, against ~9% on read speech.
|
| 190 |
- **Spacing variants still cost a little WER:** compound words written with or without a space (`چندتا` / `چند تا`) count as errors.
|
|
|
|
| 191 |
- **Not evaluated on dialects:** the training data is Iranian media and podcast speech. Performance on regional accents, Dari or Tajik hasn't been measured.
|
| 192 |
- **Evaluation coverage:** FLEURS Persian test speakers are all male.
|
| 193 |
|
|
|
|
| 19 |
- FastConformer
|
| 20 |
- RNNT
|
| 21 |
- NeMo
|
| 22 |
+
- transformers
|
| 23 |
- persian
|
| 24 |
- farsi
|
| 25 |
model-index:
|
|
|
|
| 86 |
|
| 87 |
## How to use
|
| 88 |
|
| 89 |
+
The model works with **🤗 Transformers** (≥ 5.18, files in this repo's root) and with **NVIDIA NeMo** (`nemotron-asr-streaming-farsi.nemo`). Both give the same results. The Transformers weights are bit-identical to the `.nemo`, and on FLEURS (852 clips) it scores **8.81% WER at 1.12 s look-ahead vs 8.77%** for NeMo, and 9.02% vs 9.13% at 0.32 s. In streaming, 98 of 100 transcripts were identical.
|
| 90 |
+
|
| 91 |
+
### 🤗 Transformers: transcribe a file
|
| 92 |
+
|
| 93 |
+
```python
|
| 94 |
+
from transformers import AutoModelForRNNT, AutoProcessor
|
| 95 |
+
from transformers.audio_utils import load_audio
|
| 96 |
+
|
| 97 |
+
model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
|
| 98 |
+
processor = AutoProcessor.from_pretrained(model_id)
|
| 99 |
+
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
|
| 100 |
+
|
| 101 |
+
processor.set_num_lookahead_tokens(13) # 1.12 s look-ahead (most accurate); 3 = 0.32 s (default), 6, 0
|
| 102 |
+
audio = load_audio("audio.mp3", sampling_rate=processor.feature_extractor.sampling_rate)
|
| 103 |
+
inputs = processor(audio, sampling_rate=processor.feature_extractor.sampling_rate, language="fa-IR")
|
| 104 |
+
inputs = inputs.to(model.device, dtype=model.dtype)
|
| 105 |
+
output = model.generate(**inputs, return_dict_in_generate=True)
|
| 106 |
+
print(processor.decode(output.sequences, skip_special_tokens=True)[0])
|
| 107 |
+
```
|
| 108 |
+
|
| 109 |
+
`language` accepts only `"fa-IR"`, `"fa"` or `"auto"`, and all three select the Persian prompt the model was trained with. The other languages of the base model aren't supported.
|
| 110 |
+
|
| 111 |
+
**Fine-tuning:** use NeMo. Transformers 5.18 can't compute this model's training loss. It has no loss entry for `Nemotron3_5AsrForRNNT` and falls back to a causal-LM loss; NVIDIA's original model has the same limitation. The processor's `labels`/`decoder_input_ids` are correct for this model, so this will work once Transformers adds the loss.
|
| 112 |
+
|
| 113 |
+
### 🤗 Transformers: live streaming
|
| 114 |
+
|
| 115 |
+
```python
|
| 116 |
+
from threading import Thread
|
| 117 |
+
import numpy as np
|
| 118 |
+
from transformers import AutoModelForRNNT, AutoProcessor, TextIteratorStreamer
|
| 119 |
+
from transformers.audio_utils import load_audio
|
| 120 |
+
|
| 121 |
+
model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
|
| 122 |
+
processor = AutoProcessor.from_pretrained(model_id)
|
| 123 |
+
model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
|
| 124 |
+
processor.set_num_lookahead_tokens(13)
|
| 125 |
+
|
| 126 |
+
sr = processor.feature_extractor.sampling_rate
|
| 127 |
+
audio = load_audio("audio.mp3", sampling_rate=sr)
|
| 128 |
+
# Chunks are processed whole and the model holds back its look-ahead frames, so append two chunks of
|
| 129 |
+
# silence; otherwise the last ~1 s of speech is never transcribed. (A live app does this when it stops.)
|
| 130 |
+
audio = np.concatenate([audio, np.zeros(2 * processor.num_samples_per_audio_chunk, dtype=np.float32)])
|
| 131 |
+
first = processor(audio[: processor.num_samples_first_audio_chunk], sampling_rate=sr, is_streaming=True,
|
| 132 |
+
is_first_audio_chunk=True, language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
|
| 133 |
+
|
| 134 |
+
def chunks(): # in a live app, feed microphone audio here as it arrives
|
| 135 |
+
yield first.input_features[:, : processor.num_mel_frames_first_audio_chunk, :]
|
| 136 |
+
mel, hop, n_fft = processor.num_mel_frames_first_audio_chunk, processor.feature_extractor.hop_length, processor.feature_extractor.n_fft
|
| 137 |
+
start = mel * hop - n_fft // 2
|
| 138 |
+
while (end := start + processor.num_samples_per_audio_chunk) < audio.shape[0]:
|
| 139 |
+
x = processor(audio[start:end], sampling_rate=sr, is_streaming=True, is_first_audio_chunk=False,
|
| 140 |
+
language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
|
| 141 |
+
yield x.input_features
|
| 142 |
+
mel += processor.num_mel_frames_per_audio_chunk
|
| 143 |
+
start = mel * hop - n_fft // 2
|
| 144 |
+
|
| 145 |
+
# group_tokens=False is required: the tokenizer's decode merges repeated tokens by default (a CTC rule),
|
| 146 |
+
# which would drop letters from RNN-T output. processor.decode sets it for you; the streamer doesn't.
|
| 147 |
+
streamer = TextIteratorStreamer(processor.tokenizer, skip_special_tokens=True, group_tokens=False)
|
| 148 |
+
Thread(target=model.generate, kwargs={**first, "input_features": chunks(), "streamer": streamer}).start()
|
| 149 |
+
for text in streamer:
|
| 150 |
+
print(text, end="", flush=True)
|
| 151 |
+
```
|
| 152 |
+
|
| 153 |
+
### NeMo
|
| 154 |
+
|
| 155 |
+
Install [NeMo](https://github.com/NVIDIA-NeMo/Speech) 3.0 (`pip install "nemo_toolkit[asr]>=3.0"`; Python ≥ 3.11). Then download the `.nemo` file:
|
| 156 |
|
| 157 |
```python
|
| 158 |
from huggingface_hub import hf_hub_download
|
| 159 |
model_path = hf_hub_download("mehdi-hf/nemotron-asr-streaming-farsi", "nemotron-asr-streaming-farsi.nemo")
|
| 160 |
```
|
| 161 |
|
| 162 |
+
> **Set the `fa-IR` prompt in NeMo.** This model was fine-tuned only with the `fa-IR` language prompt.
|
| 163 |
> - In streaming, call `model.set_inference_prompt("fa-IR")`.
|
| 164 |
> - Don't use `model.transcribe()` or NeMo's `speech_to_text_eval.py` as they are. In NeMo 3.0 their dataloader picks the prompt **at random per utterance** (the untrained `auto` prompt about half the time), which gives much worse, non-reproducible output.
|
| 165 |
+
> - The Transformers version doesn't have this problem: it always uses the Persian prompt.
|
| 166 |
+
|
| 167 |
+
#### NeMo streaming (Python)
|
| 168 |
|
|
|
|
| 169 |
|
| 170 |
```python
|
| 171 |
import librosa, soundfile as sf, torch
|
|
|
|
| 202 |
|
| 203 |
The loop prints the transcript as it grows, chunk by chunk. On an Apple M1 Pro (`mps`) it runs about 5–10× faster than real time.
|
| 204 |
|
| 205 |
+
#### NeMo streaming (command line)
|
| 206 |
|
| 207 |
```bash
|
| 208 |
git clone --depth 1 --branch v3.0.0 https://github.com/NVIDIA-NeMo/Speech.git nemo-speech
|
|
|
|
| 257 |
|
| 258 |
- **Conversational speech is still hard:** ~26% WER on spontaneous film and YouTube speech with music and noise, against ~9% on read speech.
|
| 259 |
- **Spacing variants still cost a little WER:** compound words written with or without a space (`چندتا` / `چند تا`) count as errors.
|
| 260 |
+
- **End of a stream (Transformers):** Transformers 5.18's streaming accepts only full-size chunks, so the end of the audio is padded with silence. A word cut off by the very end of a recording can then be dropped; whole-file mode and NeMo's streaming keep it.
|
| 261 |
- **Not evaluated on dialects:** the training data is Iranian media and podcast speech. Performance on regional accents, Dari or Tajik hasn't been measured.
|
| 262 |
- **Evaluation coverage:** FLEURS Persian test speakers are all male.
|
| 263 |
|
config.json
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Nemotron3_5AsrForRNNT"
|
| 4 |
+
],
|
| 5 |
+
"blank_token_id": 1024,
|
| 6 |
+
"decoder_hidden_size": 640,
|
| 7 |
+
"default_prompt_id": 38,
|
| 8 |
+
"dtype": "float32",
|
| 9 |
+
"encoder_config": {
|
| 10 |
+
"activation_dropout": 0.1,
|
| 11 |
+
"attention_bias": false,
|
| 12 |
+
"attention_dropout": 0.1,
|
| 13 |
+
"conv_kernel_size": 9,
|
| 14 |
+
"convolution_bias": false,
|
| 15 |
+
"default_num_lookahead_tokens": 3,
|
| 16 |
+
"dropout": 0.1,
|
| 17 |
+
"dropout_positions": 0.0,
|
| 18 |
+
"hidden_act": "silu",
|
| 19 |
+
"hidden_size": 1024,
|
| 20 |
+
"initializer_range": 0.02,
|
| 21 |
+
"intermediate_size": 4096,
|
| 22 |
+
"layerdrop": 0.1,
|
| 23 |
+
"max_position_embeddings": 5000,
|
| 24 |
+
"model_type": "nemotron_asr_streaming_encoder",
|
| 25 |
+
"num_attention_heads": 8,
|
| 26 |
+
"num_hidden_layers": 24,
|
| 27 |
+
"num_key_value_heads": 8,
|
| 28 |
+
"num_mel_bins": 128,
|
| 29 |
+
"scale_input": false,
|
| 30 |
+
"sliding_window": 57,
|
| 31 |
+
"subsampling_conv_channels": 256,
|
| 32 |
+
"subsampling_conv_kernel_size": 3,
|
| 33 |
+
"subsampling_conv_stride": 2,
|
| 34 |
+
"subsampling_factor": 8,
|
| 35 |
+
"supported_num_lookahead_tokens": [
|
| 36 |
+
3,
|
| 37 |
+
0,
|
| 38 |
+
6,
|
| 39 |
+
13
|
| 40 |
+
]
|
| 41 |
+
},
|
| 42 |
+
"hidden_act": "relu",
|
| 43 |
+
"is_encoder_decoder": true,
|
| 44 |
+
"max_symbols_per_step": 10,
|
| 45 |
+
"model_type": "nemotron3_5_asr",
|
| 46 |
+
"num_decoder_layers": 2,
|
| 47 |
+
"num_prompts": 128,
|
| 48 |
+
"pad_token_id": 0,
|
| 49 |
+
"prompt_intermediate_size": 2048,
|
| 50 |
+
"transformers_version": "5.18.0",
|
| 51 |
+
"vocab_size": 1025
|
| 52 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_from_model_config": true,
|
| 3 |
+
"decoder_start_token_id": 1024,
|
| 4 |
+
"output_attentions": false,
|
| 5 |
+
"output_hidden_states": false,
|
| 6 |
+
"pad_token_id": 0,
|
| 7 |
+
"transformers_version": "5.18.0"
|
| 8 |
+
}
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:2eee08fb901cc2e7b2ba096b79173d4e872278b7169447cb37e24ed66df953f5
|
| 3 |
+
size 2490252108
|
processor_config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"blank_token": "<blank>",
|
| 3 |
+
"default_num_lookahead_tokens": 3,
|
| 4 |
+
"feature_extractor": {
|
| 5 |
+
"feature_extractor_type": "NemotronAsrStreamingFeatureExtractor",
|
| 6 |
+
"feature_size": 128,
|
| 7 |
+
"hop_length": 160,
|
| 8 |
+
"n_fft": 512,
|
| 9 |
+
"padding_side": "right",
|
| 10 |
+
"padding_value": 0.0,
|
| 11 |
+
"preemphasis": 0.97,
|
| 12 |
+
"return_attention_mask": true,
|
| 13 |
+
"sampling_rate": 16000,
|
| 14 |
+
"win_length": 400
|
| 15 |
+
},
|
| 16 |
+
"num_prompts": 128,
|
| 17 |
+
"processor_class": "Nemotron3_5AsrProcessor",
|
| 18 |
+
"prompt_dictionary": {
|
| 19 |
+
"auto": 38,
|
| 20 |
+
"fa": 38,
|
| 21 |
+
"fa-IR": 38
|
| 22 |
+
},
|
| 23 |
+
"supported_num_lookahead_tokens": [
|
| 24 |
+
3,
|
| 25 |
+
0,
|
| 26 |
+
6,
|
| 27 |
+
13
|
| 28 |
+
]
|
| 29 |
+
}
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"backend": "tokenizers",
|
| 3 |
+
"clean_up_tokenization_spaces": false,
|
| 4 |
+
"model_max_length": 1000000000000000019884624838656,
|
| 5 |
+
"pad_token": "<unk>",
|
| 6 |
+
"processor_class": "Nemotron3_5AsrProcessor",
|
| 7 |
+
"tokenizer_class": "ParakeetTokenizer",
|
| 8 |
+
"unk_token": "<unk>"
|
| 9 |
+
}
|