mehdi-hf commited on
Commit
5201feb
·
verified ·
1 Parent(s): 1b768f5

Add 🤗 Transformers weights/processor (converted + verified against the .nemo); card: Transformers usage

Browse files
README.md CHANGED
@@ -19,6 +19,7 @@ tags:
19
  - FastConformer
20
  - RNNT
21
  - NeMo
 
22
  - persian
23
  - farsi
24
  model-index:
@@ -85,18 +86,86 @@ Word error rate (WER) and character error rate (CER), in %. Lower is better. All
85
 
86
  ## How to use
87
 
88
- Install [NeMo](https://github.com/NVIDIA-NeMo/Speech) 3.0 (`pip install "nemo_toolkit[asr]>=3.0"`; Python ≥ 3.11). Then download `nemotron-asr-streaming-farsi.nemo` from this repo:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
89
 
90
  ```python
91
  from huggingface_hub import hf_hub_download
92
  model_path = hf_hub_download("mehdi-hf/nemotron-asr-streaming-farsi", "nemotron-asr-streaming-farsi.nemo")
93
  ```
94
 
95
- > **Set the `fa-IR` prompt.** This model was fine-tuned only with the `fa-IR` language prompt.
96
  > - In streaming, call `model.set_inference_prompt("fa-IR")`.
97
  > - Don't use `model.transcribe()` or NeMo's `speech_to_text_eval.py` as they are. In NeMo 3.0 their dataloader picks the prompt **at random per utterance** (the untrained `auto` prompt about half the time), which gives much worse, non-reproducible output.
 
 
 
98
 
99
- ### Streaming (Python)
100
 
101
  ```python
102
  import librosa, soundfile as sf, torch
@@ -133,7 +202,7 @@ with torch.inference_mode():
133
 
134
  The loop prints the transcript as it grows, chunk by chunk. On an Apple M1 Pro (`mps`) it runs about 5–10× faster than real time.
135
 
136
- ### Streaming (command line, NeMo script)
137
 
138
  ```bash
139
  git clone --depth 1 --branch v3.0.0 https://github.com/NVIDIA-NeMo/Speech.git nemo-speech
@@ -188,6 +257,7 @@ The raw data was cut and filtered to 1,181 h:
188
 
189
  - **Conversational speech is still hard:** ~26% WER on spontaneous film and YouTube speech with music and noise, against ~9% on read speech.
190
  - **Spacing variants still cost a little WER:** compound words written with or without a space (`چندتا` / `چند تا`) count as errors.
 
191
  - **Not evaluated on dialects:** the training data is Iranian media and podcast speech. Performance on regional accents, Dari or Tajik hasn't been measured.
192
  - **Evaluation coverage:** FLEURS Persian test speakers are all male.
193
 
 
19
  - FastConformer
20
  - RNNT
21
  - NeMo
22
+ - transformers
23
  - persian
24
  - farsi
25
  model-index:
 
86
 
87
  ## How to use
88
 
89
+ The model works with **🤗 Transformers** (≥ 5.18, files in this repo's root) and with **NVIDIA NeMo** (`nemotron-asr-streaming-farsi.nemo`). Both give the same results. The Transformers weights are bit-identical to the `.nemo`, and on FLEURS (852 clips) it scores **8.81% WER at 1.12 s look-ahead vs 8.77%** for NeMo, and 9.02% vs 9.13% at 0.32 s. In streaming, 98 of 100 transcripts were identical.
90
+
91
+ ### 🤗 Transformers: transcribe a file
92
+
93
+ ```python
94
+ from transformers import AutoModelForRNNT, AutoProcessor
95
+ from transformers.audio_utils import load_audio
96
+
97
+ model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
98
+ processor = AutoProcessor.from_pretrained(model_id)
99
+ model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
100
+
101
+ processor.set_num_lookahead_tokens(13) # 1.12 s look-ahead (most accurate); 3 = 0.32 s (default), 6, 0
102
+ audio = load_audio("audio.mp3", sampling_rate=processor.feature_extractor.sampling_rate)
103
+ inputs = processor(audio, sampling_rate=processor.feature_extractor.sampling_rate, language="fa-IR")
104
+ inputs = inputs.to(model.device, dtype=model.dtype)
105
+ output = model.generate(**inputs, return_dict_in_generate=True)
106
+ print(processor.decode(output.sequences, skip_special_tokens=True)[0])
107
+ ```
108
+
109
+ `language` accepts only `"fa-IR"`, `"fa"` or `"auto"`, and all three select the Persian prompt the model was trained with. The other languages of the base model aren't supported.
110
+
111
+ **Fine-tuning:** use NeMo. Transformers 5.18 can't compute this model's training loss. It has no loss entry for `Nemotron3_5AsrForRNNT` and falls back to a causal-LM loss; NVIDIA's original model has the same limitation. The processor's `labels`/`decoder_input_ids` are correct for this model, so this will work once Transformers adds the loss.
112
+
113
+ ### 🤗 Transformers: live streaming
114
+
115
+ ```python
116
+ from threading import Thread
117
+ import numpy as np
118
+ from transformers import AutoModelForRNNT, AutoProcessor, TextIteratorStreamer
119
+ from transformers.audio_utils import load_audio
120
+
121
+ model_id = "mehdi-hf/nemotron-asr-streaming-farsi"
122
+ processor = AutoProcessor.from_pretrained(model_id)
123
+ model = AutoModelForRNNT.from_pretrained(model_id, device_map="auto")
124
+ processor.set_num_lookahead_tokens(13)
125
+
126
+ sr = processor.feature_extractor.sampling_rate
127
+ audio = load_audio("audio.mp3", sampling_rate=sr)
128
+ # Chunks are processed whole and the model holds back its look-ahead frames, so append two chunks of
129
+ # silence; otherwise the last ~1 s of speech is never transcribed. (A live app does this when it stops.)
130
+ audio = np.concatenate([audio, np.zeros(2 * processor.num_samples_per_audio_chunk, dtype=np.float32)])
131
+ first = processor(audio[: processor.num_samples_first_audio_chunk], sampling_rate=sr, is_streaming=True,
132
+ is_first_audio_chunk=True, language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
133
+
134
+ def chunks(): # in a live app, feed microphone audio here as it arrives
135
+ yield first.input_features[:, : processor.num_mel_frames_first_audio_chunk, :]
136
+ mel, hop, n_fft = processor.num_mel_frames_first_audio_chunk, processor.feature_extractor.hop_length, processor.feature_extractor.n_fft
137
+ start = mel * hop - n_fft // 2
138
+ while (end := start + processor.num_samples_per_audio_chunk) < audio.shape[0]:
139
+ x = processor(audio[start:end], sampling_rate=sr, is_streaming=True, is_first_audio_chunk=False,
140
+ language="fa-IR", return_tensors="pt").to(model.device, dtype=model.dtype)
141
+ yield x.input_features
142
+ mel += processor.num_mel_frames_per_audio_chunk
143
+ start = mel * hop - n_fft // 2
144
+
145
+ # group_tokens=False is required: the tokenizer's decode merges repeated tokens by default (a CTC rule),
146
+ # which would drop letters from RNN-T output. processor.decode sets it for you; the streamer doesn't.
147
+ streamer = TextIteratorStreamer(processor.tokenizer, skip_special_tokens=True, group_tokens=False)
148
+ Thread(target=model.generate, kwargs={**first, "input_features": chunks(), "streamer": streamer}).start()
149
+ for text in streamer:
150
+ print(text, end="", flush=True)
151
+ ```
152
+
153
+ ### NeMo
154
+
155
+ Install [NeMo](https://github.com/NVIDIA-NeMo/Speech) 3.0 (`pip install "nemo_toolkit[asr]>=3.0"`; Python ≥ 3.11). Then download the `.nemo` file:
156
 
157
  ```python
158
  from huggingface_hub import hf_hub_download
159
  model_path = hf_hub_download("mehdi-hf/nemotron-asr-streaming-farsi", "nemotron-asr-streaming-farsi.nemo")
160
  ```
161
 
162
+ > **Set the `fa-IR` prompt in NeMo.** This model was fine-tuned only with the `fa-IR` language prompt.
163
  > - In streaming, call `model.set_inference_prompt("fa-IR")`.
164
  > - Don't use `model.transcribe()` or NeMo's `speech_to_text_eval.py` as they are. In NeMo 3.0 their dataloader picks the prompt **at random per utterance** (the untrained `auto` prompt about half the time), which gives much worse, non-reproducible output.
165
+ > - The Transformers version doesn't have this problem: it always uses the Persian prompt.
166
+
167
+ #### NeMo streaming (Python)
168
 
 
169
 
170
  ```python
171
  import librosa, soundfile as sf, torch
 
202
 
203
  The loop prints the transcript as it grows, chunk by chunk. On an Apple M1 Pro (`mps`) it runs about 5–10× faster than real time.
204
 
205
+ #### NeMo streaming (command line)
206
 
207
  ```bash
208
  git clone --depth 1 --branch v3.0.0 https://github.com/NVIDIA-NeMo/Speech.git nemo-speech
 
257
 
258
  - **Conversational speech is still hard:** ~26% WER on spontaneous film and YouTube speech with music and noise, against ~9% on read speech.
259
  - **Spacing variants still cost a little WER:** compound words written with or without a space (`چندتا` / `چند تا`) count as errors.
260
+ - **End of a stream (Transformers):** Transformers 5.18's streaming accepts only full-size chunks, so the end of the audio is padded with silence. A word cut off by the very end of a recording can then be dropped; whole-file mode and NeMo's streaming keep it.
261
  - **Not evaluated on dialects:** the training data is Iranian media and podcast speech. Performance on regional accents, Dari or Tajik hasn't been measured.
262
  - **Evaluation coverage:** FLEURS Persian test speakers are all male.
263
 
config.json ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Nemotron3_5AsrForRNNT"
4
+ ],
5
+ "blank_token_id": 1024,
6
+ "decoder_hidden_size": 640,
7
+ "default_prompt_id": 38,
8
+ "dtype": "float32",
9
+ "encoder_config": {
10
+ "activation_dropout": 0.1,
11
+ "attention_bias": false,
12
+ "attention_dropout": 0.1,
13
+ "conv_kernel_size": 9,
14
+ "convolution_bias": false,
15
+ "default_num_lookahead_tokens": 3,
16
+ "dropout": 0.1,
17
+ "dropout_positions": 0.0,
18
+ "hidden_act": "silu",
19
+ "hidden_size": 1024,
20
+ "initializer_range": 0.02,
21
+ "intermediate_size": 4096,
22
+ "layerdrop": 0.1,
23
+ "max_position_embeddings": 5000,
24
+ "model_type": "nemotron_asr_streaming_encoder",
25
+ "num_attention_heads": 8,
26
+ "num_hidden_layers": 24,
27
+ "num_key_value_heads": 8,
28
+ "num_mel_bins": 128,
29
+ "scale_input": false,
30
+ "sliding_window": 57,
31
+ "subsampling_conv_channels": 256,
32
+ "subsampling_conv_kernel_size": 3,
33
+ "subsampling_conv_stride": 2,
34
+ "subsampling_factor": 8,
35
+ "supported_num_lookahead_tokens": [
36
+ 3,
37
+ 0,
38
+ 6,
39
+ 13
40
+ ]
41
+ },
42
+ "hidden_act": "relu",
43
+ "is_encoder_decoder": true,
44
+ "max_symbols_per_step": 10,
45
+ "model_type": "nemotron3_5_asr",
46
+ "num_decoder_layers": 2,
47
+ "num_prompts": 128,
48
+ "pad_token_id": 0,
49
+ "prompt_intermediate_size": 2048,
50
+ "transformers_version": "5.18.0",
51
+ "vocab_size": 1025
52
+ }
generation_config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "decoder_start_token_id": 1024,
4
+ "output_attentions": false,
5
+ "output_hidden_states": false,
6
+ "pad_token_id": 0,
7
+ "transformers_version": "5.18.0"
8
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2eee08fb901cc2e7b2ba096b79173d4e872278b7169447cb37e24ed66df953f5
3
+ size 2490252108
processor_config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "blank_token": "<blank>",
3
+ "default_num_lookahead_tokens": 3,
4
+ "feature_extractor": {
5
+ "feature_extractor_type": "NemotronAsrStreamingFeatureExtractor",
6
+ "feature_size": 128,
7
+ "hop_length": 160,
8
+ "n_fft": 512,
9
+ "padding_side": "right",
10
+ "padding_value": 0.0,
11
+ "preemphasis": 0.97,
12
+ "return_attention_mask": true,
13
+ "sampling_rate": 16000,
14
+ "win_length": 400
15
+ },
16
+ "num_prompts": 128,
17
+ "processor_class": "Nemotron3_5AsrProcessor",
18
+ "prompt_dictionary": {
19
+ "auto": 38,
20
+ "fa": 38,
21
+ "fa-IR": 38
22
+ },
23
+ "supported_num_lookahead_tokens": [
24
+ 3,
25
+ 0,
26
+ 6,
27
+ 13
28
+ ]
29
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "clean_up_tokenization_spaces": false,
4
+ "model_max_length": 1000000000000000019884624838656,
5
+ "pad_token": "<unk>",
6
+ "processor_class": "Nemotron3_5AsrProcessor",
7
+ "tokenizer_class": "ParakeetTokenizer",
8
+ "unk_token": "<unk>"
9
+ }