cstr commited on
Commit
3f418f1
·
verified ·
1 Parent(s): 07fbd00

Add model card with architecture, languages, and usage

Browse files
Files changed (1) hide show
  1. README.md +152 -0
README.md ADDED
@@ -0,0 +1,152 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: openmdw-1.1
3
+ language:
4
+ - ar
5
+ - bg
6
+ - cs
7
+ - da
8
+ - de
9
+ - el
10
+ - en
11
+ - es
12
+ - et
13
+ - fi
14
+ - fr
15
+ - he
16
+ - hi
17
+ - hr
18
+ - hu
19
+ - it
20
+ - ja
21
+ - ko
22
+ - lt
23
+ - lv
24
+ - nb
25
+ - nl
26
+ - nn
27
+ - pl
28
+ - pt
29
+ - ro
30
+ - ru
31
+ - sk
32
+ - sl
33
+ - sv
34
+ - th
35
+ - tr
36
+ - uk
37
+ - vi
38
+ - zh
39
+ tags:
40
+ - asr
41
+ - speech-recognition
42
+ - gguf
43
+ - streaming
44
+ - fastconformer
45
+ - rnnt
46
+ - multilingual
47
+ - crispasr
48
+ pipeline_tag: automatic-speech-recognition
49
+ base_model: nvidia/nemotron-3.5-asr-streaming-0.6b
50
+ ---
51
+
52
+ # Nemotron-3.5-ASR-Streaming-0.6B GGUF
53
+
54
+ GGUF conversion of [nvidia/nemotron-3.5-asr-streaming-0.6b](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b) for use with [CrispASR](https://github.com/CrispStrobe/CrispASR).
55
+
56
+ ## Model details
57
+
58
+ **Architecture:** Cache-Aware Streaming FastConformer encoder (24 layers, d=1024, 8 heads) + RNN-T decoder (2-layer LSTM, hidden=640) + joint network (640 → 13088 vocab).
59
+
60
+ **Languages:** 39 languages, selected via `prompt_kernel` MLP conditioning:
61
+
62
+ | Code | Language | Code | Language | Code | Language |
63
+ |------|----------|------|----------|------|----------|
64
+ | ar-AR | Arabic | fr-CA | French (CA) | nn-NO | Norwegian (NN) |
65
+ | bg-BG | Bulgarian | fr-FR | French (FR) | pl-PL | Polish |
66
+ | cs-CZ | Czech | he-IL | Hebrew | pt-BR | Portuguese (BR) |
67
+ | da-DK | Danish | hi-IN | Hindi | pt-PT | Portuguese (PT) |
68
+ | de-DE | German | hr-HR | Croatian | ro-RO | Romanian |
69
+ | el-GR | Greek | hu-HU | Hungarian | ru-RU | Russian |
70
+ | en-GB | English (GB) | it-IT | Italian | sk-SK | Slovak |
71
+ | en-US | English (US) | ja-JP | Japanese | sl-SL | Slovenian |
72
+ | es-ES | Spanish (ES) | ko-KR | Korean | sv-SE | Swedish |
73
+ | es-US | Spanish (US) | lt-LT | Lithuanian | th-TH | Thai |
74
+ | et-EE | Estonian | lv-LV | Latvian | tr-TR | Turkish |
75
+ | fi-FI | Finnish | nb-NO | Norwegian (NB) | uk-UA | Ukrainian |
76
+ | | | nl-NL | Dutch | vi-VN | Vietnamese |
77
+ | | | | | zh-CN | Chinese |
78
+
79
+ **Key properties:**
80
+ - Sample rate: 16 kHz
81
+ - 128 mel filterbank features, n_fft=512, hop=160 (10ms), win=400 (25ms)
82
+ - 8× time downsampling (causal) → 80ms frame duration
83
+ - Streaming: cache-aware attention with `att_context_size=[[56,3],[56,0],[56,6],[56,13]]`
84
+ - Vocab: 13087 SentencePiece tokens + 1 blank (pure RNN-T, no TDT durations)
85
+ - License: [OpenMDW-1.1](https://github.com/linux-foundation/open-model-developer-weight-license/blob/main/LICENSE.md) (permissive)
86
+
87
+ ## Files
88
+
89
+ | File | Size | Description |
90
+ |------|------|-------------|
91
+ | `nemotron-3.5-asr-streaming-0.6b-f16.gguf` | ~1.2 GB | F16 weights (full precision) |
92
+ | `nemotron-3.5-asr-streaming-ref.gguf` | ~1.2 GB | Reference GGUF (for parity testing) |
93
+
94
+ ## Usage with CrispASR
95
+
96
+ ```bash
97
+ # Download
98
+ huggingface-cli download cstr/nemotron-3.5-asr-streaming-GGUF \
99
+ nemotron-3.5-asr-streaming-0.6b-f16.gguf --local-dir models/
100
+
101
+ # Transcribe (English, default)
102
+ crispasr --backend nemotron \
103
+ -m models/nemotron-3.5-asr-streaming-0.6b-f16.gguf \
104
+ -f audio.wav
105
+
106
+ # Transcribe in German
107
+ crispasr --backend nemotron \
108
+ -m models/nemotron-3.5-asr-streaming-0.6b-f16.gguf \
109
+ -f audio.wav --language de-DE
110
+
111
+ # Streaming mode
112
+ crispasr --backend nemotron \
113
+ -m models/nemotron-3.5-asr-streaming-0.6b-f16.gguf \
114
+ -f audio.wav --stream
115
+ ```
116
+
117
+ ## Conversion
118
+
119
+ Converted from the original NeMo `.nemo` checkpoint using:
120
+
121
+ ```bash
122
+ python models/convert-nemotron-to-gguf.py \
123
+ --nemo nvidia/nemotron-3.5-asr-streaming-0.6b \
124
+ --output nemotron-3.5-asr-streaming-0.6b-f16.gguf
125
+ ```
126
+
127
+ Quantized variants can be produced with:
128
+
129
+ ```bash
130
+ crispasr-quantize models/nemotron-3.5-asr-streaming-0.6b-f16.gguf \
131
+ models/nemotron-3.5-asr-streaming-0.6b-q4_k.gguf Q4_K
132
+ ```
133
+
134
+ ## Architecture
135
+
136
+ ```
137
+ Audio (16kHz) → Mel (128 bins, 10ms hop)
138
+ → Pre-encode (3× causal Conv2d, 8× downsample, Linear 4352→1024)
139
+ → 24× Cache-Aware FastConformer block:
140
+ FFN1(½) → MHA(rel_pos, cache-aware) → DWConv(k=9, causal, LN) → FFN2(½) → LN
141
+ → Prompt kernel (MLP: concat(enc[1024], lang_onehot[128]) → 2048 → 1024)
142
+ → RNN-T decoder:
143
+ Prediction: Embed(13088, 640) + 2-layer LSTM(640)
144
+ Joint: enc(1024→640) + pred(640→640) → ReLU → Linear(640→13088)
145
+ → Greedy / beam search decode
146
+ ```
147
+
148
+ ## Original model
149
+
150
+ - **Paper:** [NVIDIA NeMo documentation](https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/asr/intro.html)
151
+ - **Source:** [nvidia/nemotron-3.5-asr-streaming-0.6b](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b)
152
+ - **License:** OpenMDW-1.1