lunks commited on
Commit
f34b38d
·
verified ·
1 Parent(s): d0e38c6

Model card in the style of the MLX conversion's

Browse files
Files changed (1) hide show
  1. README.md +66 -134
README.md CHANGED
@@ -1,157 +1,89 @@
1
  ---
2
  license: other
3
  license_name: netease-model-use-license-agreement
4
- license_link: https://raw.githubusercontent.com/netease-youdao/Confucius4-R2T2/refs/heads/master/MODEL_LICENSE
5
  base_model: netease-youdao/Confucius4-R2T2
6
  pipeline_tag: automatic-speech-recognition
7
  library_name: coreml
8
  tags:
9
  - coreml
10
  - apple-neural-engine
11
- - apple-silicon
12
  - asr
13
- - speech-recognition
14
  - streaming
15
- - real-time
16
- - qwen3-asr
17
  - confucius4
18
  - r2t2
19
- language:
20
- - en
21
- - pt
22
- - zh
23
- - es
24
- - fr
25
- - de
26
- - it
27
- - ja
28
- - ko
29
- - ru
30
  ---
31
 
32
- # Confucius4-R2T2 for Core ML and the Apple Neural Engine
33
 
34
- > Any modifications made to the original model in this Derivative Work are not endorsed, warranted,
35
- > or guaranteed by the original right-holder of the original model, and the original right-holder
36
- > disclaims all liability related to this Derivative Work.
37
-
38
- This is [NetEase Youdao's Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2)
39
  (a real-time speech-recognition fine-tune of Qwen3-ASR-1.7B, 30 languages) converted to Core ML
40
- programs that run almost entirely on the Apple Neural Engine: every op of the encoder, 99 % of the
41
- decoder's, and the head's, with the single-row argmax and a dozen index ops per decoder function on
42
- the CPU. No re-training or fine-tuning was done; the weights are the original ones, converted and
43
- palettised.
 
 
 
 
 
 
44
 
45
- It is the model used by the R2T2 provider in a fork of [VoiceInk](https://github.com/Beingpax/VoiceInk),
46
- whose runtime and conversion pipeline are published with that fork. The files here are not a
47
- general-purpose Core ML model with a single input and output: the decoder is a stateful,
48
- multifunction program that the runtime drives one pass at a time (see *How it is run*).
49
 
50
  ## Contents
51
 
52
- | file | what | size |
53
  |---|---|---|
54
- | `R2T2AudioEncoder.mlmodelc` | audio encoder, fp16, one fixed 800-frame (8 s) mel window with a key mask → 104 rows | 607 MB |
55
- | `r2t2_FFN_PF_lut8_chunk_01of02.mlmodelc` | decoder layers 0–13, LUT8, functions `prefill` (128 rows) and `infer` (1 row), stateful KV cache | 692 MB |
56
- | `r2t2_FFN_PF_lut8_chunk_02of02.mlmodelc` | decoder layers 14–27 + final norm, same functions | 692 MB |
57
- | `r2t2_lm_head_lut8.mlmodelc` | language-model head, LUT8, 16-way split: `infer` (one row → logits) and `verify` (128 rows → 128 argmaxes) | 306 MB |
58
- | `embed_tokens.f16.bin` | token embedding table, fp16, 151 936 × 2048, memory-mapped by the runtime | 622 MB |
59
  | `tokenizer.json`, `tokenizer_config.json` | the original tokenizer | 11 MB |
60
- | `MODEL_LICENSE`, `LICENSE-Qwen3-ASR.txt`, `NOTICE`, `SHA256SUMS` | licences, attribution, checksums | |
61
-
62
- Requirements: Apple silicon and macOS 15 or later (the decoder uses Core ML stateful models).
63
- While loaded, the Neural Engine holds the weights dequantised to fp16, about 4.5 GB, outside the
64
- process; the process itself uses about 200 MB. The first load by a given application compiles the
65
- programs for the Neural Engine, which takes about 50 s; the compiled programs are cached per
66
- application binary, so later loads take about a second, and a rebuilt or updated application pays
67
- the compile once more. The cache also needs free disk: with a nearly full disk (under ~20 GB) every
68
- load recompiled.
69
-
70
- ## Interface
71
-
72
- The files are meant for a runtime that drives them; they are not a drop-in `MLModel` with audio in
73
- and text out.
74
-
75
- - **Encoder** `R2T2AudioEncoder.mlmodelc`: input `[1, 128, 800]` log-mel frames (Whisper's recipe:
76
- 16 kHz, n_fft 400, hop 160, Slaney filterbank, `log10`, clamped to 8 dB below the buffer's maximum,
77
- `(x + 4) / 4`) plus a key mask for windows shorter than 800 frames (masked keys get −10 000);
78
- output `[1, 104, 2048]` rows, 13 per 100 frames.
79
- - **Decoder chunks**: functions `prefill` (B = 128) and `infer` (B = 1) with inputs
80
- `hidden_states [1, B, 2048]` fp16, `position_ids [B]` int32, `causal_mask [1, 1, B, 1024]` fp16
81
- (0 to attend, −10 000 otherwise), `current_pos [1]` int32; output `output_hidden_states [1, B, 2048]`
82
- (chunk 2's is final-normalised). Both chunks share one `MLState` of shape `(56, 8, 1024, 128)` fp16:
83
- layer *l* of chunk *c* uses slots *l* (K) and *28 + l* (V). Rows are written by absolute position,
84
- so a caller can rewind and overwrite. Context length 1024.
85
- - **Head**: `infer` takes `hidden_states [1, 1, 2048]` and returns `logits1…logits16`, each
86
- `[1, 1, 9496]` (the vocabulary of 151 936 in 16 slices); `verify` takes `[1, 128, 2048]` and returns
87
- per row and per slice `argmax_val`, `argmax_hi`, `argmax_lo` (`[128, 16]` fp16 each), the slice's
88
- best logit and its index as `hi × 64 + lo`, exact in fp16.
89
- - **Embeddings** `embed_tokens.f16.bin`: 151 936 rows of 2048 fp16 values, row-major, no header.
90
- - The decoder programs load under `.cpuAndNeuralEngine` only: `.all` and `.cpuAndGPU` fail to
91
- compile, and `.cpuOnly` cannot load the multifunction packages. Python `coremltools` cannot open
92
- the combined packages either; drive them from Swift or Objective-C.
93
-
94
- ## Accuracy
95
-
96
- Word error rate of this conversion, decoding whole utterances, on 100-utterance subsets:
97
-
98
- | corpus | this conversion |
99
- |---|---|
100
- | LibriSpeech test-clean (English) | 2.40 % |
101
- | FLEURS pt_br (Brazilian Portuguese) | 3.50 % |
102
-
103
- Against the unconverted BF16 weights run on MLX, paired on 300 utterances of each corpus, the
104
- conversion measured 2.47 % vs 2.42 % (ratio 1.02, 95 % CI 0.96–1.10) on LibriSpeech and
105
- 4.01 % vs 4.01 % (ratio 1.00, CI 0.95–1.05) on FLEURS. The streaming runtime described below ends
106
- on the same transcript as the whole-file decode.
107
-
108
- ## How it is run
109
-
110
- The runtime that uses these files decodes the whole current piece of audio (up to 30 s) on every
111
- pass, so the live text converges on the same result as an offline decode. What keeps that cheap
112
- is what stays in the decoder's KV state between passes: the prompt prefix and the rows of every
113
- completed encoder window. A pass encodes the partial last window, prefills the new audio rows,
114
- the prompt scaffold and the previous pass's tokens as a draft, checks the draft with one call of
115
- the head's `verify` function, and decodes one token at a time only from the first divergence.
116
- Passes take about 150–250 ms on an M5 Pro; a 30 s piece decodes offline in about 2.5 s
117
- (20 ms per token).
118
-
119
- The prompt is the Qwen3-ASR chat template: system text (hotwords may go here), the user turn with
120
- `<|audio_start|>`, the encoder's output rows in place of the `<|audio_pad|>` token embeddings,
121
- `<|audio_end|>`, and the assistant turn, optionally prefixed with `language <Name><asr_text>` to
122
- force a language. The model answers `language <Name><asr_text>` followed by the text; `|` marks
123
- where it stops trusting its own output.
124
-
125
- ## Conversion
126
-
127
- Encoder: fp16, the network rewritten in the layout the Neural Engine compiler keeps resident
128
- (the fused `gelu` op replaced by a tanh formulation, because its constant absolute error is large
129
- against this encoder's small activations), fixed 800-frame window with a key mask for shorter
130
- input; 4786/4786 ops on the Neural Engine, 55 dB SNR against the PyTorch encoder.
131
-
132
- Decoder: through [ANEMLL](https://github.com/Anemll/Anemll) (patched to return the normalised
133
- hidden states of every prefill row), two chunks, context 1024, 8-bit palettised weights, a
134
- 128-row `prefill` and a 1-row `infer` function sharing one KV state; prefill and single-step paths
135
- agree to cosine 1.0. A 6-bit build was measured and rejected (WER ratio 1.16 on Portuguese).
136
-
137
- Head: the 16-way split head of the original conversion plus a `verify` function that returns the
138
- argmax of 128 hidden rows in one call, computed exactly on the Neural Engine.
139
-
140
- ## Licence
141
-
142
- Dual licensing, as for the original model:
143
-
144
- - **Weights** (everything in this repository derived from the model): the
145
- [NetEase Youdao Model Use License Agreement](MODEL_LICENSE). Using these files means accepting
146
- it. In short: royalty-free use including commercial use, with a separate licence required above
147
- 100 million monthly active users or RMB 1 billion in annual revenue; no use to improve other AI
148
- models; no high-risk deployments; keep the NOTICE and MODEL_LICENSE with every copy; anyone you
149
- redistribute to is bound by the same terms. The Chinese version of the agreement prevails.
150
- - **Base model** Qwen3-ASR-1.7B: Apache License 2.0 ([LICENSE-Qwen3-ASR.txt](LICENSE-Qwen3-ASR.txt)).
151
- - **Conversion pipeline and runtime code**: part of the VoiceInk fork, which is GPL-3.0 like
152
- VoiceInk itself. Nothing in this repository is code.
153
-
154
- ## Attribution
155
-
156
- Confucius4-R2T2 by NetEase Youdao. Qwen3-ASR by Alibaba Cloud. Core ML conversion tooling built on
157
- ANEMLL and coremltools. See [NOTICE](NOTICE).
 
1
  ---
2
  license: other
3
  license_name: netease-model-use-license-agreement
4
+ license_link: https://github.com/netease-youdao/Confucius4-R2T2/blob/master/MODEL_LICENSE
5
  base_model: netease-youdao/Confucius4-R2T2
6
  pipeline_tag: automatic-speech-recognition
7
  library_name: coreml
8
  tags:
9
  - coreml
10
  - apple-neural-engine
11
+ - speech-to-text
12
  - asr
13
+ - stt
14
  - streaming
 
 
15
  - confucius4
16
  - r2t2
 
 
 
 
 
 
 
 
 
 
 
17
  ---
18
 
19
+ # Confucius4-R2T2, Core ML for the Apple Neural Engine
20
 
21
+ [`netease-youdao/Confucius4-R2T2`](https://huggingface.co/netease-youdao/Confucius4-R2T2)
 
 
 
 
22
  (a real-time speech-recognition fine-tune of Qwen3-ASR-1.7B, 30 languages) converted to Core ML
23
+ programs that run on the Apple Neural Engine: the audio encoder in fp16, the Qwen3 decoder and the
24
+ LM head palettised to 8 bits. 4.09 GB of bfloat16 becomes 2.8 GB on disk; while loaded, the
25
+ Neural Engine holds the weights dequantised to fp16, about 4.5 GB, outside the process.
26
+
27
+ Converted with the pipeline in the VoiceInk fork's `tools/r2t2-coreml`: the encoder through
28
+ coremltools 9 with a fixed 800-frame window (the fused `gelu` replaced by a tanh formulation, whose
29
+ constant absolute error otherwise swamps this encoder's small activations), the decoder through
30
+ [ANEMLL](https://github.com/Anemll/Anemll) as two stateful chunks with a 128-row `prefill` and a
31
+ 1-row `infer` function, and a head with a `verify` function that returns the argmax of 128 rows
32
+ in one call. Requires macOS 15 or later (Core ML stateful models) on Apple silicon.
33
 
34
+ Word error rate, whole utterances, 100-utterance subsets: LibriSpeech test-clean 2.40 %, FLEURS
35
+ pt_br 3.50 %. Against the unconverted bfloat16 weights on MLX, paired on 300 utterances of each:
36
+ 2.47 % vs 2.42 % and 4.01 % vs 4.01 %.
 
37
 
38
  ## Contents
39
 
40
+ | file | | size |
41
  |---|---|---|
42
+ | `R2T2AudioEncoder.mlmodelc` | encoder, fp16, `[1, 128, 800]` mel window + key mask → `[1, 104, 2048]` | 607 MB |
43
+ | `r2t2_FFN_PF_lut8_chunk_01of02.mlmodelc` | decoder layers 0–13, functions `prefill` (128 rows) and `infer` (1 row) | 692 MB |
44
+ | `r2t2_FFN_PF_lut8_chunk_02of02.mlmodelc` | decoder layers 14–27 and the final norm, same functions | 692 MB |
45
+ | `r2t2_lm_head_lut8.mlmodelc` | head, 16-way split: `infer` (one row → `logits1…16`), `verify` (128 rows → argmaxes) | 306 MB |
46
+ | `embed_tokens.f16.bin` | token embeddings, 151 936 × 2048 fp16, row-major, no header | 622 MB |
47
  | `tokenizer.json`, `tokenizer_config.json` | the original tokenizer | 11 MB |
48
+ | `MODEL_LICENSE`, `LICENSE-Qwen3-ASR.txt`, `NOTICE`, `SHA256SUMS` | | |
49
+
50
+ ## Use
51
+
52
+ These files are driven by the R2T2 provider of a VoiceInk fork, which downloads them from here
53
+ after the licence is accepted. The runtime re-decodes the whole current piece of audio (up to 30 s)
54
+ on every pass, so the live text converges on the same result as an offline decode; what makes that
55
+ cheap is keeping the prompt prefix and every completed encoder window in the decoder's KV state,
56
+ prefilling only the new rows plus the previous pass's tokens as a draft, confirming the draft with
57
+ one `verify` call and decoding only from the first divergence. Passes take 150–250 ms on an M5 Pro;
58
+ a 30 s piece decodes offline in about 2.5 s.
59
+
60
+ For another runtime, the interface:
61
+
62
+ - Decoder inputs `hidden_states [1, B, 2048]` fp16, `position_ids [B]` int32,
63
+ `causal_mask [1, 1, B, 1024]` fp16 (0 to attend, −10 000 otherwise), `current_pos [1]` int32;
64
+ output `output_hidden_states [1, B, 2048]`. Both chunks share one `MLState` of shape
65
+ `(56, 8, 1024, 128)` fp16, written by absolute position, so a caller can rewind and overwrite.
66
+ - Head `verify` output per row and slice: `argmax_val`, `argmax_hi`, `argmax_lo` (`[128, 16]` fp16),
67
+ the slice's best logit and its index as `hi × 64 + lo`.
68
+ - Prompt: the Qwen3-ASR chat template, with the encoder rows in place of the `<|audio_pad|>`
69
+ embeddings and, to force a language, the assistant turn opened with `language <Name><asr_text>`.
70
+ The model's `|` marks where it stops trusting its own output.
71
+ - Mel: Whisper's recipe (16 kHz, n_fft 400, hop 160, 128 Slaney bins, `log10`, clamped to 8 dB below
72
+ the buffer's maximum, `(x + 4) / 4`).
73
+ - The decoder loads under `.cpuAndNeuralEngine` only, and Python coremltools cannot open the
74
+ multifunction packages. The first load by an application compiles the programs (about 50 s);
75
+ they are cached per application binary afterwards.
76
+
77
+ ## Notice
78
+
79
+ Any modifications made to the original model in this Derivative Work are not endorsed, warranted,
80
+ or guaranteed by the original right-holder of the original model, and the original right-holder
81
+ disclaims all liability related to this Derivative Work.
82
+
83
+ This is a conversion of NetEase Youdao's Confucius4-R2T2 and is governed by the original
84
+ [MODEL_LICENSE](MODEL_LICENSE). That licence is royalty-free for most users, including commercial
85
+ use, but requires a separate licence from NetEase Youdao above 100 million monthly active users or
86
+ RMB 1 billion in annual revenue, forbids using the model to improve other AI models and use in the
87
+ high-risk scenarios it lists, and binds anyone you redistribute it to. Using these files means
88
+ accepting it; keep `MODEL_LICENSE` and `NOTICE` with every copy. The base model, Qwen3-ASR-1.7B, is
89
+ Apache 2.0 (`LICENSE-Qwen3-ASR.txt`).