Model card in the style of the MLX conversion's
Browse files
README.md
CHANGED
|
@@ -1,157 +1,89 @@
|
|
| 1 |
---
|
| 2 |
license: other
|
| 3 |
license_name: netease-model-use-license-agreement
|
| 4 |
-
license_link: https://
|
| 5 |
base_model: netease-youdao/Confucius4-R2T2
|
| 6 |
pipeline_tag: automatic-speech-recognition
|
| 7 |
library_name: coreml
|
| 8 |
tags:
|
| 9 |
- coreml
|
| 10 |
- apple-neural-engine
|
| 11 |
-
-
|
| 12 |
- asr
|
| 13 |
-
-
|
| 14 |
- streaming
|
| 15 |
-
- real-time
|
| 16 |
-
- qwen3-asr
|
| 17 |
- confucius4
|
| 18 |
- r2t2
|
| 19 |
-
language:
|
| 20 |
-
- en
|
| 21 |
-
- pt
|
| 22 |
-
- zh
|
| 23 |
-
- es
|
| 24 |
-
- fr
|
| 25 |
-
- de
|
| 26 |
-
- it
|
| 27 |
-
- ja
|
| 28 |
-
- ko
|
| 29 |
-
- ru
|
| 30 |
---
|
| 31 |
|
| 32 |
-
# Confucius4-R2T2
|
| 33 |
|
| 34 |
-
|
| 35 |
-
> or guaranteed by the original right-holder of the original model, and the original right-holder
|
| 36 |
-
> disclaims all liability related to this Derivative Work.
|
| 37 |
-
|
| 38 |
-
This is [NetEase Youdao's Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2)
|
| 39 |
(a real-time speech-recognition fine-tune of Qwen3-ASR-1.7B, 30 languages) converted to Core ML
|
| 40 |
-
programs that run
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 44 |
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
multifunction program that the runtime drives one pass at a time (see *How it is run*).
|
| 49 |
|
| 50 |
## Contents
|
| 51 |
|
| 52 |
-
| file |
|
| 53 |
|---|---|---|
|
| 54 |
-
| `R2T2AudioEncoder.mlmodelc` |
|
| 55 |
-
| `r2t2_FFN_PF_lut8_chunk_01of02.mlmodelc` | decoder layers 0–13,
|
| 56 |
-
| `r2t2_FFN_PF_lut8_chunk_02of02.mlmodelc` | decoder layers 14–27
|
| 57 |
-
| `r2t2_lm_head_lut8.mlmodelc` |
|
| 58 |
-
| `embed_tokens.f16.bin` | token
|
| 59 |
| `tokenizer.json`, `tokenizer_config.json` | the original tokenizer | 11 MB |
|
| 60 |
-
| `MODEL_LICENSE`, `LICENSE-Qwen3-ASR.txt`, `NOTICE`, `SHA256SUMS` |
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
`(
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
Against the unconverted BF16 weights run on MLX, paired on 300 utterances of each corpus, the
|
| 104 |
-
conversion measured 2.47 % vs 2.42 % (ratio 1.02, 95 % CI 0.96–1.10) on LibriSpeech and
|
| 105 |
-
4.01 % vs 4.01 % (ratio 1.00, CI 0.95–1.05) on FLEURS. The streaming runtime described below ends
|
| 106 |
-
on the same transcript as the whole-file decode.
|
| 107 |
-
|
| 108 |
-
## How it is run
|
| 109 |
-
|
| 110 |
-
The runtime that uses these files decodes the whole current piece of audio (up to 30 s) on every
|
| 111 |
-
pass, so the live text converges on the same result as an offline decode. What keeps that cheap
|
| 112 |
-
is what stays in the decoder's KV state between passes: the prompt prefix and the rows of every
|
| 113 |
-
completed encoder window. A pass encodes the partial last window, prefills the new audio rows,
|
| 114 |
-
the prompt scaffold and the previous pass's tokens as a draft, checks the draft with one call of
|
| 115 |
-
the head's `verify` function, and decodes one token at a time only from the first divergence.
|
| 116 |
-
Passes take about 150–250 ms on an M5 Pro; a 30 s piece decodes offline in about 2.5 s
|
| 117 |
-
(20 ms per token).
|
| 118 |
-
|
| 119 |
-
The prompt is the Qwen3-ASR chat template: system text (hotwords may go here), the user turn with
|
| 120 |
-
`<|audio_start|>`, the encoder's output rows in place of the `<|audio_pad|>` token embeddings,
|
| 121 |
-
`<|audio_end|>`, and the assistant turn, optionally prefixed with `language <Name><asr_text>` to
|
| 122 |
-
force a language. The model answers `language <Name><asr_text>` followed by the text; `|` marks
|
| 123 |
-
where it stops trusting its own output.
|
| 124 |
-
|
| 125 |
-
## Conversion
|
| 126 |
-
|
| 127 |
-
Encoder: fp16, the network rewritten in the layout the Neural Engine compiler keeps resident
|
| 128 |
-
(the fused `gelu` op replaced by a tanh formulation, because its constant absolute error is large
|
| 129 |
-
against this encoder's small activations), fixed 800-frame window with a key mask for shorter
|
| 130 |
-
input; 4786/4786 ops on the Neural Engine, 55 dB SNR against the PyTorch encoder.
|
| 131 |
-
|
| 132 |
-
Decoder: through [ANEMLL](https://github.com/Anemll/Anemll) (patched to return the normalised
|
| 133 |
-
hidden states of every prefill row), two chunks, context 1024, 8-bit palettised weights, a
|
| 134 |
-
128-row `prefill` and a 1-row `infer` function sharing one KV state; prefill and single-step paths
|
| 135 |
-
agree to cosine 1.0. A 6-bit build was measured and rejected (WER ratio 1.16 on Portuguese).
|
| 136 |
-
|
| 137 |
-
Head: the 16-way split head of the original conversion plus a `verify` function that returns the
|
| 138 |
-
argmax of 128 hidden rows in one call, computed exactly on the Neural Engine.
|
| 139 |
-
|
| 140 |
-
## Licence
|
| 141 |
-
|
| 142 |
-
Dual licensing, as for the original model:
|
| 143 |
-
|
| 144 |
-
- **Weights** (everything in this repository derived from the model): the
|
| 145 |
-
[NetEase Youdao Model Use License Agreement](MODEL_LICENSE). Using these files means accepting
|
| 146 |
-
it. In short: royalty-free use including commercial use, with a separate licence required above
|
| 147 |
-
100 million monthly active users or RMB 1 billion in annual revenue; no use to improve other AI
|
| 148 |
-
models; no high-risk deployments; keep the NOTICE and MODEL_LICENSE with every copy; anyone you
|
| 149 |
-
redistribute to is bound by the same terms. The Chinese version of the agreement prevails.
|
| 150 |
-
- **Base model** Qwen3-ASR-1.7B: Apache License 2.0 ([LICENSE-Qwen3-ASR.txt](LICENSE-Qwen3-ASR.txt)).
|
| 151 |
-
- **Conversion pipeline and runtime code**: part of the VoiceInk fork, which is GPL-3.0 like
|
| 152 |
-
VoiceInk itself. Nothing in this repository is code.
|
| 153 |
-
|
| 154 |
-
## Attribution
|
| 155 |
-
|
| 156 |
-
Confucius4-R2T2 by NetEase Youdao. Qwen3-ASR by Alibaba Cloud. Core ML conversion tooling built on
|
| 157 |
-
ANEMLL and coremltools. See [NOTICE](NOTICE).
|
|
|
|
| 1 |
---
|
| 2 |
license: other
|
| 3 |
license_name: netease-model-use-license-agreement
|
| 4 |
+
license_link: https://github.com/netease-youdao/Confucius4-R2T2/blob/master/MODEL_LICENSE
|
| 5 |
base_model: netease-youdao/Confucius4-R2T2
|
| 6 |
pipeline_tag: automatic-speech-recognition
|
| 7 |
library_name: coreml
|
| 8 |
tags:
|
| 9 |
- coreml
|
| 10 |
- apple-neural-engine
|
| 11 |
+
- speech-to-text
|
| 12 |
- asr
|
| 13 |
+
- stt
|
| 14 |
- streaming
|
|
|
|
|
|
|
| 15 |
- confucius4
|
| 16 |
- r2t2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
---
|
| 18 |
|
| 19 |
+
# Confucius4-R2T2, Core ML for the Apple Neural Engine
|
| 20 |
|
| 21 |
+
[`netease-youdao/Confucius4-R2T2`](https://huggingface.co/netease-youdao/Confucius4-R2T2)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
(a real-time speech-recognition fine-tune of Qwen3-ASR-1.7B, 30 languages) converted to Core ML
|
| 23 |
+
programs that run on the Apple Neural Engine: the audio encoder in fp16, the Qwen3 decoder and the
|
| 24 |
+
LM head palettised to 8 bits. 4.09 GB of bfloat16 becomes 2.8 GB on disk; while loaded, the
|
| 25 |
+
Neural Engine holds the weights dequantised to fp16, about 4.5 GB, outside the process.
|
| 26 |
+
|
| 27 |
+
Converted with the pipeline in the VoiceInk fork's `tools/r2t2-coreml`: the encoder through
|
| 28 |
+
coremltools 9 with a fixed 800-frame window (the fused `gelu` replaced by a tanh formulation, whose
|
| 29 |
+
constant absolute error otherwise swamps this encoder's small activations), the decoder through
|
| 30 |
+
[ANEMLL](https://github.com/Anemll/Anemll) as two stateful chunks with a 128-row `prefill` and a
|
| 31 |
+
1-row `infer` function, and a head with a `verify` function that returns the argmax of 128 rows
|
| 32 |
+
in one call. Requires macOS 15 or later (Core ML stateful models) on Apple silicon.
|
| 33 |
|
| 34 |
+
Word error rate, whole utterances, 100-utterance subsets: LibriSpeech test-clean 2.40 %, FLEURS
|
| 35 |
+
pt_br 3.50 %. Against the unconverted bfloat16 weights on MLX, paired on 300 utterances of each:
|
| 36 |
+
2.47 % vs 2.42 % and 4.01 % vs 4.01 %.
|
|
|
|
| 37 |
|
| 38 |
## Contents
|
| 39 |
|
| 40 |
+
| file | | size |
|
| 41 |
|---|---|---|
|
| 42 |
+
| `R2T2AudioEncoder.mlmodelc` | encoder, fp16, `[1, 128, 800]` mel window + key mask → `[1, 104, 2048]` | 607 MB |
|
| 43 |
+
| `r2t2_FFN_PF_lut8_chunk_01of02.mlmodelc` | decoder layers 0–13, functions `prefill` (128 rows) and `infer` (1 row) | 692 MB |
|
| 44 |
+
| `r2t2_FFN_PF_lut8_chunk_02of02.mlmodelc` | decoder layers 14–27 and the final norm, same functions | 692 MB |
|
| 45 |
+
| `r2t2_lm_head_lut8.mlmodelc` | head, 16-way split: `infer` (one row → `logits1…16`), `verify` (128 rows → argmaxes) | 306 MB |
|
| 46 |
+
| `embed_tokens.f16.bin` | token embeddings, 151 936 × 2048 fp16, row-major, no header | 622 MB |
|
| 47 |
| `tokenizer.json`, `tokenizer_config.json` | the original tokenizer | 11 MB |
|
| 48 |
+
| `MODEL_LICENSE`, `LICENSE-Qwen3-ASR.txt`, `NOTICE`, `SHA256SUMS` | | |
|
| 49 |
+
|
| 50 |
+
## Use
|
| 51 |
+
|
| 52 |
+
These files are driven by the R2T2 provider of a VoiceInk fork, which downloads them from here
|
| 53 |
+
after the licence is accepted. The runtime re-decodes the whole current piece of audio (up to 30 s)
|
| 54 |
+
on every pass, so the live text converges on the same result as an offline decode; what makes that
|
| 55 |
+
cheap is keeping the prompt prefix and every completed encoder window in the decoder's KV state,
|
| 56 |
+
prefilling only the new rows plus the previous pass's tokens as a draft, confirming the draft with
|
| 57 |
+
one `verify` call and decoding only from the first divergence. Passes take 150–250 ms on an M5 Pro;
|
| 58 |
+
a 30 s piece decodes offline in about 2.5 s.
|
| 59 |
+
|
| 60 |
+
For another runtime, the interface:
|
| 61 |
+
|
| 62 |
+
- Decoder inputs `hidden_states [1, B, 2048]` fp16, `position_ids [B]` int32,
|
| 63 |
+
`causal_mask [1, 1, B, 1024]` fp16 (0 to attend, −10 000 otherwise), `current_pos [1]` int32;
|
| 64 |
+
output `output_hidden_states [1, B, 2048]`. Both chunks share one `MLState` of shape
|
| 65 |
+
`(56, 8, 1024, 128)` fp16, written by absolute position, so a caller can rewind and overwrite.
|
| 66 |
+
- Head `verify` output per row and slice: `argmax_val`, `argmax_hi`, `argmax_lo` (`[128, 16]` fp16),
|
| 67 |
+
the slice's best logit and its index as `hi × 64 + lo`.
|
| 68 |
+
- Prompt: the Qwen3-ASR chat template, with the encoder rows in place of the `<|audio_pad|>`
|
| 69 |
+
embeddings and, to force a language, the assistant turn opened with `language <Name><asr_text>`.
|
| 70 |
+
The model's `|` marks where it stops trusting its own output.
|
| 71 |
+
- Mel: Whisper's recipe (16 kHz, n_fft 400, hop 160, 128 Slaney bins, `log10`, clamped to 8 dB below
|
| 72 |
+
the buffer's maximum, `(x + 4) / 4`).
|
| 73 |
+
- The decoder loads under `.cpuAndNeuralEngine` only, and Python coremltools cannot open the
|
| 74 |
+
multifunction packages. The first load by an application compiles the programs (about 50 s);
|
| 75 |
+
they are cached per application binary afterwards.
|
| 76 |
+
|
| 77 |
+
## Notice
|
| 78 |
+
|
| 79 |
+
Any modifications made to the original model in this Derivative Work are not endorsed, warranted,
|
| 80 |
+
or guaranteed by the original right-holder of the original model, and the original right-holder
|
| 81 |
+
disclaims all liability related to this Derivative Work.
|
| 82 |
+
|
| 83 |
+
This is a conversion of NetEase Youdao's Confucius4-R2T2 and is governed by the original
|
| 84 |
+
[MODEL_LICENSE](MODEL_LICENSE). That licence is royalty-free for most users, including commercial
|
| 85 |
+
use, but requires a separate licence from NetEase Youdao above 100 million monthly active users or
|
| 86 |
+
RMB 1 billion in annual revenue, forbids using the model to improve other AI models and use in the
|
| 87 |
+
high-risk scenarios it lists, and binds anyone you redistribute it to. Using these files means
|
| 88 |
+
accepting it; keep `MODEL_LICENSE` and `NOTICE` with every copy. The base model, Qwen3-ASR-1.7B, is
|
| 89 |
+
Apache 2.0 (`LICENSE-Qwen3-ASR.txt`).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|