|
Download README.md from lunks/Confucius4-R2T2-CoreML-ANE: direct link, hf CLI and curl.
- Browser
- Download file 6.21 kB
-
https://huggingface.co/lunks/Confucius4-R2T2-CoreML-ANE/resolve/main/README.md
- Command line
-
hf download hf://lunks/Confucius4-R2T2-CoreML-ANE/README.md
-
curl -L -o README.md https://huggingface.co/lunks/Confucius4-R2T2-CoreML-ANE/resolve/main/README.md
6.21 kB
| license: other | |
| license_name: netease-model-use-license-agreement | |
| license_link: https://github.com/netease-youdao/Confucius4-R2T2/blob/master/MODEL_LICENSE | |
| base_model: netease-youdao/Confucius4-R2T2 | |
| base_model_relation: quantized | |
| pipeline_tag: automatic-speech-recognition | |
| library_name: coreml | |
| tags: | |
| - coreml | |
| - apple-neural-engine | |
| - speech-to-text | |
| - asr | |
| - stt | |
| - streaming | |
| - confucius4 | |
| - r2t2 | |
| # Confucius4-R2T2, Core ML for the Apple Neural Engine | |
| [`netease-youdao/Confucius4-R2T2`](https://huggingface.co/netease-youdao/Confucius4-R2T2) | |
| (a real-time speech-recognition fine-tune of Qwen3-ASR-1.7B, 30 languages) converted to Core ML | |
| programs that run on the Apple Neural Engine: the audio encoder in fp16, the Qwen3 decoder and the | |
| LM head palettised to 8 bits. 4.09 GB of bfloat16 becomes 2.8 GB on disk; while loaded, the | |
| Neural Engine holds the weights dequantised to fp16, about 4.5 GB, outside the process. | |
| Converted with the pipeline in the VoiceInk fork's `tools/r2t2-coreml`: the encoder through | |
| coremltools 9 with a fixed 800-frame window (the fused `gelu` replaced by a tanh formulation, whose | |
| constant absolute error otherwise swamps this encoder's small activations), the decoder through | |
| [ANEMLL](https://github.com/Anemll/Anemll) as two stateful chunks whose KV state holds two | |
| sequences side by side ("lanes"), with a 128-row `prefill` of one lane and an `infer` that | |
| decodes one token of each lane in a single pass over the weights, and a head with a two-row | |
| `infer` and a `verify` function that returns the argmax of 128 rows in one call. Requires macOS 15 | |
| or later (Core ML stateful models) on Apple silicon. | |
| Word error rate, whole utterances, 100-utterance subsets: LibriSpeech test-clean 2.40 %, FLEURS | |
| pt_br 3.50 %. Against the unconverted bfloat16 weights on MLX, paired on 300 utterances of each: | |
| 2.47 % vs 2.42 % and 4.01 % vs 4.01 %. | |
| ## Contents | |
| | file | | size | | |
| |---|---|---| | |
| | `R2T2AudioEncoder.mlmodelc` | encoder, fp16, `[1, 128, 800]` mel window + key mask → `[1, 104, 2048]` | 607 MB | | |
| | `r2t2_FFN_PF_lut8_chunk_01of02.mlmodelc` | decoder layers 0–13, functions `prefill` (128 rows of one lane) and `infer` (one row per lane) | 692 MB | | |
| | `r2t2_FFN_PF_lut8_chunk_02of02.mlmodelc` | decoder layers 14–27 and the final norm, same functions | 692 MB | | |
| | `r2t2_lm_head_lut8.mlmodelc` | head, 16-way split: `infer` (two rows → `logits1…16`), `verify` (128 rows → argmaxes) | 306 MB | | |
| | `embed_tokens.f16.bin` | token embeddings, 151 936 × 2048 fp16, row-major, no header | 622 MB | | |
| | `tokenizer.json`, `tokenizer_config.json` | the original tokenizer | 11 MB | | |
| | `MODEL_LICENSE`, `LICENSE-Qwen3-ASR.txt`, `NOTICE`, `SHA256SUMS` | | | | |
| ## Use | |
| These files are driven by the R2T2 provider of a VoiceInk fork, which downloads them from here | |
| after the licence is accepted. The runtime re-decodes the whole current piece of audio (up to 30 s) | |
| on every pass, so the live text converges on the same result as an offline decode; what makes that | |
| cheap is keeping the prompt prefix and every completed encoder window in the decoder's KV state, | |
| prefilling only the new rows plus the previous pass's tokens as a draft, confirming the draft with | |
| one `verify` call and decoding only from the first divergence. Passes take 150–250 ms on an M5 Pro. | |
| Offline, two 30 s pieces decode together, one token of each per `infer` call: a 120 s recording | |
| takes about 6 s (8.7 ms per token; 17 ms for a single piece). | |
| For another runtime, the interface: | |
| - Both chunks share one `MLState` of shape `(56, 8, 2048, 128)` fp16: lane *l* owns cache positions | |
| `l·1024 ..< (l+1)·1024`, written by absolute position, so a caller can rewind and overwrite. The | |
| causal mask is 0 to attend, −10 000 otherwise; output `output_hidden_states [1, B, 2048]`. | |
| - `prefill`: `hidden_states [1, 128, 2048]` fp16, `position_ids [128]` int32 (the lane's own | |
| positions), `causal_mask [1, 1, 128, 1024]`, `current_pos [1]` int32 (absolute: `l·1024 + p`), | |
| `lane_weights [2, 1, 1]` fp16 one-hot selecting the lane. | |
| - `infer`: `hidden_states [1, 2, 2048]` (one row per lane), `position_ids [2]`, | |
| `causal_mask [1, 1, 4, 1024]` (each lane's row twice: the two query heads of a K/V head share a | |
| matmul row axis), `current_pos [1]` and `current_pos_b [1]` (lane 0 and lane 1, absolute). An | |
| idle lane takes a zero row, a fully masked row and a scratch position its next real row overwrites. | |
| - Head `infer` takes `hidden_states [1, 2, 2048]` and returns `logits1…16` for both rows; `verify` | |
| takes 128 rows and returns per row and slice `argmax_val`, `argmax_hi`, `argmax_lo` (`[128, 16]` | |
| fp16), the slice's best logit and its index as `hi × 64 + lo`. | |
| - Prompt: the Qwen3-ASR chat template, with the encoder rows in place of the `<|audio_pad|>` | |
| embeddings and, to force a language, the assistant turn opened with `language <Name><asr_text>`. | |
| The model's `|` marks where it stops trusting its own output. | |
| - Mel: Whisper's recipe (16 kHz, n_fft 400, hop 160, 128 Slaney bins, `log10`, clamped to 8 dB below | |
| the buffer's maximum, `(x + 4) / 4`). | |
| - The decoder loads under `.cpuAndNeuralEngine` only, and Python coremltools cannot open the | |
| multifunction packages. The first load by an application compiles the programs (about 50 s); | |
| they are cached per application binary afterwards. | |
| ## Notice | |
| Any modifications made to the original model in this Derivative Work are not endorsed, warranted, | |
| or guaranteed by the original right-holder of the original model, and the original right-holder | |
| disclaims all liability related to this Derivative Work. | |
| This is a conversion of NetEase Youdao's Confucius4-R2T2 and is governed by the original | |
| [MODEL_LICENSE](MODEL_LICENSE). That licence is royalty-free for most users, including commercial | |
| use, but requires a separate licence from NetEase Youdao above 100 million monthly active users or | |
| RMB 1 billion in annual revenue, forbids using the model to improve other AI models and use in the | |
| high-risk scenarios it lists, and binds anyone you redistribute it to. Using these files means | |
| accepting it; keep `MODEL_LICENSE` and `NOTICE` with every copy. The base model, Qwen3-ASR-1.7B, is | |
| Apache 2.0 (`LICENSE-Qwen3-ASR.txt`). | |