File size: 6,206 Bytes
d0e38c6
 
 
f34b38d
d0e38c6
c54339a
d0e38c6
 
 
 
 
f34b38d
d0e38c6
f34b38d
d0e38c6
 
 
 
 
f34b38d
d0e38c6
f34b38d
d0e38c6
f34b38d
 
 
 
 
 
 
b8b7f8c
 
 
 
 
d0e38c6
f34b38d
 
 
d0e38c6
 
 
f34b38d
d0e38c6
f34b38d
b8b7f8c
f34b38d
b8b7f8c
f34b38d
d0e38c6
f34b38d
 
 
 
 
 
 
 
 
b8b7f8c
 
 
f34b38d
 
 
b8b7f8c
 
 
 
 
 
 
 
 
 
 
 
 
f34b38d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
---
license: other
license_name: netease-model-use-license-agreement
license_link: https://github.com/netease-youdao/Confucius4-R2T2/blob/master/MODEL_LICENSE
base_model: netease-youdao/Confucius4-R2T2
base_model_relation: quantized
pipeline_tag: automatic-speech-recognition
library_name: coreml
tags:
- coreml
- apple-neural-engine
- speech-to-text
- asr
- stt
- streaming
- confucius4
- r2t2
---

# Confucius4-R2T2, Core ML for the Apple Neural Engine

[`netease-youdao/Confucius4-R2T2`](https://huggingface.co/netease-youdao/Confucius4-R2T2)
(a real-time speech-recognition fine-tune of Qwen3-ASR-1.7B, 30 languages) converted to Core ML
programs that run on the Apple Neural Engine: the audio encoder in fp16, the Qwen3 decoder and the
LM head palettised to 8 bits. 4.09 GB of bfloat16 becomes 2.8 GB on disk; while loaded, the
Neural Engine holds the weights dequantised to fp16, about 4.5 GB, outside the process.

Converted with the pipeline in the VoiceInk fork's `tools/r2t2-coreml`: the encoder through
coremltools 9 with a fixed 800-frame window (the fused `gelu` replaced by a tanh formulation, whose
constant absolute error otherwise swamps this encoder's small activations), the decoder through
[ANEMLL](https://github.com/Anemll/Anemll) as two stateful chunks whose KV state holds two
sequences side by side ("lanes"), with a 128-row `prefill` of one lane and an `infer` that
decodes one token of each lane in a single pass over the weights, and a head with a two-row
`infer` and a `verify` function that returns the argmax of 128 rows in one call. Requires macOS 15
or later (Core ML stateful models) on Apple silicon.

Word error rate, whole utterances, 100-utterance subsets: LibriSpeech test-clean 2.40 %, FLEURS
pt_br 3.50 %. Against the unconverted bfloat16 weights on MLX, paired on 300 utterances of each:
2.47 % vs 2.42 % and 4.01 % vs 4.01 %.

## Contents

| file | | size |
|---|---|---|
| `R2T2AudioEncoder.mlmodelc` | encoder, fp16, `[1, 128, 800]` mel window + key mask → `[1, 104, 2048]` | 607 MB |
| `r2t2_FFN_PF_lut8_chunk_01of02.mlmodelc` | decoder layers 0–13, functions `prefill` (128 rows of one lane) and `infer` (one row per lane) | 692 MB |
| `r2t2_FFN_PF_lut8_chunk_02of02.mlmodelc` | decoder layers 14–27 and the final norm, same functions | 692 MB |
| `r2t2_lm_head_lut8.mlmodelc` | head, 16-way split: `infer` (two rows → `logits1…16`), `verify` (128 rows → argmaxes) | 306 MB |
| `embed_tokens.f16.bin` | token embeddings, 151 936 × 2048 fp16, row-major, no header | 622 MB |
| `tokenizer.json`, `tokenizer_config.json` | the original tokenizer | 11 MB |
| `MODEL_LICENSE`, `LICENSE-Qwen3-ASR.txt`, `NOTICE`, `SHA256SUMS` | | |

## Use

These files are driven by the R2T2 provider of a VoiceInk fork, which downloads them from here
after the licence is accepted. The runtime re-decodes the whole current piece of audio (up to 30 s)
on every pass, so the live text converges on the same result as an offline decode; what makes that
cheap is keeping the prompt prefix and every completed encoder window in the decoder's KV state,
prefilling only the new rows plus the previous pass's tokens as a draft, confirming the draft with
one `verify` call and decoding only from the first divergence. Passes take 150–250 ms on an M5 Pro.
Offline, two 30 s pieces decode together, one token of each per `infer` call: a 120 s recording
takes about 6 s (8.7 ms per token; 17 ms for a single piece).

For another runtime, the interface:

- Both chunks share one `MLState` of shape `(56, 8, 2048, 128)` fp16: lane *l* owns cache positions
  `l·1024 ..< (l+1)·1024`, written by absolute position, so a caller can rewind and overwrite. The
  causal mask is 0 to attend, −10 000 otherwise; output `output_hidden_states [1, B, 2048]`.
  - `prefill`: `hidden_states [1, 128, 2048]` fp16, `position_ids [128]` int32 (the lane's own
    positions), `causal_mask [1, 1, 128, 1024]`, `current_pos [1]` int32 (absolute: `l·1024 + p`),
    `lane_weights [2, 1, 1]` fp16 one-hot selecting the lane.
  - `infer`: `hidden_states [1, 2, 2048]` (one row per lane), `position_ids [2]`,
    `causal_mask [1, 1, 4, 1024]` (each lane's row twice: the two query heads of a K/V head share a
    matmul row axis), `current_pos [1]` and `current_pos_b [1]` (lane 0 and lane 1, absolute). An
    idle lane takes a zero row, a fully masked row and a scratch position its next real row overwrites.
- Head `infer` takes `hidden_states [1, 2, 2048]` and returns `logits1…16` for both rows; `verify`
  takes 128 rows and returns per row and slice `argmax_val`, `argmax_hi`, `argmax_lo` (`[128, 16]`
  fp16), the slice's best logit and its index as `hi × 64 + lo`.
- Prompt: the Qwen3-ASR chat template, with the encoder rows in place of the `<|audio_pad|>`
  embeddings and, to force a language, the assistant turn opened with `language <Name><asr_text>`.
  The model's `|` marks where it stops trusting its own output.
- Mel: Whisper's recipe (16 kHz, n_fft 400, hop 160, 128 Slaney bins, `log10`, clamped to 8 dB below
  the buffer's maximum, `(x + 4) / 4`).
- The decoder loads under `.cpuAndNeuralEngine` only, and Python coremltools cannot open the
  multifunction packages. The first load by an application compiles the programs (about 50 s);
  they are cached per application binary afterwards.

## Notice

Any modifications made to the original model in this Derivative Work are not endorsed, warranted,
or guaranteed by the original right-holder of the original model, and the original right-holder
disclaims all liability related to this Derivative Work.

This is a conversion of NetEase Youdao's Confucius4-R2T2 and is governed by the original
[MODEL_LICENSE](MODEL_LICENSE). That licence is royalty-free for most users, including commercial
use, but requires a separate licence from NetEase Youdao above 100 million monthly active users or
RMB 1 billion in annual revenue, forbids using the model to improve other AI models and use in the
high-risk scenarios it lists, and binds anyone you redistribute it to. Using these files means
accepting it; keep `MODEL_LICENSE` and `NOTICE` with every copy. The base model, Qwen3-ASR-1.7B, is
Apache 2.0 (`LICENSE-Qwen3-ASR.txt`).