Nairod785 commited on
Commit
0516975
Β·
verified Β·
1 Parent(s): a83f1b9

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +27 -60
README.md CHANGED
@@ -52,80 +52,45 @@ tags:
52
  - transcribe.cpp
53
  ---
54
 
55
- # Confucius4-R2T2 β€” Q4_K_M GGUF (Arm M)
56
 
57
  Optimized **Q4_K_M GGUF** quantization of **NetEase Youdao Confucius4-R2T2**, a low-latency, append-only streaming ASR model (fine-tuned from Qwen3-ASR-1.7B with Longest Stable Prefix decoding), ready for native execution in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp) and [audio.cpp](https://github.com/0xShug0/audio.cpp).
58
 
59
  * **Original model:** [netease-youdao/Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2) Β· [source code](https://github.com/netease-youdao/Confucius4-R2T2) Β· [demo](https://r2t2.youdao.com/demo)
60
  * **Base Architecture:** Qwen3-ASR audio encoder (24 layers, 1024 width, 2048 projected width) + Causal LM decoder (28 layers, 16 query heads, 8 KV heads, 128 head width).
61
- * **Reference GGUF repository & inspiration:** [davidxifeng/Confucius4-R2T2-gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf) by David Xi Feng
62
- * **Quantization & Benchmarking:** Researched, validated, and tuned in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp).
63
 
64
  ---
65
 
66
- ## The Breakthrough: Solving the Low-Bit Quantization Barrier
67
 
68
- In previous GGUF releases (e.g. `davidxifeng/Confucius4-R2T2-gguf`), precision was restricted:
69
- > *"Precision policy: Q8_0 and higher are supported. 4-bit and 5-bit quantizations (legacy and k-quant) are rejected at load time with an actionable error, because this graph's kernels are not validated below Q8_0 and otherwise decode to empty text."*
70
 
71
- Uniform 4-bit and 5-bit quantization suffered from two severe failure modes:
72
- 1. **The Empty Output Cliff:** Low-bit matmuls below Q4_K (e.g., tower or gate/up at Q3_K) produce abrupt total collapseβ€”emitting empty output rather than degraded text.
73
- 2. **Multilingual Language Drift:** Standard Q4_K and Q5_K quantizations passed simple English (`jfk`) and Chinese (`zh`) tests, but silently flipped to English when fed German or Russian audio.
74
 
75
- ### The Arm M Recipe
 
 
 
76
 
77
- Through exhaustive per-tensor ablation and cross-lingual probing across 20+ experimental arms, the optimal configuration (**Arm M**) was isolated:
78
 
79
- - **Filename:** `r2t2-q4_k_m.gguf`
80
- - **File size:** **1.187 GB** (1,186,939,968 bytes)
81
- - **Footprint reduction:** **52% smaller** than the 2.478 GB Q8_0 reference.
82
- - **Multilingual stability:** 100% passes on English, Chinese, German, and French.
83
-
84
- | Component | Tensors | Target Dtype | Architectural Mechanism |
85
- |---|---|---|---|
86
- | **Audio Tower** | 147 | `BF16 / Q4_K` | Audio encoder layers quantized to Q4_K while keeping converter-selected projection weights at BF16. |
87
- | **Attention Projections** | 112 (`q`, `k`, `v`, `o_proj`) | `Q4_K` | Quantized within the LM error budget when paired with Q6_K down projections. |
88
- | **MLP Gate / Up** | 56 (`gate_proj`, `up_proj`) | `Q4_K` | Errors pass through SwiGLU non-linearities, bounding error propagation. |
89
- | **MLP Down** | 28 (`down_proj`) | **`Q6_K` (Floor)** | **Non-negotiable hard floor.** `down_proj` writes straight into the residual stream, compounding errors across all subsequent layers. Quantizing below Q6_K triggers severe English language drift on non-English audio. |
90
- | **Embeddings** | 1 (`embed_tokens`) | **`Q2_K`** | Lookup table error enters once and does not compound. Slices 311M parameters down with no measurable fidelity penalty. |
91
-
92
- ---
93
-
94
- ## Why Arm M Works
95
-
96
- ### 1. Cumulative Error Budget
97
- A transformer block cannot be quantized in isolation; its tolerance depends on what the rest of the model has spent:
98
- - **Da4** (Attention @ Q4_K, `down_proj` @ Q6_K) passes.
99
- - **E** (Attention @ Q6_K, `down_proj` @ Q4_K) passes simple probes.
100
- - **G** (Both Attention & `down_proj` @ Q4_K) **fails completely**.
101
-
102
- The LM decoder can afford **exactly one** of {attention, `down_proj`} at Q4_K. Arm M allocates that budget to attention while pinning `down_proj` strictly to Q6_K.
103
-
104
- ### 2. `down_proj` @ Q6_K is the Hard Floor
105
- Every arm evaluated that placed `down_proj` at Q5_K or Q4_K (D5, E, L, N, K, G, G3) emitted English when given German audio. No arm with `down_proj` below Q6_K passes multilingual validation.
106
-
107
- ### 3. Embeddings are the Cheapest Lever
108
- `embed_tokens` has 311M parameters. Quantizing it to **Q2_K** drops the total model size from 1.26 GB (Da4) to **1.187 GB** without triggering degradation.
109
 
110
  ---
111
 
112
- ## Arm Evaluation & Benchmarks
113
-
114
- Measured on CPU (pooled ratio against Q8_0 baseline on JFK and ZH test audio):
115
 
116
- | Arm | Attention | Gate/Up | `down_proj` | Embed | Tower | Size (GB) | vs Q8_0 Speed | Status |
117
- |---|---|---|---|---|---|---|---|---|
118
- | `r2t2-q8_0` | Q8_0 | Q8_0 | Q8_0 | F16 | BF16/Q8_0 | 2.478 | 0% (baseline) | Reference |
119
- | `r2t2-q6_k` | Q6_K | Q6_K | Q6_K | F16 | BF16/Q6_K | 2.060 | +5% | Passes |
120
- | `r2t2-q5_k` | Q5_K | Q5_K | Q5_K | F16 | BF16/Q5_K | 1.832 | **-20% (slower)** | Fails Russian (flips to English) |
121
- | `r2t2-Da4` | Q4_K | Q4_K | Q6_K | Q4_K | BF16/Q4_K | 1.260 | β€” | Passes |
122
- | **`r2t2-q4_k_m` (M)** | **Q4_K** | **Q4_K** | **Q6_K** | **Q2_K** | **BF16/Q4_K** | **1.187** | **+20%+** | **Floor β€” Passes All** |
123
- | `r2t2-D5` | Q6_K | Q4_K | **Q5_K** | Q4_K | BF16/Q4_K | 1.304 | +18% | Fails German |
124
- | `r2t2-E` | Q6_K | Q4_K | **Q4_K** | Q4_K | BF16/Q4_K | 1.260 | +22% | Fails German |
125
- | `r2t2-G` | Q4_K | Q4_K | **Q4_K** | Q4_K | BF16/Q4_K | 1.169 | β€” | Fails German |
126
- | `r2t2-P` | Q4_K | **Q3_K** | Q6_K | Q2_K | BF16/Q4_K | 1.093 | β€” | Empty output cliff |
127
 
128
- *Note: Q5_K is strictly dominated β€” it runs ~20% slower than Q8_0 due to non-vectorized kernel routines and fails multilingual validation. Arm M provides both higher throughput and minimal size.*
129
 
130
  ---
131
 
@@ -133,7 +98,10 @@ Measured on CPU (pooled ratio against Q8_0 baseline on JFK and ZH test audio):
133
 
134
  | File | Quantization | Size | Description |
135
  |---|---|---|---|
136
- | `r2t2-q4_k_m.gguf` | Q4_K_M (Arm M) | 1.187 GB | Highly optimized floor quant; passes multilingual validation |
 
 
 
137
 
138
  ---
139
 
@@ -156,11 +124,10 @@ audiocpp_cli --task asr --mode streaming --family confucius4_r2t2 --model r2t2
156
  ### Server (OpenAI-compatible)
157
  ```bash
158
  curl http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F file=@speech.wav
159
- curl -N http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F stream=true -F file=@speech.wav # SSE transcript deltas
160
  ```
161
 
162
- ### Streaming Parameters
163
- Chunk sizes from **80 ms to 2000 ms** are supported (320 ms recommended default). Pass `--language Chinese` to skip language detection and `--text "hotword, term"` for context/hotword prompting.
164
 
165
  ---
166
 
 
52
  - transcribe.cpp
53
  ---
54
 
55
+ # Confucius4-R2T2 β€” Q4_K_M GGUF
56
 
57
  Optimized **Q4_K_M GGUF** quantization of **NetEase Youdao Confucius4-R2T2**, a low-latency, append-only streaming ASR model (fine-tuned from Qwen3-ASR-1.7B with Longest Stable Prefix decoding), ready for native execution in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp) and [audio.cpp](https://github.com/0xShug0/audio.cpp).
58
 
59
  * **Original model:** [netease-youdao/Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2) Β· [source code](https://github.com/netease-youdao/Confucius4-R2T2) Β· [demo](https://r2t2.youdao.com/demo)
60
  * **Base Architecture:** Qwen3-ASR audio encoder (24 layers, 1024 width, 2048 projected width) + Causal LM decoder (28 layers, 16 query heads, 8 KV heads, 128 head width).
61
+ * **Reference GGUF repo & inspiration:** [davidxifeng/Confucius4-R2T2-gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf) by David Xi Feng
62
+ * **Quantization & Benchmarking:** Researched and benchmarked in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp).
63
 
64
  ---
65
 
66
+ ## Technical Overview
67
 
68
+ Previous GGUF releases were constrained to Q8_0 or higher because uniform 4-bit and 5-bit quantizations either suffered from catastrophic empty-output cliffs or exhibited silent language drift (e.g. flipping to English when transcribing German or Russian).
 
69
 
70
+ This release introduces an **asymmetric per-tensor quantization recipe** that breaks through that barrier:
 
 
71
 
72
+ - **File:** `r2t2-q4_k_m.gguf`
73
+ - **File size:** **1.187 GB** (1,186,939,968 bytes) β€” **52% smaller** than the 2.478 GB Q8_0 reference.
74
+ - **Accuracy & Parity:** Preserves multilingual stability (verified across English, Chinese, German, and French).
75
+ - **Speed:** Faster than Q8_0 on CPU, avoiding the slow non-vectorized paths found in Q5_K.
76
 
77
+ ### Quantization Composition
78
 
79
+ | Block | Target Dtype | Architectural Role |
80
+ |---|---|---|
81
+ | **Audio Tower** | `BF16 / Q4_K` | Audio encoder layers quantized to Q4_K with sensitive projection weights preserved at BF16. |
82
+ | **Attention Projections** | `Q4_K` | Fits within the LM error budget when down projections are held at higher precision. |
83
+ | **MLP Gate / Up** | `Q4_K` | Errors are bounded through SwiGLU activation. |
84
+ | **MLP Down** | **`Q6_K`** | **Hard Floor.** Writes into residual stream; lower precision triggers multilingual degradation. |
85
+ | **Embeddings** | **`Q2_K`** | Lookup table error does not compound; cuts 311M parameters with negligible loss. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86
 
87
  ---
88
 
89
+ ## Comprehensive Quantization Report
 
 
90
 
91
+ For the full ablation study, complete empirical test tables across 20+ quantization arms, failure mode analyses, speed benchmarks, and conversion reproduction instructions:
 
 
 
 
 
 
 
 
 
 
92
 
93
+ πŸ‘‰ **Read the full report in [QUANTIZATION.md](QUANTIZATION.md)**
94
 
95
  ---
96
 
 
98
 
99
  | File | Quantization | Size | Description |
100
  |---|---|---|---|
101
+ | `r2t2-q4_k_m.gguf` | Q4_K_M | 1.187 GB | High-performance floor quantization; passes multilingual validation |
102
+ | `QUANTIZATION.md` | β€” | β€” | Detailed technical report, ablation study, and benchmarks |
103
+ | `NOTICE` | β€” | β€” | Attribution and derivative work disclaimers |
104
+ | `LICENSE` / `LICENSE_zh` | β€” | β€” | NetEase Youdao Model Use License Agreement |
105
 
106
  ---
107
 
 
124
  ### Server (OpenAI-compatible)
125
  ```bash
126
  curl http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F file=@speech.wav
127
+ curl -N http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F stream=true -F file=@speech.wav
128
  ```
129
 
130
+ Streaming chunk sizes from **80 ms to 2000 ms** are supported (320 ms recommended default).
 
131
 
132
  ---
133