Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -52,80 +52,45 @@ tags:
|
|
| 52 |
- transcribe.cpp
|
| 53 |
---
|
| 54 |
|
| 55 |
-
# Confucius4-R2T2 β Q4_K_M GGUF
|
| 56 |
|
| 57 |
Optimized **Q4_K_M GGUF** quantization of **NetEase Youdao Confucius4-R2T2**, a low-latency, append-only streaming ASR model (fine-tuned from Qwen3-ASR-1.7B with Longest Stable Prefix decoding), ready for native execution in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp) and [audio.cpp](https://github.com/0xShug0/audio.cpp).
|
| 58 |
|
| 59 |
* **Original model:** [netease-youdao/Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2) Β· [source code](https://github.com/netease-youdao/Confucius4-R2T2) Β· [demo](https://r2t2.youdao.com/demo)
|
| 60 |
* **Base Architecture:** Qwen3-ASR audio encoder (24 layers, 1024 width, 2048 projected width) + Causal LM decoder (28 layers, 16 query heads, 8 KV heads, 128 head width).
|
| 61 |
-
* **Reference GGUF
|
| 62 |
-
* **Quantization & Benchmarking:** Researched
|
| 63 |
|
| 64 |
---
|
| 65 |
|
| 66 |
-
##
|
| 67 |
|
| 68 |
-
|
| 69 |
-
> *"Precision policy: Q8_0 and higher are supported. 4-bit and 5-bit quantizations (legacy and k-quant) are rejected at load time with an actionable error, because this graph's kernels are not validated below Q8_0 and otherwise decode to empty text."*
|
| 70 |
|
| 71 |
-
|
| 72 |
-
1. **The Empty Output Cliff:** Low-bit matmuls below Q4_K (e.g., tower or gate/up at Q3_K) produce abrupt total collapseβemitting empty output rather than degraded text.
|
| 73 |
-
2. **Multilingual Language Drift:** Standard Q4_K and Q5_K quantizations passed simple English (`jfk`) and Chinese (`zh`) tests, but silently flipped to English when fed German or Russian audio.
|
| 74 |
|
| 75 |
-
|
|
|
|
|
|
|
|
|
|
| 76 |
|
| 77 |
-
|
| 78 |
|
| 79 |
-
|
| 80 |
-
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
|
| 85 |
-
|
|
| 86 |
-
| **Audio Tower** | 147 | `BF16 / Q4_K` | Audio encoder layers quantized to Q4_K while keeping converter-selected projection weights at BF16. |
|
| 87 |
-
| **Attention Projections** | 112 (`q`, `k`, `v`, `o_proj`) | `Q4_K` | Quantized within the LM error budget when paired with Q6_K down projections. |
|
| 88 |
-
| **MLP Gate / Up** | 56 (`gate_proj`, `up_proj`) | `Q4_K` | Errors pass through SwiGLU non-linearities, bounding error propagation. |
|
| 89 |
-
| **MLP Down** | 28 (`down_proj`) | **`Q6_K` (Floor)** | **Non-negotiable hard floor.** `down_proj` writes straight into the residual stream, compounding errors across all subsequent layers. Quantizing below Q6_K triggers severe English language drift on non-English audio. |
|
| 90 |
-
| **Embeddings** | 1 (`embed_tokens`) | **`Q2_K`** | Lookup table error enters once and does not compound. Slices 311M parameters down with no measurable fidelity penalty. |
|
| 91 |
-
|
| 92 |
-
---
|
| 93 |
-
|
| 94 |
-
## Why Arm M Works
|
| 95 |
-
|
| 96 |
-
### 1. Cumulative Error Budget
|
| 97 |
-
A transformer block cannot be quantized in isolation; its tolerance depends on what the rest of the model has spent:
|
| 98 |
-
- **Da4** (Attention @ Q4_K, `down_proj` @ Q6_K) passes.
|
| 99 |
-
- **E** (Attention @ Q6_K, `down_proj` @ Q4_K) passes simple probes.
|
| 100 |
-
- **G** (Both Attention & `down_proj` @ Q4_K) **fails completely**.
|
| 101 |
-
|
| 102 |
-
The LM decoder can afford **exactly one** of {attention, `down_proj`} at Q4_K. Arm M allocates that budget to attention while pinning `down_proj` strictly to Q6_K.
|
| 103 |
-
|
| 104 |
-
### 2. `down_proj` @ Q6_K is the Hard Floor
|
| 105 |
-
Every arm evaluated that placed `down_proj` at Q5_K or Q4_K (D5, E, L, N, K, G, G3) emitted English when given German audio. No arm with `down_proj` below Q6_K passes multilingual validation.
|
| 106 |
-
|
| 107 |
-
### 3. Embeddings are the Cheapest Lever
|
| 108 |
-
`embed_tokens` has 311M parameters. Quantizing it to **Q2_K** drops the total model size from 1.26 GB (Da4) to **1.187 GB** without triggering degradation.
|
| 109 |
|
| 110 |
---
|
| 111 |
|
| 112 |
-
##
|
| 113 |
-
|
| 114 |
-
Measured on CPU (pooled ratio against Q8_0 baseline on JFK and ZH test audio):
|
| 115 |
|
| 116 |
-
|
| 117 |
-
|---|---|---|---|---|---|---|---|---|
|
| 118 |
-
| `r2t2-q8_0` | Q8_0 | Q8_0 | Q8_0 | F16 | BF16/Q8_0 | 2.478 | 0% (baseline) | Reference |
|
| 119 |
-
| `r2t2-q6_k` | Q6_K | Q6_K | Q6_K | F16 | BF16/Q6_K | 2.060 | +5% | Passes |
|
| 120 |
-
| `r2t2-q5_k` | Q5_K | Q5_K | Q5_K | F16 | BF16/Q5_K | 1.832 | **-20% (slower)** | Fails Russian (flips to English) |
|
| 121 |
-
| `r2t2-Da4` | Q4_K | Q4_K | Q6_K | Q4_K | BF16/Q4_K | 1.260 | β | Passes |
|
| 122 |
-
| **`r2t2-q4_k_m` (M)** | **Q4_K** | **Q4_K** | **Q6_K** | **Q2_K** | **BF16/Q4_K** | **1.187** | **+20%+** | **Floor β Passes All** |
|
| 123 |
-
| `r2t2-D5` | Q6_K | Q4_K | **Q5_K** | Q4_K | BF16/Q4_K | 1.304 | +18% | Fails German |
|
| 124 |
-
| `r2t2-E` | Q6_K | Q4_K | **Q4_K** | Q4_K | BF16/Q4_K | 1.260 | +22% | Fails German |
|
| 125 |
-
| `r2t2-G` | Q4_K | Q4_K | **Q4_K** | Q4_K | BF16/Q4_K | 1.169 | β | Fails German |
|
| 126 |
-
| `r2t2-P` | Q4_K | **Q3_K** | Q6_K | Q2_K | BF16/Q4_K | 1.093 | β | Empty output cliff |
|
| 127 |
|
| 128 |
-
*
|
| 129 |
|
| 130 |
---
|
| 131 |
|
|
@@ -133,7 +98,10 @@ Measured on CPU (pooled ratio against Q8_0 baseline on JFK and ZH test audio):
|
|
| 133 |
|
| 134 |
| File | Quantization | Size | Description |
|
| 135 |
|---|---|---|---|
|
| 136 |
-
| `r2t2-q4_k_m.gguf` | Q4_K_M
|
|
|
|
|
|
|
|
|
|
| 137 |
|
| 138 |
---
|
| 139 |
|
|
@@ -156,11 +124,10 @@ audiocpp_cli --task asr --mode streaming --family confucius4_r2t2 --model r2t2
|
|
| 156 |
### Server (OpenAI-compatible)
|
| 157 |
```bash
|
| 158 |
curl http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F file=@speech.wav
|
| 159 |
-
curl -N http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F stream=true -F file=@speech.wav
|
| 160 |
```
|
| 161 |
|
| 162 |
-
|
| 163 |
-
Chunk sizes from **80 ms to 2000 ms** are supported (320 ms recommended default). Pass `--language Chinese` to skip language detection and `--text "hotword, term"` for context/hotword prompting.
|
| 164 |
|
| 165 |
---
|
| 166 |
|
|
|
|
| 52 |
- transcribe.cpp
|
| 53 |
---
|
| 54 |
|
| 55 |
+
# Confucius4-R2T2 β Q4_K_M GGUF
|
| 56 |
|
| 57 |
Optimized **Q4_K_M GGUF** quantization of **NetEase Youdao Confucius4-R2T2**, a low-latency, append-only streaming ASR model (fine-tuned from Qwen3-ASR-1.7B with Longest Stable Prefix decoding), ready for native execution in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp) and [audio.cpp](https://github.com/0xShug0/audio.cpp).
|
| 58 |
|
| 59 |
* **Original model:** [netease-youdao/Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2) Β· [source code](https://github.com/netease-youdao/Confucius4-R2T2) Β· [demo](https://r2t2.youdao.com/demo)
|
| 60 |
* **Base Architecture:** Qwen3-ASR audio encoder (24 layers, 1024 width, 2048 projected width) + Causal LM decoder (28 layers, 16 query heads, 8 KV heads, 128 head width).
|
| 61 |
+
* **Reference GGUF repo & inspiration:** [davidxifeng/Confucius4-R2T2-gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf) by David Xi Feng
|
| 62 |
+
* **Quantization & Benchmarking:** Researched and benchmarked in [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp).
|
| 63 |
|
| 64 |
---
|
| 65 |
|
| 66 |
+
## Technical Overview
|
| 67 |
|
| 68 |
+
Previous GGUF releases were constrained to Q8_0 or higher because uniform 4-bit and 5-bit quantizations either suffered from catastrophic empty-output cliffs or exhibited silent language drift (e.g. flipping to English when transcribing German or Russian).
|
|
|
|
| 69 |
|
| 70 |
+
This release introduces an **asymmetric per-tensor quantization recipe** that breaks through that barrier:
|
|
|
|
|
|
|
| 71 |
|
| 72 |
+
- **File:** `r2t2-q4_k_m.gguf`
|
| 73 |
+
- **File size:** **1.187 GB** (1,186,939,968 bytes) β **52% smaller** than the 2.478 GB Q8_0 reference.
|
| 74 |
+
- **Accuracy & Parity:** Preserves multilingual stability (verified across English, Chinese, German, and French).
|
| 75 |
+
- **Speed:** Faster than Q8_0 on CPU, avoiding the slow non-vectorized paths found in Q5_K.
|
| 76 |
|
| 77 |
+
### Quantization Composition
|
| 78 |
|
| 79 |
+
| Block | Target Dtype | Architectural Role |
|
| 80 |
+
|---|---|---|
|
| 81 |
+
| **Audio Tower** | `BF16 / Q4_K` | Audio encoder layers quantized to Q4_K with sensitive projection weights preserved at BF16. |
|
| 82 |
+
| **Attention Projections** | `Q4_K` | Fits within the LM error budget when down projections are held at higher precision. |
|
| 83 |
+
| **MLP Gate / Up** | `Q4_K` | Errors are bounded through SwiGLU activation. |
|
| 84 |
+
| **MLP Down** | **`Q6_K`** | **Hard Floor.** Writes into residual stream; lower precision triggers multilingual degradation. |
|
| 85 |
+
| **Embeddings** | **`Q2_K`** | Lookup table error does not compound; cuts 311M parameters with negligible loss. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
|
| 87 |
---
|
| 88 |
|
| 89 |
+
## Comprehensive Quantization Report
|
|
|
|
|
|
|
| 90 |
|
| 91 |
+
For the full ablation study, complete empirical test tables across 20+ quantization arms, failure mode analyses, speed benchmarks, and conversion reproduction instructions:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
|
| 93 |
+
π **Read the full report in [QUANTIZATION.md](QUANTIZATION.md)**
|
| 94 |
|
| 95 |
---
|
| 96 |
|
|
|
|
| 98 |
|
| 99 |
| File | Quantization | Size | Description |
|
| 100 |
|---|---|---|---|
|
| 101 |
+
| `r2t2-q4_k_m.gguf` | Q4_K_M | 1.187 GB | High-performance floor quantization; passes multilingual validation |
|
| 102 |
+
| `QUANTIZATION.md` | β | β | Detailed technical report, ablation study, and benchmarks |
|
| 103 |
+
| `NOTICE` | β | β | Attribution and derivative work disclaimers |
|
| 104 |
+
| `LICENSE` / `LICENSE_zh` | β | β | NetEase Youdao Model Use License Agreement |
|
| 105 |
|
| 106 |
---
|
| 107 |
|
|
|
|
| 124 |
### Server (OpenAI-compatible)
|
| 125 |
```bash
|
| 126 |
curl http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F file=@speech.wav
|
| 127 |
+
curl -N http://127.0.0.1:8488/v1/audio/transcriptions -F model=r2t2-asr -F stream=true -F file=@speech.wav
|
| 128 |
```
|
| 129 |
|
| 130 |
+
Streaming chunk sizes from **80 ms to 2000 ms** are supported (320 ms recommended default).
|
|
|
|
| 131 |
|
| 132 |
---
|
| 133 |
|