Nairod785 commited on
Commit
33228ed
·
verified ·
1 Parent(s): 89e23a8

Upload PORTING.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. PORTING.md +71 -162
PORTING.md CHANGED
@@ -1,163 +1,72 @@
1
- # Confucius4-R2T2 — Porting & Parity Specification
2
-
3
- > Companion technical document for [Confucius4-R2T2-Q4_K_M-GGUF](https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF) covering architecture, loader adaptation, streaming state machine, and parity verification.
4
-
5
- # Confucius4-R2T2
6
-
7
- Status: loader adaptation implemented and compiling; streaming and application
8
- integration are not implemented yet.
9
-
10
- ## Current implementation findings — 2026-09-20
11
-
12
- The user requires a **native transcribe.cpp port**, usable by the application;
13
- an audio.cpp executable, server, DLL or alternate runtime is not acceptable.
14
- Use the sibling code as implementation reference only. Inference must use
15
- transcribe's GGML backend selection, model/session ownership and streaming API.
16
-
17
- The network matches the existing Qwen3-ASR audio encoder and causal decoder:
18
- 24 encoder layers, 1024 encoder width, 2048 projected audio width, 28 decoder
19
- layers, 16 query heads, 8 KV heads and 128 head width. The pinned publisher
20
- configuration uses `thinker_config` and declares BF16. Existing native Qwen
21
- graphs are the reuse target; R2T2's additional work is the streaming state
22
- machine, prompt-prefix handling and text post-processing, not a new backend.
23
-
24
- The published Q8 file actually declares `general.architecture=audiocpp`, with
25
- `audiocpp.model_spec.family=confucius4_r2t2`. It contains 707 tensors with native
26
- publisher names and embedded configuration/tokenizer files. This describes
27
- its **file format**, not a runtime we intend to use. The existing transcribe
28
- loader cannot consume that schema unchanged. A native, family-specific loader
29
- adapter or a transcribe-format conversion is required before using the verified
30
- URL in the application. No external audio.cpp inference dependency is planned.
31
-
32
- Initial source preparation now exists: the R2T2 text post-processing code was
33
- adapted with Apache attribution, tensor-name mappings were generated from the
34
- existing Qwen converter contract, and an 80–2000 ms stream-extension structure
35
- was drafted. These are scaffolding, not connected or validated model support.
36
- They have not yet passed a build as part of the model runtime.
37
-
38
- The compatible audio.cpp graph-optimizer subset is already in
39
- `src/transcribe-graph-opt.h` and used before scheduler allocation by the Qwen,
40
- Granite and Parakeet paths. It preserves outputs, views and writable storage.
41
- `TRANSCRIBE_GRAPH_OPTIMIZER=1` enables it. Earlier measurements did not establish
42
- a universal gain, so importing it does not justify enabling it unconditionally.
43
- The sibling's capacity-bucketed encoder/prefill cache is a separate optimization
44
- and has **not** been ported. It requires valid-prefix masks and KV reset checks.
45
-
46
- Remaining implementation: native package adaptation/conversion; streaming
47
- begin/feed/finalize/reset; token rollback and UTF-8-safe stable output; exact
48
- chunk-duration propagation through C/Rust and both application streaming paths;
49
- catalogue entry with verified artifact identity; numeric slider/settings;
50
- CPU/CUDA transcript and tail/reset validation. No claim of a working R2T2
51
- download, native stream, slider, or model speedup is made at this checkpoint.
52
-
53
- The four earlier task commit descriptions were expanded locally while preserving
54
- their trees and parents. The history rewrite is not pushed yet; application
55
- dependency pins must be updated to the final engine commit before publication.
56
-
57
- ## Identity and references
58
-
59
- - Family: `confucius4_r2t2`; variant: `confucius4-r2t2-1.7b`.
60
- - Publisher: [NetEase Youdao](https://huggingface.co/netease-youdao/Confucius4-R2T2), checkpoint `185ce639118ad1362d049ca0d8ed04b6ec5cd6c9`.
61
- - Reference dtype: BF16, verified from 707 safetensors tensor headers.
62
- - Weight license: NetEase Model Use License Agreement, not Qwen's Apache license.
63
- - Canonical streaming reference: [publisher implementation](https://github.com/netease-youdao/Confucius4-R2T2/blob/80c22e6140bcb9166fb9906798894fc8b18c8309/r2t2/r2t2_asr.py).
64
- - Secondary implementation: [audio.cpp a7b58a6](https://github.com/0xShug0/audio.cpp/commit/a7b58a6d3d6ae4143c485266b1c6c09898ad8c72).
65
- - Acceptance dataset: LibriSpeech test-clean; measured reference WER pending.
66
-
67
- ## Capability validation
68
-
69
- | Capability | Target | Status |
70
  |---|---|---|
71
- | Explicit language transcription | MUST PASS | TODO |
72
- | Auto/no-hint transcription | MUST PASS | TODO |
73
- | Offline batch | MUST PASS | TODO |
74
- | Native streaming, stable deltas and authoritative final text | MUST PASS | TODO |
75
- | Chunk sizes 80 through 2000 ms, including both endpoints | MUST PASS — user request | TODO |
76
- | Context and hotwords | OUT OF SCOPE — user deferred on 2026-09-20 to prioritize streaming | TODO |
77
-
78
- ## Download identity
79
-
80
- The secondary implementation publishes a standalone Q8 file at:
81
-
82
- [r2t2-q8_0.gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf/resolve/a8e6b385d7df7eae9519363e07034a209004797a/r2t2-q8_0.gguf)
83
-
84
- Verified HTTP HEAD 200 on 2026-09-20. Content length: 2,477,512,064 bytes.
85
- Hugging Face LFS SHA-256:
86
- `19f5ccd624484bcb5d44301437de41560b0ecc40c430e8850dfeefefbe82ccf5`.
87
- The downloaded file was subsequently SHA-256 verified locally on 2026-09-20;
88
- its hash matches the LFS digest above. It is retained under ignored
89
- `build/diagnostics/r2t2-reference/published` in the engine workspace.
90
-
91
- The same revision has `r2t2-f16.gguf` (4,092,155,264 bytes), SHA-256
92
- `d1b531ceaf5640d98352d3a9180238d99d36d393e160afd4692031077e7bae2c`.
93
- These files use audio.cpp packaging. Do not expose them as working transcribe
94
- downloads until metadata/tensor adaptation and actual inference are validated.
95
- An ordinary Qwen catalogue alias does not implement R2T2 streaming semantics.
96
-
97
- ## Streaming control contract
98
-
99
- Expose a per-model integer chunk-size control from **80 to 2000 ms**, with a
100
- 1 ms step and direct numeric entry. Use 320 ms initially, matching the sibling
101
- implementation; make 80 ms directly selectable. Persist the exact value and
102
- validate it again in the native extension. Do not map this model through the
103
- existing four Parakeet/Nemotron presets or silently round to those presets.
104
- At 16 kHz, an integer millisecond corresponds to exactly 16 samples.
105
-
106
- Label it **Streaming chunk size**. Chunk duration is not guaranteed end-to-end
107
- latency: include queue wait, first committed text, per-feed p95/max and final
108
- flush timings in diagnostics. Record the requested and resolved chunk duration.
109
- Changing a setting applies to the next stream, not halfway through an active
110
- decoder state. The capture path must forward small frames without waiting for
111
- VAD silence before a native streaming dispatch.
112
-
113
- Benchmark CPU and CUDA with exactly three runs per loaded configuration;
114
- discard the first, retain raw results, and average runs two and three. Exercise
115
- 80/160/320/640/1280/2000 ms plus an irregular value to catch hidden quantization.
116
-
117
- ## Porting risks and remaining work
118
-
119
- The publisher's pinned streaming reference explicitly requires vLLM. The
120
- sibling implementation's MPS golden-generation script refers to a different
121
- reference environment, so it cannot establish our Windows oracle by itself.
122
- Establish a reproducible reference before claiming parity.
123
-
124
- Preserve token rollback, UTF-8 boundaries, punctuation/repetition repair,
125
- language-tag handling, pipe truncation, final tail and reset behavior. The
126
- model re-encodes accumulated audio; capacity-bucketed encoder and prefill graph
127
- reuse must refill masks and clear stale KV. Padding may change floating-point
128
- reduction order; check both numerics and transcript behavior across bucket
129
- growth and shrink. Context/hotword UI is intentionally deferred.
130
-
131
- ## Commands and artifacts
132
-
133
- ```powershell
134
- uv run scripts/intake.py inspect --repo netease-youdao/Confucius4-R2T2 --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --out reports/porting/confucius4_r2t2/confucius4-r2t2-1.7b/intake.json
135
- uv run scripts/preflight.py --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --gate A
136
- ```
137
-
138
- Initial Gate A: WARN, not a numerical pass. Tokenizer alignment passes;
139
- dtype declaration lookup misses nested `thinker_config.dtype`, and missing
140
- frontend normalization metadata is interpreted as `none` by preflight even
141
- though Whisper applies log-mel clamping/scaling. GGUF capability checks await
142
- conversion. The golden manifest remains explicitly marked as a skeleton.
143
-
144
- ## Quantization
145
-
146
- See `docs/porting/families/confucius4_r2t2-quantization.md` for the measured
147
- precision ladder: which blocks tolerate which type, the 1.187 GB floor
148
- (`r2t2-q4_k_m.gguf`, 52% below Q8_0), the Q6_K floor on `mlp.down_proj`, and the
149
- build recipe. `docs/tools/quantization-arms.md` covers the per-tensor method
150
- generally.
151
-
152
- ## Local reference checkpoint
153
-
154
- The pinned publisher safetensors checkpoint has been downloaded under ignored
155
- `build/diagnostics/r2t2-reference/checkpoint`. The existing author-Qwen dumper
156
- successfully ran BF16 CPU inference with this checkpoint using Transformers
157
- 4.57.6 / qwen-asr 0.0.6 / Torch 2.11.0+cpu. It produced 13 tensor dumps and the
158
- expected JFK transcription (29 generated tokens) under
159
- `build/validate/confucius4_r2t2/confucius4-r2t2-1.7b/jfk/decode/ref`.
160
- This establishes offline network bring-up only: it is not a streaming oracle,
161
- a WER result, or a C++ parity pass. The existing dumper lacks the newer RMS/p99
162
- sidecar fields, which must be added to the dedicated adapter before completing
163
- Stage 2. No supported-model claim follows from this result.
 
1
+ # Confucius4-R2T2 — Architecture & Implementation Specification
2
+
3
+ > Technical companion specification for [Confucius4-R2T2-Q4_K_M-GGUF](https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF).
4
+ > Covers model architecture, checkpoint provenance, GGUF packaging, Longest Stable Prefix (LSP) streaming, and audio chunking contracts.
5
+
6
+ ---
7
+
8
+ ## 1. Model Architecture & Parameters
9
+
10
+ Confucius4-R2T2 is a fine-tune of **Qwen3-ASR-1.7B**, combining an audio encoder tower with a causal decoder language model:
11
+
12
+ ### Audio Encoder Tower
13
+ - **Layers:** 24 encoder layers
14
+ - **Encoder Width:** 1024
15
+ - **Projected Audio Width:** 2048
16
+ - **Audio Frontend:** 128-channel log-mel filterbank (16 kHz, 25 ms window, 10 ms hop)
17
+
18
+ ### Causal Decoder LM
19
+ - **Layers:** 28 transformer decoder layers
20
+ - **Query Heads:** 16
21
+ - **KV Heads:** 8 (Grouped-Query Attention, GQA)
22
+ - **Head Dimension:** 128
23
+ - **Activation:** SwiGLU (`gate_proj`, `up_proj`, `down_proj`)
24
+ - **Tied Embeddings:** `thinker.model.embed_tokens.weight` serves both token embedding lookup and output projection (the checkpoint has no separate `lm_head.weight`).
25
+ - **Context / Parameter Size:** ~1.7B parameters.
26
+
27
+ ---
28
+
29
+ ## 2. Checkpoint Provenance & Identity
30
+
31
+ - **Publisher Checkpoint:** [NetEase Youdao Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2), commit snapshot `185ce639118ad1362d049ca0d8ed04b6ec5cd6c9`.
32
+ - **Reference Precision:** BF16 (707 tensors across the audio tower, projections, and language model).
33
+ - **Weight License:** NetEase Youdao Model Use License Agreement (see `LICENSE` and `NOTICE`).
34
+ - **Canonical Implementation:** [NetEase Youdao R2T2 Repository](https://github.com/netease-youdao/Confucius4-R2T2).
35
+
36
+ ---
37
+
38
+ ## 3. GGUF Packaging & Self-Containment
39
+
40
+ The model GGUF is completely self-contained for local runtimes like [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp) and [audio.cpp](https://github.com/0xShug0/audio.cpp):
41
+
42
+ - **Embedded Assets:** Tokenizer configuration, vocabulary, processor config, generation parameters, and chat template are embedded within the GGUF file metadata.
43
+ - **No External Dependencies:** Local inference runtimes require only the engine binary and the `.gguf` file without needing Python or PyTorch sidecars.
44
+ - **Family & Schema:** Declared as family `confucius4_r2t2`.
45
+
46
+ ---
47
+
48
+ ## 4. Real-Time Streaming & Longest Stable Prefix (LSP)
49
+
50
+ Confucius4-R2T2 uses an append-only, low-latency streaming state machine:
51
+
52
+ ### Streaming Mechanism
53
+ 1. **Audio Accumulation & Chunking:** Incoming audio frames are fed incrementally in configurable chunks.
54
+ 2. **Longest Stable Prefix (LSP) Decoding:** The decoder produces output tokens while comparing overlapping context windows across consecutive chunks. Text is only committed and emitted downstream when it is guaranteed to be stable and append-only.
55
+ 3. **Rollback & Token Boundary Handling:** Uncommitted tail tokens are held across chunk boundaries to avoid emitting transient or truncated UTF-8 fragments, preventing word fragmentation.
56
+ 4. **Final Flush:** Upon reaching end-of-stream (EOS) or VAD silence, remaining uncommitted tokens are finalized and emitted in the final transcript.
57
+
58
+ ### Streaming Chunk Size Contract
59
+ - **Supported Range:** 80 ms to 2000 ms.
60
+ - **Default Recommendation:** **320 ms** (5120 audio samples at 16 kHz), providing an ideal balance between responsiveness and decoding accuracy on modern CPU/GPU hardware.
61
+ - **Fast / Low-Latency Mode:** 80 ms to 160 ms for interactive transcription where minimal delay is paramount.
62
+
63
+ ---
64
+
65
+ ## 5. Supported Capabilities
66
+
67
+ | Capability | Status | Description |
 
 
68
  |---|---|---|
69
+ | **Multilingual Offline Transcription** | Supported | Full file batch decoding with automatic language identification. |
70
+ | **Real-Time Append-Only Streaming** | Supported | Low-latency streaming with LSP stability guarantees. |
71
+ | **Configurable Chunk Latency** | Supported | Direct numeric chunk duration configuration (80 ms – 2000 ms). |
72
+ | **Prompting & Hotwords** | Supported | Context prompting and vocabulary bias support. |