Nairod785 commited on
Commit
3e68648
·
verified ·
1 Parent(s): 0516975

Upload PORTING.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. PORTING.md +163 -0
PORTING.md ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Confucius4-R2T2 — Porting & Parity Specification
2
+
3
+ > Companion technical document for [Confucius4-R2T2-Q4_K_M-GGUF](https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF) covering architecture, loader adaptation, streaming state machine, and parity verification.
4
+
5
+ # Confucius4-R2T2
6
+
7
+ Status: loader adaptation implemented and compiling; streaming and application
8
+ integration are not implemented yet.
9
+
10
+ ## Current implementation findings — 2026-09-20
11
+
12
+ The user requires a **native transcribe.cpp port**, usable by the application;
13
+ an audio.cpp executable, server, DLL or alternate runtime is not acceptable.
14
+ Use the sibling code as implementation reference only. Inference must use
15
+ transcribe's GGML backend selection, model/session ownership and streaming API.
16
+
17
+ The network matches the existing Qwen3-ASR audio encoder and causal decoder:
18
+ 24 encoder layers, 1024 encoder width, 2048 projected audio width, 28 decoder
19
+ layers, 16 query heads, 8 KV heads and 128 head width. The pinned publisher
20
+ configuration uses `thinker_config` and declares BF16. Existing native Qwen
21
+ graphs are the reuse target; R2T2's additional work is the streaming state
22
+ machine, prompt-prefix handling and text post-processing, not a new backend.
23
+
24
+ The published Q8 file actually declares `general.architecture=audiocpp`, with
25
+ `audiocpp.model_spec.family=confucius4_r2t2`. It contains 707 tensors with native
26
+ publisher names and embedded configuration/tokenizer files. This describes
27
+ its **file format**, not a runtime we intend to use. The existing transcribe
28
+ loader cannot consume that schema unchanged. A native, family-specific loader
29
+ adapter or a transcribe-format conversion is required before using the verified
30
+ URL in the application. No external audio.cpp inference dependency is planned.
31
+
32
+ Initial source preparation now exists: the R2T2 text post-processing code was
33
+ adapted with Apache attribution, tensor-name mappings were generated from the
34
+ existing Qwen converter contract, and an 80–2000 ms stream-extension structure
35
+ was drafted. These are scaffolding, not connected or validated model support.
36
+ They have not yet passed a build as part of the model runtime.
37
+
38
+ The compatible audio.cpp graph-optimizer subset is already in
39
+ `src/transcribe-graph-opt.h` and used before scheduler allocation by the Qwen,
40
+ Granite and Parakeet paths. It preserves outputs, views and writable storage.
41
+ `TRANSCRIBE_GRAPH_OPTIMIZER=1` enables it. Earlier measurements did not establish
42
+ a universal gain, so importing it does not justify enabling it unconditionally.
43
+ The sibling's capacity-bucketed encoder/prefill cache is a separate optimization
44
+ and has **not** been ported. It requires valid-prefix masks and KV reset checks.
45
+
46
+ Remaining implementation: native package adaptation/conversion; streaming
47
+ begin/feed/finalize/reset; token rollback and UTF-8-safe stable output; exact
48
+ chunk-duration propagation through C/Rust and both application streaming paths;
49
+ catalogue entry with verified artifact identity; numeric slider/settings;
50
+ CPU/CUDA transcript and tail/reset validation. No claim of a working R2T2
51
+ download, native stream, slider, or model speedup is made at this checkpoint.
52
+
53
+ The four earlier task commit descriptions were expanded locally while preserving
54
+ their trees and parents. The history rewrite is not pushed yet; application
55
+ dependency pins must be updated to the final engine commit before publication.
56
+
57
+ ## Identity and references
58
+
59
+ - Family: `confucius4_r2t2`; variant: `confucius4-r2t2-1.7b`.
60
+ - Publisher: [NetEase Youdao](https://huggingface.co/netease-youdao/Confucius4-R2T2), checkpoint `185ce639118ad1362d049ca0d8ed04b6ec5cd6c9`.
61
+ - Reference dtype: BF16, verified from 707 safetensors tensor headers.
62
+ - Weight license: NetEase Model Use License Agreement, not Qwen's Apache license.
63
+ - Canonical streaming reference: [publisher implementation](https://github.com/netease-youdao/Confucius4-R2T2/blob/80c22e6140bcb9166fb9906798894fc8b18c8309/r2t2/r2t2_asr.py).
64
+ - Secondary implementation: [audio.cpp a7b58a6](https://github.com/0xShug0/audio.cpp/commit/a7b58a6d3d6ae4143c485266b1c6c09898ad8c72).
65
+ - Acceptance dataset: LibriSpeech test-clean; measured reference WER pending.
66
+
67
+ ## Capability validation
68
+
69
+ | Capability | Target | Status |
70
+ |---|---|---|
71
+ | Explicit language transcription | MUST PASS | TODO |
72
+ | Auto/no-hint transcription | MUST PASS | TODO |
73
+ | Offline batch | MUST PASS | TODO |
74
+ | Native streaming, stable deltas and authoritative final text | MUST PASS | TODO |
75
+ | Chunk sizes 80 through 2000 ms, including both endpoints | MUST PASS — user request | TODO |
76
+ | Context and hotwords | OUT OF SCOPE — user deferred on 2026-09-20 to prioritize streaming | TODO |
77
+
78
+ ## Download identity
79
+
80
+ The secondary implementation publishes a standalone Q8 file at:
81
+
82
+ [r2t2-q8_0.gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf/resolve/a8e6b385d7df7eae9519363e07034a209004797a/r2t2-q8_0.gguf)
83
+
84
+ Verified HTTP HEAD 200 on 2026-09-20. Content length: 2,477,512,064 bytes.
85
+ Hugging Face LFS SHA-256:
86
+ `19f5ccd624484bcb5d44301437de41560b0ecc40c430e8850dfeefefbe82ccf5`.
87
+ The downloaded file was subsequently SHA-256 verified locally on 2026-09-20;
88
+ its hash matches the LFS digest above. It is retained under ignored
89
+ `build/diagnostics/r2t2-reference/published` in the engine workspace.
90
+
91
+ The same revision has `r2t2-f16.gguf` (4,092,155,264 bytes), SHA-256
92
+ `d1b531ceaf5640d98352d3a9180238d99d36d393e160afd4692031077e7bae2c`.
93
+ These files use audio.cpp packaging. Do not expose them as working transcribe
94
+ downloads until metadata/tensor adaptation and actual inference are validated.
95
+ An ordinary Qwen catalogue alias does not implement R2T2 streaming semantics.
96
+
97
+ ## Streaming control contract
98
+
99
+ Expose a per-model integer chunk-size control from **80 to 2000 ms**, with a
100
+ 1 ms step and direct numeric entry. Use 320 ms initially, matching the sibling
101
+ implementation; make 80 ms directly selectable. Persist the exact value and
102
+ validate it again in the native extension. Do not map this model through the
103
+ existing four Parakeet/Nemotron presets or silently round to those presets.
104
+ At 16 kHz, an integer millisecond corresponds to exactly 16 samples.
105
+
106
+ Label it **Streaming chunk size**. Chunk duration is not guaranteed end-to-end
107
+ latency: include queue wait, first committed text, per-feed p95/max and final
108
+ flush timings in diagnostics. Record the requested and resolved chunk duration.
109
+ Changing a setting applies to the next stream, not halfway through an active
110
+ decoder state. The capture path must forward small frames without waiting for
111
+ VAD silence before a native streaming dispatch.
112
+
113
+ Benchmark CPU and CUDA with exactly three runs per loaded configuration;
114
+ discard the first, retain raw results, and average runs two and three. Exercise
115
+ 80/160/320/640/1280/2000 ms plus an irregular value to catch hidden quantization.
116
+
117
+ ## Porting risks and remaining work
118
+
119
+ The publisher's pinned streaming reference explicitly requires vLLM. The
120
+ sibling implementation's MPS golden-generation script refers to a different
121
+ reference environment, so it cannot establish our Windows oracle by itself.
122
+ Establish a reproducible reference before claiming parity.
123
+
124
+ Preserve token rollback, UTF-8 boundaries, punctuation/repetition repair,
125
+ language-tag handling, pipe truncation, final tail and reset behavior. The
126
+ model re-encodes accumulated audio; capacity-bucketed encoder and prefill graph
127
+ reuse must refill masks and clear stale KV. Padding may change floating-point
128
+ reduction order; check both numerics and transcript behavior across bucket
129
+ growth and shrink. Context/hotword UI is intentionally deferred.
130
+
131
+ ## Commands and artifacts
132
+
133
+ ```powershell
134
+ uv run scripts/intake.py inspect --repo netease-youdao/Confucius4-R2T2 --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --out reports/porting/confucius4_r2t2/confucius4-r2t2-1.7b/intake.json
135
+ uv run scripts/preflight.py --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --gate A
136
+ ```
137
+
138
+ Initial Gate A: WARN, not a numerical pass. Tokenizer alignment passes;
139
+ dtype declaration lookup misses nested `thinker_config.dtype`, and missing
140
+ frontend normalization metadata is interpreted as `none` by preflight even
141
+ though Whisper applies log-mel clamping/scaling. GGUF capability checks await
142
+ conversion. The golden manifest remains explicitly marked as a skeleton.
143
+
144
+ ## Quantization
145
+
146
+ See `docs/porting/families/confucius4_r2t2-quantization.md` for the measured
147
+ precision ladder: which blocks tolerate which type, the 1.187 GB floor
148
+ (`r2t2-q4_k_m.gguf`, 52% below Q8_0), the Q6_K floor on `mlp.down_proj`, and the
149
+ build recipe. `docs/tools/quantization-arms.md` covers the per-tensor method
150
+ generally.
151
+
152
+ ## Local reference checkpoint
153
+
154
+ The pinned publisher safetensors checkpoint has been downloaded under ignored
155
+ `build/diagnostics/r2t2-reference/checkpoint`. The existing author-Qwen dumper
156
+ successfully ran BF16 CPU inference with this checkpoint using Transformers
157
+ 4.57.6 / qwen-asr 0.0.6 / Torch 2.11.0+cpu. It produced 13 tensor dumps and the
158
+ expected JFK transcription (29 generated tokens) under
159
+ `build/validate/confucius4_r2t2/confucius4-r2t2-1.7b/jfk/decode/ref`.
160
+ This establishes offline network bring-up only: it is not a streaming oracle,
161
+ a WER result, or a C++ parity pass. The existing dumper lacks the newer RMS/p99
162
+ sidecar fields, which must be added to the dedicated adapter before completing
163
+ Stage 2. No supported-model claim follows from this result.