Upload PORTING.md with huggingface_hub
Browse files- PORTING.md +163 -0
PORTING.md
ADDED
|
@@ -0,0 +1,163 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Confucius4-R2T2 — Porting & Parity Specification
|
| 2 |
+
|
| 3 |
+
> Companion technical document for [Confucius4-R2T2-Q4_K_M-GGUF](https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF) covering architecture, loader adaptation, streaming state machine, and parity verification.
|
| 4 |
+
|
| 5 |
+
# Confucius4-R2T2
|
| 6 |
+
|
| 7 |
+
Status: loader adaptation implemented and compiling; streaming and application
|
| 8 |
+
integration are not implemented yet.
|
| 9 |
+
|
| 10 |
+
## Current implementation findings — 2026-09-20
|
| 11 |
+
|
| 12 |
+
The user requires a **native transcribe.cpp port**, usable by the application;
|
| 13 |
+
an audio.cpp executable, server, DLL or alternate runtime is not acceptable.
|
| 14 |
+
Use the sibling code as implementation reference only. Inference must use
|
| 15 |
+
transcribe's GGML backend selection, model/session ownership and streaming API.
|
| 16 |
+
|
| 17 |
+
The network matches the existing Qwen3-ASR audio encoder and causal decoder:
|
| 18 |
+
24 encoder layers, 1024 encoder width, 2048 projected audio width, 28 decoder
|
| 19 |
+
layers, 16 query heads, 8 KV heads and 128 head width. The pinned publisher
|
| 20 |
+
configuration uses `thinker_config` and declares BF16. Existing native Qwen
|
| 21 |
+
graphs are the reuse target; R2T2's additional work is the streaming state
|
| 22 |
+
machine, prompt-prefix handling and text post-processing, not a new backend.
|
| 23 |
+
|
| 24 |
+
The published Q8 file actually declares `general.architecture=audiocpp`, with
|
| 25 |
+
`audiocpp.model_spec.family=confucius4_r2t2`. It contains 707 tensors with native
|
| 26 |
+
publisher names and embedded configuration/tokenizer files. This describes
|
| 27 |
+
its **file format**, not a runtime we intend to use. The existing transcribe
|
| 28 |
+
loader cannot consume that schema unchanged. A native, family-specific loader
|
| 29 |
+
adapter or a transcribe-format conversion is required before using the verified
|
| 30 |
+
URL in the application. No external audio.cpp inference dependency is planned.
|
| 31 |
+
|
| 32 |
+
Initial source preparation now exists: the R2T2 text post-processing code was
|
| 33 |
+
adapted with Apache attribution, tensor-name mappings were generated from the
|
| 34 |
+
existing Qwen converter contract, and an 80–2000 ms stream-extension structure
|
| 35 |
+
was drafted. These are scaffolding, not connected or validated model support.
|
| 36 |
+
They have not yet passed a build as part of the model runtime.
|
| 37 |
+
|
| 38 |
+
The compatible audio.cpp graph-optimizer subset is already in
|
| 39 |
+
`src/transcribe-graph-opt.h` and used before scheduler allocation by the Qwen,
|
| 40 |
+
Granite and Parakeet paths. It preserves outputs, views and writable storage.
|
| 41 |
+
`TRANSCRIBE_GRAPH_OPTIMIZER=1` enables it. Earlier measurements did not establish
|
| 42 |
+
a universal gain, so importing it does not justify enabling it unconditionally.
|
| 43 |
+
The sibling's capacity-bucketed encoder/prefill cache is a separate optimization
|
| 44 |
+
and has **not** been ported. It requires valid-prefix masks and KV reset checks.
|
| 45 |
+
|
| 46 |
+
Remaining implementation: native package adaptation/conversion; streaming
|
| 47 |
+
begin/feed/finalize/reset; token rollback and UTF-8-safe stable output; exact
|
| 48 |
+
chunk-duration propagation through C/Rust and both application streaming paths;
|
| 49 |
+
catalogue entry with verified artifact identity; numeric slider/settings;
|
| 50 |
+
CPU/CUDA transcript and tail/reset validation. No claim of a working R2T2
|
| 51 |
+
download, native stream, slider, or model speedup is made at this checkpoint.
|
| 52 |
+
|
| 53 |
+
The four earlier task commit descriptions were expanded locally while preserving
|
| 54 |
+
their trees and parents. The history rewrite is not pushed yet; application
|
| 55 |
+
dependency pins must be updated to the final engine commit before publication.
|
| 56 |
+
|
| 57 |
+
## Identity and references
|
| 58 |
+
|
| 59 |
+
- Family: `confucius4_r2t2`; variant: `confucius4-r2t2-1.7b`.
|
| 60 |
+
- Publisher: [NetEase Youdao](https://huggingface.co/netease-youdao/Confucius4-R2T2), checkpoint `185ce639118ad1362d049ca0d8ed04b6ec5cd6c9`.
|
| 61 |
+
- Reference dtype: BF16, verified from 707 safetensors tensor headers.
|
| 62 |
+
- Weight license: NetEase Model Use License Agreement, not Qwen's Apache license.
|
| 63 |
+
- Canonical streaming reference: [publisher implementation](https://github.com/netease-youdao/Confucius4-R2T2/blob/80c22e6140bcb9166fb9906798894fc8b18c8309/r2t2/r2t2_asr.py).
|
| 64 |
+
- Secondary implementation: [audio.cpp a7b58a6](https://github.com/0xShug0/audio.cpp/commit/a7b58a6d3d6ae4143c485266b1c6c09898ad8c72).
|
| 65 |
+
- Acceptance dataset: LibriSpeech test-clean; measured reference WER pending.
|
| 66 |
+
|
| 67 |
+
## Capability validation
|
| 68 |
+
|
| 69 |
+
| Capability | Target | Status |
|
| 70 |
+
|---|---|---|
|
| 71 |
+
| Explicit language transcription | MUST PASS | TODO |
|
| 72 |
+
| Auto/no-hint transcription | MUST PASS | TODO |
|
| 73 |
+
| Offline batch | MUST PASS | TODO |
|
| 74 |
+
| Native streaming, stable deltas and authoritative final text | MUST PASS | TODO |
|
| 75 |
+
| Chunk sizes 80 through 2000 ms, including both endpoints | MUST PASS — user request | TODO |
|
| 76 |
+
| Context and hotwords | OUT OF SCOPE — user deferred on 2026-09-20 to prioritize streaming | TODO |
|
| 77 |
+
|
| 78 |
+
## Download identity
|
| 79 |
+
|
| 80 |
+
The secondary implementation publishes a standalone Q8 file at:
|
| 81 |
+
|
| 82 |
+
[r2t2-q8_0.gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf/resolve/a8e6b385d7df7eae9519363e07034a209004797a/r2t2-q8_0.gguf)
|
| 83 |
+
|
| 84 |
+
Verified HTTP HEAD 200 on 2026-09-20. Content length: 2,477,512,064 bytes.
|
| 85 |
+
Hugging Face LFS SHA-256:
|
| 86 |
+
`19f5ccd624484bcb5d44301437de41560b0ecc40c430e8850dfeefefbe82ccf5`.
|
| 87 |
+
The downloaded file was subsequently SHA-256 verified locally on 2026-09-20;
|
| 88 |
+
its hash matches the LFS digest above. It is retained under ignored
|
| 89 |
+
`build/diagnostics/r2t2-reference/published` in the engine workspace.
|
| 90 |
+
|
| 91 |
+
The same revision has `r2t2-f16.gguf` (4,092,155,264 bytes), SHA-256
|
| 92 |
+
`d1b531ceaf5640d98352d3a9180238d99d36d393e160afd4692031077e7bae2c`.
|
| 93 |
+
These files use audio.cpp packaging. Do not expose them as working transcribe
|
| 94 |
+
downloads until metadata/tensor adaptation and actual inference are validated.
|
| 95 |
+
An ordinary Qwen catalogue alias does not implement R2T2 streaming semantics.
|
| 96 |
+
|
| 97 |
+
## Streaming control contract
|
| 98 |
+
|
| 99 |
+
Expose a per-model integer chunk-size control from **80 to 2000 ms**, with a
|
| 100 |
+
1 ms step and direct numeric entry. Use 320 ms initially, matching the sibling
|
| 101 |
+
implementation; make 80 ms directly selectable. Persist the exact value and
|
| 102 |
+
validate it again in the native extension. Do not map this model through the
|
| 103 |
+
existing four Parakeet/Nemotron presets or silently round to those presets.
|
| 104 |
+
At 16 kHz, an integer millisecond corresponds to exactly 16 samples.
|
| 105 |
+
|
| 106 |
+
Label it **Streaming chunk size**. Chunk duration is not guaranteed end-to-end
|
| 107 |
+
latency: include queue wait, first committed text, per-feed p95/max and final
|
| 108 |
+
flush timings in diagnostics. Record the requested and resolved chunk duration.
|
| 109 |
+
Changing a setting applies to the next stream, not halfway through an active
|
| 110 |
+
decoder state. The capture path must forward small frames without waiting for
|
| 111 |
+
VAD silence before a native streaming dispatch.
|
| 112 |
+
|
| 113 |
+
Benchmark CPU and CUDA with exactly three runs per loaded configuration;
|
| 114 |
+
discard the first, retain raw results, and average runs two and three. Exercise
|
| 115 |
+
80/160/320/640/1280/2000 ms plus an irregular value to catch hidden quantization.
|
| 116 |
+
|
| 117 |
+
## Porting risks and remaining work
|
| 118 |
+
|
| 119 |
+
The publisher's pinned streaming reference explicitly requires vLLM. The
|
| 120 |
+
sibling implementation's MPS golden-generation script refers to a different
|
| 121 |
+
reference environment, so it cannot establish our Windows oracle by itself.
|
| 122 |
+
Establish a reproducible reference before claiming parity.
|
| 123 |
+
|
| 124 |
+
Preserve token rollback, UTF-8 boundaries, punctuation/repetition repair,
|
| 125 |
+
language-tag handling, pipe truncation, final tail and reset behavior. The
|
| 126 |
+
model re-encodes accumulated audio; capacity-bucketed encoder and prefill graph
|
| 127 |
+
reuse must refill masks and clear stale KV. Padding may change floating-point
|
| 128 |
+
reduction order; check both numerics and transcript behavior across bucket
|
| 129 |
+
growth and shrink. Context/hotword UI is intentionally deferred.
|
| 130 |
+
|
| 131 |
+
## Commands and artifacts
|
| 132 |
+
|
| 133 |
+
```powershell
|
| 134 |
+
uv run scripts/intake.py inspect --repo netease-youdao/Confucius4-R2T2 --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --out reports/porting/confucius4_r2t2/confucius4-r2t2-1.7b/intake.json
|
| 135 |
+
uv run scripts/preflight.py --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --gate A
|
| 136 |
+
```
|
| 137 |
+
|
| 138 |
+
Initial Gate A: WARN, not a numerical pass. Tokenizer alignment passes;
|
| 139 |
+
dtype declaration lookup misses nested `thinker_config.dtype`, and missing
|
| 140 |
+
frontend normalization metadata is interpreted as `none` by preflight even
|
| 141 |
+
though Whisper applies log-mel clamping/scaling. GGUF capability checks await
|
| 142 |
+
conversion. The golden manifest remains explicitly marked as a skeleton.
|
| 143 |
+
|
| 144 |
+
## Quantization
|
| 145 |
+
|
| 146 |
+
See `docs/porting/families/confucius4_r2t2-quantization.md` for the measured
|
| 147 |
+
precision ladder: which blocks tolerate which type, the 1.187 GB floor
|
| 148 |
+
(`r2t2-q4_k_m.gguf`, 52% below Q8_0), the Q6_K floor on `mlp.down_proj`, and the
|
| 149 |
+
build recipe. `docs/tools/quantization-arms.md` covers the per-tensor method
|
| 150 |
+
generally.
|
| 151 |
+
|
| 152 |
+
## Local reference checkpoint
|
| 153 |
+
|
| 154 |
+
The pinned publisher safetensors checkpoint has been downloaded under ignored
|
| 155 |
+
`build/diagnostics/r2t2-reference/checkpoint`. The existing author-Qwen dumper
|
| 156 |
+
successfully ran BF16 CPU inference with this checkpoint using Transformers
|
| 157 |
+
4.57.6 / qwen-asr 0.0.6 / Torch 2.11.0+cpu. It produced 13 tensor dumps and the
|
| 158 |
+
expected JFK transcription (29 generated tokens) under
|
| 159 |
+
`build/validate/confucius4_r2t2/confucius4-r2t2-1.7b/jfk/decode/ref`.
|
| 160 |
+
This establishes offline network bring-up only: it is not a streaming oracle,
|
| 161 |
+
a WER result, or a C++ parity pass. The existing dumper lacks the newer RMS/p99
|
| 162 |
+
sidecar fields, which must be added to the dedicated adapter before completing
|
| 163 |
+
Stage 2. No supported-model claim follows from this result.
|