Upload PORTING.md with huggingface_hub
Browse files- PORTING.md +71 -162
PORTING.md
CHANGED
|
@@ -1,163 +1,72 @@
|
|
| 1 |
-
# Confucius4-R2T2 —
|
| 2 |
-
|
| 3 |
-
>
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
-
|
| 60 |
-
-
|
| 61 |
-
-
|
| 62 |
-
|
| 63 |
-
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
| Capability | Target | Status |
|
| 70 |
|---|---|---|
|
| 71 |
-
|
|
| 72 |
-
|
|
| 73 |
-
|
|
| 74 |
-
|
|
| 75 |
-
| Chunk sizes 80 through 2000 ms, including both endpoints | MUST PASS — user request | TODO |
|
| 76 |
-
| Context and hotwords | OUT OF SCOPE — user deferred on 2026-09-20 to prioritize streaming | TODO |
|
| 77 |
-
|
| 78 |
-
## Download identity
|
| 79 |
-
|
| 80 |
-
The secondary implementation publishes a standalone Q8 file at:
|
| 81 |
-
|
| 82 |
-
[r2t2-q8_0.gguf](https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf/resolve/a8e6b385d7df7eae9519363e07034a209004797a/r2t2-q8_0.gguf)
|
| 83 |
-
|
| 84 |
-
Verified HTTP HEAD 200 on 2026-09-20. Content length: 2,477,512,064 bytes.
|
| 85 |
-
Hugging Face LFS SHA-256:
|
| 86 |
-
`19f5ccd624484bcb5d44301437de41560b0ecc40c430e8850dfeefefbe82ccf5`.
|
| 87 |
-
The downloaded file was subsequently SHA-256 verified locally on 2026-09-20;
|
| 88 |
-
its hash matches the LFS digest above. It is retained under ignored
|
| 89 |
-
`build/diagnostics/r2t2-reference/published` in the engine workspace.
|
| 90 |
-
|
| 91 |
-
The same revision has `r2t2-f16.gguf` (4,092,155,264 bytes), SHA-256
|
| 92 |
-
`d1b531ceaf5640d98352d3a9180238d99d36d393e160afd4692031077e7bae2c`.
|
| 93 |
-
These files use audio.cpp packaging. Do not expose them as working transcribe
|
| 94 |
-
downloads until metadata/tensor adaptation and actual inference are validated.
|
| 95 |
-
An ordinary Qwen catalogue alias does not implement R2T2 streaming semantics.
|
| 96 |
-
|
| 97 |
-
## Streaming control contract
|
| 98 |
-
|
| 99 |
-
Expose a per-model integer chunk-size control from **80 to 2000 ms**, with a
|
| 100 |
-
1 ms step and direct numeric entry. Use 320 ms initially, matching the sibling
|
| 101 |
-
implementation; make 80 ms directly selectable. Persist the exact value and
|
| 102 |
-
validate it again in the native extension. Do not map this model through the
|
| 103 |
-
existing four Parakeet/Nemotron presets or silently round to those presets.
|
| 104 |
-
At 16 kHz, an integer millisecond corresponds to exactly 16 samples.
|
| 105 |
-
|
| 106 |
-
Label it **Streaming chunk size**. Chunk duration is not guaranteed end-to-end
|
| 107 |
-
latency: include queue wait, first committed text, per-feed p95/max and final
|
| 108 |
-
flush timings in diagnostics. Record the requested and resolved chunk duration.
|
| 109 |
-
Changing a setting applies to the next stream, not halfway through an active
|
| 110 |
-
decoder state. The capture path must forward small frames without waiting for
|
| 111 |
-
VAD silence before a native streaming dispatch.
|
| 112 |
-
|
| 113 |
-
Benchmark CPU and CUDA with exactly three runs per loaded configuration;
|
| 114 |
-
discard the first, retain raw results, and average runs two and three. Exercise
|
| 115 |
-
80/160/320/640/1280/2000 ms plus an irregular value to catch hidden quantization.
|
| 116 |
-
|
| 117 |
-
## Porting risks and remaining work
|
| 118 |
-
|
| 119 |
-
The publisher's pinned streaming reference explicitly requires vLLM. The
|
| 120 |
-
sibling implementation's MPS golden-generation script refers to a different
|
| 121 |
-
reference environment, so it cannot establish our Windows oracle by itself.
|
| 122 |
-
Establish a reproducible reference before claiming parity.
|
| 123 |
-
|
| 124 |
-
Preserve token rollback, UTF-8 boundaries, punctuation/repetition repair,
|
| 125 |
-
language-tag handling, pipe truncation, final tail and reset behavior. The
|
| 126 |
-
model re-encodes accumulated audio; capacity-bucketed encoder and prefill graph
|
| 127 |
-
reuse must refill masks and clear stale KV. Padding may change floating-point
|
| 128 |
-
reduction order; check both numerics and transcript behavior across bucket
|
| 129 |
-
growth and shrink. Context/hotword UI is intentionally deferred.
|
| 130 |
-
|
| 131 |
-
## Commands and artifacts
|
| 132 |
-
|
| 133 |
-
```powershell
|
| 134 |
-
uv run scripts/intake.py inspect --repo netease-youdao/Confucius4-R2T2 --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --out reports/porting/confucius4_r2t2/confucius4-r2t2-1.7b/intake.json
|
| 135 |
-
uv run scripts/preflight.py --family confucius4_r2t2 --variant confucius4-r2t2-1.7b --gate A
|
| 136 |
-
```
|
| 137 |
-
|
| 138 |
-
Initial Gate A: WARN, not a numerical pass. Tokenizer alignment passes;
|
| 139 |
-
dtype declaration lookup misses nested `thinker_config.dtype`, and missing
|
| 140 |
-
frontend normalization metadata is interpreted as `none` by preflight even
|
| 141 |
-
though Whisper applies log-mel clamping/scaling. GGUF capability checks await
|
| 142 |
-
conversion. The golden manifest remains explicitly marked as a skeleton.
|
| 143 |
-
|
| 144 |
-
## Quantization
|
| 145 |
-
|
| 146 |
-
See `docs/porting/families/confucius4_r2t2-quantization.md` for the measured
|
| 147 |
-
precision ladder: which blocks tolerate which type, the 1.187 GB floor
|
| 148 |
-
(`r2t2-q4_k_m.gguf`, 52% below Q8_0), the Q6_K floor on `mlp.down_proj`, and the
|
| 149 |
-
build recipe. `docs/tools/quantization-arms.md` covers the per-tensor method
|
| 150 |
-
generally.
|
| 151 |
-
|
| 152 |
-
## Local reference checkpoint
|
| 153 |
-
|
| 154 |
-
The pinned publisher safetensors checkpoint has been downloaded under ignored
|
| 155 |
-
`build/diagnostics/r2t2-reference/checkpoint`. The existing author-Qwen dumper
|
| 156 |
-
successfully ran BF16 CPU inference with this checkpoint using Transformers
|
| 157 |
-
4.57.6 / qwen-asr 0.0.6 / Torch 2.11.0+cpu. It produced 13 tensor dumps and the
|
| 158 |
-
expected JFK transcription (29 generated tokens) under
|
| 159 |
-
`build/validate/confucius4_r2t2/confucius4-r2t2-1.7b/jfk/decode/ref`.
|
| 160 |
-
This establishes offline network bring-up only: it is not a streaming oracle,
|
| 161 |
-
a WER result, or a C++ parity pass. The existing dumper lacks the newer RMS/p99
|
| 162 |
-
sidecar fields, which must be added to the dedicated adapter before completing
|
| 163 |
-
Stage 2. No supported-model claim follows from this result.
|
|
|
|
| 1 |
+
# Confucius4-R2T2 — Architecture & Implementation Specification
|
| 2 |
+
|
| 3 |
+
> Technical companion specification for [Confucius4-R2T2-Q4_K_M-GGUF](https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF).
|
| 4 |
+
> Covers model architecture, checkpoint provenance, GGUF packaging, Longest Stable Prefix (LSP) streaming, and audio chunking contracts.
|
| 5 |
+
|
| 6 |
+
---
|
| 7 |
+
|
| 8 |
+
## 1. Model Architecture & Parameters
|
| 9 |
+
|
| 10 |
+
Confucius4-R2T2 is a fine-tune of **Qwen3-ASR-1.7B**, combining an audio encoder tower with a causal decoder language model:
|
| 11 |
+
|
| 12 |
+
### Audio Encoder Tower
|
| 13 |
+
- **Layers:** 24 encoder layers
|
| 14 |
+
- **Encoder Width:** 1024
|
| 15 |
+
- **Projected Audio Width:** 2048
|
| 16 |
+
- **Audio Frontend:** 128-channel log-mel filterbank (16 kHz, 25 ms window, 10 ms hop)
|
| 17 |
+
|
| 18 |
+
### Causal Decoder LM
|
| 19 |
+
- **Layers:** 28 transformer decoder layers
|
| 20 |
+
- **Query Heads:** 16
|
| 21 |
+
- **KV Heads:** 8 (Grouped-Query Attention, GQA)
|
| 22 |
+
- **Head Dimension:** 128
|
| 23 |
+
- **Activation:** SwiGLU (`gate_proj`, `up_proj`, `down_proj`)
|
| 24 |
+
- **Tied Embeddings:** `thinker.model.embed_tokens.weight` serves both token embedding lookup and output projection (the checkpoint has no separate `lm_head.weight`).
|
| 25 |
+
- **Context / Parameter Size:** ~1.7B parameters.
|
| 26 |
+
|
| 27 |
+
---
|
| 28 |
+
|
| 29 |
+
## 2. Checkpoint Provenance & Identity
|
| 30 |
+
|
| 31 |
+
- **Publisher Checkpoint:** [NetEase Youdao Confucius4-R2T2](https://huggingface.co/netease-youdao/Confucius4-R2T2), commit snapshot `185ce639118ad1362d049ca0d8ed04b6ec5cd6c9`.
|
| 32 |
+
- **Reference Precision:** BF16 (707 tensors across the audio tower, projections, and language model).
|
| 33 |
+
- **Weight License:** NetEase Youdao Model Use License Agreement (see `LICENSE` and `NOTICE`).
|
| 34 |
+
- **Canonical Implementation:** [NetEase Youdao R2T2 Repository](https://github.com/netease-youdao/Confucius4-R2T2).
|
| 35 |
+
|
| 36 |
+
---
|
| 37 |
+
|
| 38 |
+
## 3. GGUF Packaging & Self-Containment
|
| 39 |
+
|
| 40 |
+
The model GGUF is completely self-contained for local runtimes like [transcribe.cpp](https://github.com/NairoDorian/transcribe.cpp) and [audio.cpp](https://github.com/0xShug0/audio.cpp):
|
| 41 |
+
|
| 42 |
+
- **Embedded Assets:** Tokenizer configuration, vocabulary, processor config, generation parameters, and chat template are embedded within the GGUF file metadata.
|
| 43 |
+
- **No External Dependencies:** Local inference runtimes require only the engine binary and the `.gguf` file without needing Python or PyTorch sidecars.
|
| 44 |
+
- **Family & Schema:** Declared as family `confucius4_r2t2`.
|
| 45 |
+
|
| 46 |
+
---
|
| 47 |
+
|
| 48 |
+
## 4. Real-Time Streaming & Longest Stable Prefix (LSP)
|
| 49 |
+
|
| 50 |
+
Confucius4-R2T2 uses an append-only, low-latency streaming state machine:
|
| 51 |
+
|
| 52 |
+
### Streaming Mechanism
|
| 53 |
+
1. **Audio Accumulation & Chunking:** Incoming audio frames are fed incrementally in configurable chunks.
|
| 54 |
+
2. **Longest Stable Prefix (LSP) Decoding:** The decoder produces output tokens while comparing overlapping context windows across consecutive chunks. Text is only committed and emitted downstream when it is guaranteed to be stable and append-only.
|
| 55 |
+
3. **Rollback & Token Boundary Handling:** Uncommitted tail tokens are held across chunk boundaries to avoid emitting transient or truncated UTF-8 fragments, preventing word fragmentation.
|
| 56 |
+
4. **Final Flush:** Upon reaching end-of-stream (EOS) or VAD silence, remaining uncommitted tokens are finalized and emitted in the final transcript.
|
| 57 |
+
|
| 58 |
+
### Streaming Chunk Size Contract
|
| 59 |
+
- **Supported Range:** 80 ms to 2000 ms.
|
| 60 |
+
- **Default Recommendation:** **320 ms** (5120 audio samples at 16 kHz), providing an ideal balance between responsiveness and decoding accuracy on modern CPU/GPU hardware.
|
| 61 |
+
- **Fast / Low-Latency Mode:** 80 ms to 160 ms for interactive transcription where minimal delay is paramount.
|
| 62 |
+
|
| 63 |
+
---
|
| 64 |
+
|
| 65 |
+
## 5. Supported Capabilities
|
| 66 |
+
|
| 67 |
+
| Capability | Status | Description |
|
|
|
|
|
|
|
| 68 |
|---|---|---|
|
| 69 |
+
| **Multilingual Offline Transcription** | Supported | Full file batch decoding with automatic language identification. |
|
| 70 |
+
| **Real-Time Append-Only Streaming** | Supported | Low-latency streaming with LSP stability guarantees. |
|
| 71 |
+
| **Configurable Chunk Latency** | Supported | Direct numeric chunk duration configuration (80 ms – 2000 ms). |
|
| 72 |
+
| **Prompting & Hotwords** | Supported | Context prompting and vocabulary bias support. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|