--- title: BlueMagpie-TTS Demo emoji: 🐦‍⬛ colorFrom: blue colorTo: indigo sdk: gradio sdk_version: 5.49.1 app_file: app.py pinned: false license: apache-2.0 short_description: 台灣華語與中英混合文字轉語音 models: - OpenFormosa/BlueMagpie-TTS --- # BlueMagpie-TTS Demo [OpenFormosa/BlueMagpie-TTS](https://huggingface.co/OpenFormosa/BlueMagpie-TTS) 的線上展示,輸出 48 kHz 單聲道語音,支援台灣華語與中英混合文字。 ## 模式 - **內建語者**:使用模型發佈內附的 speaker centroid;預設為已重新稽核的 `female_voice`(介面中的內建語者 B)。 - **參考音色**:從至少 3 秒的授權參考音檔抽取 ECAPA speaker embedding;參考內容不需要逐字稿。 - **穩定長文**:使用和短句相同的 speaker embedding,切段後輸出一個完整音檔。 參考音色模式只把參考音檔轉成固定長度 speaker embedding,不把 raw reference-audio tokens 送入 TTS。Embedding 會在每個請求開始時抽取一次,並由該請求的所有文字切段共用。 ## Production Profile Demo 預設採用目前通過長文穩定性評估的推論設定: 模型 snapshot 固定在 `aaf1a0878e37875382bb0e5c8a3a2ba43be67297`; speaker-reference 模式使用的 ECAPA encoder 固定在 `0f99f2d0ebe89ac095bcc5903c4dd8f72b367286`;TTS runtime 以 `ce384c8cc54efea1aaba7b9f1d7ded6c1c99aa9a` 的已驗證原始碼隨 Space 保存, Barbet 另固定在 `6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8`, Mamba2 所需的 Triton runtime 以 `v2.3.2.post1` 的 Apache-2.0 上游原始碼隨 Space 保存(不綁定 `selective_scan_cuda` 編譯產物), 避免 Space 重啟後在沒有程式版本變更的情況下改變模型或 speaker embedding 空間。 候選語意驗證固定使用 Whisper large-v3-turbo revision `41f01f3fe87f28c78e2fbf8b568835947dd65ed9`;exact-assembled whole-output 語意驗證另固定使用 Whisper large-v3 revision `06f233fe06e710322aca913c1bc4249a0d71fce1`,兩者必須同時通過。 | 設定 | 值 | |---|---:| | CFG | fixed per-chunk schedule: row ordinal 0/even = 3.0, positive odd = 2.0; UI primary CFG is locked | | NFE steps | fixed at the validated value 10 | | Target pace | 4.0 speech units/sec | | Initial whole-trajectory policy | base: 5.2 CJK / 4.6 ASCII-mixed units/sec + 1 latent step | | Low-coverage chunk-refill policy | safe duration: 4.6 CJK / 4.0 ASCII-mixed units/sec + 1 latent step | | Sparse completion-headroom policy | every fourth retry: 4.2 CJK / 3.6 ASCII-mixed units/sec + 1 latent step | | Short-text guidance | minimum CFG 3.0 at no more than 6 speech units | | Generation guidance | fixed mixed-CFG assignment within the bounded cascade; public NFE fixed at 10 | | Email / URL frontend | email and generic URL fallback keep the hybrid explicit grammar; a bounded reversible HTTP(S) + lowercase short-path subset uses natural scheme/host/path cues with exact raw↔spoken atom-range proof | | Stop policy | 0.50 → 0.05 from 75% to 95% predicted progress, 1 hit | | Endpoint cue | append terminal punctuation for model input when missing, except very short text | | Hard stop | native-pace target steps, independent of playback pace; URL/email chunks use a separate conservative endpoint-duration counter (ASCII alnum run ÷ 2) | | Pace correction | derive total/active pace after completion; probe the corrected waveform and, when needed, re-render once from the untouched model waveform toward 3.95 active units/sec (ordinary `1536/384`; URL/email code-switch `2048/512`); no active-interval slicing | | Pace / speaker eligibility duration | union of active 25 ms RMS frames at 10 ms hop; internal pauses excluded | | Semantic verification | orthographic whole CER + tone-aware acoustic first/last 6 exact + no lexical tail; range-proven URL/email locals and every exact whole output require turbo/full-large-v3 hard intersection; fail closed | | Local endpoint verification | internal chunk first/last 6 may contain at most one substitution and zero deletion; absolute request onset/ending remain exact | | Acoustic quality gate | pinned CPU SQUIM objective SHA-256 `2c54586fea83fb5eb5394d710038ee89f55cab7011a5bf730bebed4c8777e828`; hard STOI ≥ 0.60 / PESQ ≥ 1.12; at ≥1.50 s, strict preferred return requires STOI ≥ 0.72 / PESQ ≥ 1.20 and speaker 0.25 / 0.05. A single-chunk exact joined waveform may also stop at the latency tier speaker 0.23 / 0.08, but only after both ASRs, preferred SQUIM, hard release speaker, pace, endpoint, and echo/smearing gates pass; local proxies cannot use this tier. SI-SDR is bounded soft evidence only | | Local speaker verification | frozen request centroid + bounded window begin/end directional drift | | Local hard speaker gate | similarity ≥ 0.10 and directional drop ≤ 0.10 | | Joined/final release-parity speaker guard | full-audio similarity ≥ 0.105; disjoint first/last-third drop ≤ 0.095 | | Local long speaker evidence | one ECAPA batch over up to four 3-second windows + begin/end | | Candidate cascade | exactly one same-seed whole trajectory, then deterministic single-chunk refills for zero/low-coverage rows only; a fully verified single-chunk exact waveform returns immediately instead of filling the preferred candidate pool | | Whole-trajectory qualification | assemble with production RMS/fade/pause/crossfade/speed, then whole-output gate | | Sequence fallback | ragged DP with up to 3 culprit-diverse paths; boundary-only local rejects may enter through a frozen 0.15 cap, then every exact assembly must pass turbo + full large-v3 and the stricter 0.105/0.095 whole gate | | Final output gate | re-verify joined/faded/RMS-matched/speed-adjusted whole waveform; fail closed | | Runtime budget | at most 32 generated TTS chunks and 800 generated speech units per request; NFE fixed at 10 | | Maximum chunk | ordinary request 80 speech units; every chunk in a URL/email request 36 units | | Minimum chunk | ordinary text 12 speech units; URL/email-bearing chunks 8 units; genuine short non-network requests/natural short sentence boundaries are preserved | | Crossfade / internal edge fade | semantic boundary 80 ms / 80 ms with at least 250 ms silence; proven URL/email internal boundary 5 ms / 5 ms with 400 ms zero-silence verifier seam | | Chunk RMS adjustment | at most 4 dB | | Generated continuation context | 0 sec | | Generated-audio retry | disabled | | Quality selection | enabled; no unverified fallback | SQUIM objective runtime follows the official [Torchaudio SQUIM documentation](https://docs.pytorch.org/audio/stable/tutorials/squim_tutorial.html). The bundled DNS-Challenge weights are attributed under [CC BY 4.0](https://github.com/microsoft/DNS-Challenge/blob/interspeech2020/master/LICENSE); the runtime additionally verifies the exact weight SHA-256 before deserialization. ### V5 expanded glyph asset release boundary The V5 candidate's expanded production glyph manifest is an immutable provenance and audio integrity object. The renderer and bundle share one exact 32-entry grammar: `letter-a` through `letter-z` are spoken as `第一個英文字母` through `第二十六個英文字母`; `symbol-at`, `symbol-dot`, `symbol-hyphen`, `symbol-underscore`, `symbol-colon` and `symbol-slash` are respectively `網路符號小老鼠`, `網路符號點`, `這個符號叫做橫線`, `這個符號叫做底線`, `這個符號叫做冒號` and `網路符號斜線`. Its recorded Whisper, ECAPA and SQUIM values are **attestations only**: successfully loading the manifest never sets `runtime_quality_verified`, never runs those evaluators, and is not a release pass. A release evaluator must independently load the pinned Whisper large-v3-turbo, Whisper large-v3 (and the frozen third ASR where required), ECAPA and SQUIM artifacts, recompute their evidence from the exact assembled PCM16 delivered to the user, and apply the unchanged frozen release gates. Silent PCM accompanied by invented passing numbers is rejected and can never be promoted by the asset loader. Every hybrid audio plan hard-binds the externally produced semantic-plan hash, inverse-proof hash and a content-free ordered request-attempt ledger. Every generator invocation has a unique ordered record containing its attempt ID, seed, text hash, endpoint-evidence hash, generated-chunk count and generated text-unit count. The commitment independently binds the ledger hash, record count, invocation total, generated-chunk total and generated-text-unit total; discarded candidates and retries therefore remain in the 32-chunk/800-unit request accounting. Schema-v7 release evaluation must reconstruct the same attempt ledger rather than deriving work only from final selected audio. Every final carrier additionally binds the selected attempt ID, seed, text hash and endpoint-evidence hash. Verification receives the trusted loaded bundle, the original ordered carrier/asset segment specifications and that external commitment; it reconstructs the plan, conditions each carrier again, checks every asset's retained WAV/PCM/core and grammar labels, rebuilds the output and compares exact PCM slices. Ledger labels and digest strings are never sufficient evidence on their own. Asset WAVs are opened once with no-follow semantics; container hashing and WAV decoding use the same retained bytes. User segment IDs cannot enter the private pause/final-silence namespace. Assets have a recomputed 3 ms builder-side leading-edge envelope, and every assembled boundary has a frozen absolute jump limit of 0.10. The assembler never fades an asset at runtime. If the final asset already contains at least 180 ms of exact zero PCM, no second 180 ms tail is appended; otherwise only the missing zero samples are added, preserving every frozen asset byte. 每個請求會取得新的隨機 root seed。服務先且只先生成一次完整 trajectory;其中所有長文切段 共用 root seed,並使用 base policy(5.2 CJK / 4.6 ASCII)。如果這條 exact whole path 未通過, 服務會保留每一段的 local evidence,再依 `(coverage, refill attempts, chunk index)` 穩定順序, 只對 zero/low-coverage chunk 逐一生成 refill。每個 chunk 依自己的 refill ordinal 輪替 policy;大多數 refill 使用 safe-duration(4.6 CJK / 4.0 ASCII), 每第四個 retry 使用 completion-headroom(4.2 CJK / 3.6 ASCII)。Refill 使用 `root seed + global generation offset`,不會重新生成其他已覆蓋 chunks。三種 policy 都只改 native-duration endpoint estimate, margin 固定 +1 latent step、`min_len` 固定為 2,不會為了播放目標語速硬撐 generation loop。 唯一例外是 completion-headroom 的一至兩單位短句:其 sparse retry 將 hard cap 下限提高到 5 steps, 讓第一個 eligible weak-stop 有一次真正的 threshold 判定機會;所有額外內容仍須通過雙 ASR 與 tail=0。 選中完整 trajectory 或 coverage-sequence-DP 混合路徑後,log 會記錄每個 chunk 的 candidate、seed、policy、實際 CFG, 並記錄每個已嘗試 row-local ordinal 的 schedule CFG。Ordinal 0 與正偶數使用 primary CFG 3.0, 正奇數使用 alternate CFG 2.0;含 URL/email component 的 chunk 仍有 CFG 3.0 的內容安全下限。 同一列 log 也記錄實際 generated chunks/text units 與各 row 的有效候選數。 每次生成前另外記錄 seed、local chunk index、duration policy、scheduled CFG, endpoint-duration units/counter,以及 network/short-text floor 後的 effective CFG;因此即使最後 fail closed 也能重建嘗試軌跡。公開 pace、chunk planner 與 32/800 work budget 仍使用原本的 speech-unit counter;只有 URL/email-conditioned native endpoint estimate 對連續 ASCII alnum run 採用除以 2 的保守 counter。這個 counter 刻意同時決定 adaptive weak-stop 的 `expected_steps` progress 與 `max_len` hard cap;兩者屬於同一份 native endpoint plan, 但不會進入播放 pace、chunk planning 或 32/800 work accounting,`min_len` 仍固定為 2。 終端 outcome 另輸出 content-free canonical evidence schema v6,包含 original chunk indices、 row-local ordinal、network provenance、CFG contract、32/800 預算用量與 selected-path 交叉檢查, 每個 chunk 的 stop reason、pre-fade endpoint energy、generated/hard-stop steps, 每個 local row 的 independent-large-v3 attempted/pass/proof-count/result attestation,以及 local、joined、independent-large-v3 與 final gate 的 bounded SQUIM scalar, 不包含 target/transcript/audio/embedding。 為了讓 release contract 與正式量測一致,UI 不提供其他 primary CFG。 內部 hosted evaluator 可注入 `[0, 2^31)` 的固定 root seed 以重現結果;UI 不暴露這個參數, 一般請求仍只在未注入 seed 時使用系統亂數。 初始完整 same-seed trajectory 永遠優先;通過 turbo + full large-v3 exact-whole gate 就立即回傳, 不會執行 refill 或 DP。若失敗,之後只做單 chunk refill,不再生成第二條完整 trajectory。 Ragged DP 可混合已通過 hard gate 的 chunks, 以及下述唯一 boundary-only、drop 不超過 0.15 的 DP-only chunks。切段保留標點,並依逗號、分號或句末標點 插入不同長度的停頓。已通過 hard gate 的候選再加入 bounded endpoint soft cost: 任何到達 hard cap(即使 threshold 恰於同一步 crossing)皆加 0.05,再加上 `0.02 ×` 最後 5 ms RMS/全波形 peak。只有在目前 preferred 候選到達 hard cap 時,才允許真正於 cap 前自然停止的候選使用最多 STOI 0.02/PESQ 0.03 的 preferred-tier slack;speaker 與 boundary 在既有長度 gate 適用時仍須完整達到 preferred 門檻,所有 hard gate 亦維持不變。 輸出最後會套用保守的 RMS floor 與 peak limit、固定 3 ms cosine-squared 收尾與 180 ms silence, 接著先量化到 PCM16 並以 WAV decoder 等價波形執行 joined/final verifier。公開 wrapper 再編碼時 逐 sample 冪等,因此驗證的就是使用者收到的 PCM codes;同時不讓 Gradio 對 float ndarray 再做 peak normalization,以免額外放大底噪。 完成 RMS matching、edge fade、pause、crossfade 與使用者 speed 後,服務會再對最終整段 waveform 執行 normalized target 的 prefix/whole/suffix/tail、pace 與 speaker anchor/boundary gate,並要求 turbo 與 full large-v3 的語意 hard intersection;這個 post-join gate 失敗時不會回傳先前已通過的 chunk 音訊。除此之外,每個帶有 planner range proof 的 URL/email local chunk 在進入 coverage pool 以前,也必須用完全相同的 `NetworkFragmentProof` 同時通過 turbo 與 pinned full large-v3; large-v3 reject 會成為不可被 boundary proxy 放寬的 semantic hard failure。一般 local chunk 維持 turbo gate;已被 turbo 的非 boundary hard gate 拒絕、原本就不可能進 pool 的 network row 不額外耗用 full large-v3。Production assembler 產生的 exact whole candidate 或 exact sequence path 仍全部 使用 dual-ASR。Initial whole rejection 會啟動低覆蓋 refill,sequence rejection 則只會前進到 下一條已排序的 exact path。 此外,初始 same-seed trajectory 只有在每個 local chunk 都通過後,才會先用同一個 `_assemble_trajectory_audio` 組成實際播放版本並跑整段 semantic/pace/speaker gate;joined gate 失敗會取消該 whole trajectory 的資格,但保留已通過的 local evidence 供 sequence DP 使用。 這個額外 dual-ASR whole gate 只對 local turbo 全數通過的候選執行,DP 選定後仍會保留上述最終 post-join gate。每個 request 只在 waveform float32 bytes、sample rate、normalized target 與 pinned verifier profile 完全相同時重用 full large-v3 evidence;任何 sample 改變都會重新驗證。ASR loader 異常、OOM、空 transcript 或無效 evidence 都 fail closed,不會降級為 turbo-only。同一 request 內若 release 與 semantic verifier 使用完全相同的精確 waveform、sample rate、Whisper model/revision、語言、token cap 與 segmentation cap,會共用一次 transcript decode;cache 不跨 request,exception 與無效結果不會寫入。 初始 whole trajectory 未通過後,refill scheduler 會優先補 coverage=0 的 rows,再補不足 3 個 可用候選的低覆蓋 rows;失敗的 refill 也完整計入 generation budget,避免單一壞段無限重試。 Multi-chunk request 一旦每個 pool 首次都有可用候選,會先建立 rank-1 ragged DP path 並執行 完全相同的 exact whole verifier;只有同時通過 hard gate 與 preferred speaker/SQUIM tier 才提前 回傳,否則繼續原本的 bounded refill。完成 refill 或觸及任一硬預算後再建立完整 ragged DP, 不要求每個 chunk 有相同候選數。 已通過 local hard gate 的 chunk 可直接進入 DP;唯一 rejection reason 是 `boundary_speaker_drop`、且 drop 不超過 0.15 的 chunk 也可作為 DP-only 候選。Single-chunk request 可使用相同的 local proxy,但仍必須重新通過 exact whole-waveform 0.105/0.095 release guard。 若 exact joined waveform 已通過 dual-ASR,single-chunk preferred early return 以該 release ECAPA/SQUIM evidence 為準,不會因較粗的 local windowed boundary proxy 再生成多個無用候選。 這個 boundary-only relaxation 不適用於 semantic、pace、speaker similarity、缺少 evidence 或 whole-trajectory accept,也不能在缺少整段 final verifier 時啟用。Local score 加上相鄰 ECAPA/RMS transition cost 會產生最多 3 條 culprit-diverse 完整路徑;除了最低總成本,後續 exact checks 會優先選擇 在初始失敗 chunks 上使用不同候選的路徑,避免三次驗證都只改動已穩定段。每條都用同一個 production assembler 和整段 semantic/pace/speaker gate 驗證;整段仍必須通過 similarity 0.105 與 boundary drop 0.095, 第一條通過才回傳。三條都失敗、callback 異常或沒有完整 finite path 時一律 fail closed。 被拒的 joined/final log 只記 normalized CER、prefix/suffix CER、tail units 與 reasons,不記逐字 transcript。 候選排序與 DP 保留低成本的 windowed speaker evidence;joined candidate、k-best path 與 最終 waveform 會再跑一次與獨立發布評估定義相同的 ECAPA guard:原始 sample rate 上以 25/10 ms RMS 取 active 外緣,將整段與三個不重疊 thirds 分開 encode。這避免短音檔的 1.5 秒首尾視窗重疊而掩蓋句尾 speaker drift。 日期、24 小時制時間、百分比、常見單位與大寫 acronym/model code 會先轉成保守的 zh-TW spoken form,例如 `2026/07/16`、`15:30`、`12.5%` 與 `10 km`。ASR 會先統一 繁簡字形再評分,避免把正確的台灣華語輸出誤判為內容錯誤。 Email 與一般 URL fallback 會以混合式可辨識讀法展開:一般 local/domain/path label 保留 lexical ASCII(例如 `tour`、`help`),scheme、全大寫 atom 與短 suffix 使用明確 ASCII letter tokens,數字逐位朗讀,分隔符也明確朗讀(例如 `.tw` 讀成「點、T、W」)。Email contract 不受下述 URL naturalization 影響。 另有一個刻意受限、可逆的 URL grammar:只接受 canonical lowercase `http`/`https`、 lowercase ASCII DNS host、至少一個能由固定 lexicon 唯一分詞的 compound label、無 userinfo/port/query/fragment,以及可選的單一 1–4 個 lowercase ASCII 字母 path;naturalized spoken form 本身也必須不超過 36 public speech units。`https` 使用「安全網址是」, `http` 使用不同的「網址是」;compound host 會自然分詞,dot 讀「點」,短 suffix 與 path 逐字母讀,最後明說「網址輸入完畢」。例如 `https://coastwatch.example.tw/tide` 會成為 `安全網址是 coast watch 點 example 點 T W,路徑是 T I D E,網址輸入完畢`。 大小寫、長 path、query、port、無唯一 compound 分詞或超限 URL 一律回到原本的 explicit frontend,不作有損猜測。 Naturalized URL 每一個 scheme/host/dot/path/terminator atom 都保存原始 raw slice 與 emitted spoken range;planner、Space runtime 與 local ASR 入口都會從 raw identifier 重新 render 並要求 proof 完全相等。若 nested proof 被移除、range/rule/reordering 被修改,或非 naturalized identifier 夾帶這組 provenance,會 fail closed。這個 naturalized phrase 在 generation planner 中是一個 不可切 component;H07 因此是 35 public units 的單一 network chunk,沒有 artificial `network_internal` seam。兩個中文逗號只是模型 pause cue,Whisper 省略它們仍可通過;scheme cue、 host、path 與 terminator 仍是 exact lexical protected range。 完整 identifier 仍保留為 joined、full-large-v3 與 final ASR 的 exact protected target。完整 URL spoken proof 不超過 36 units 時仍保持單一 atomic chunk;URL 後的逗號也留在同一 chunk。 Hosted profile 對 bounded email 只開放一個更窄的例外:同一 strong sentence 必須恰好只有一個 email、email 前的 prose 至少 12 units、domain fragment 至少 8 units,且在 range-proven `@`/ `小老鼠` separator 切開後兩側都介於 8–36 units,才把 local part 與 domain 分成兩個 generation rows。短 email、裸 email、多 email、URL 與不滿足上述 bounds 的輸入仍保持 atomic。這個切分只 改變 generation trajectory;joined exact target、proof coverage 與最終 ASR gate 都不變。 Proof 超過 36 units 時,planner 仍只在由原始 ASCII grammar 證明的 scheme/domain/path/query/email component 邊界切段,並維持 32 units 目標、36 units 硬上限與 8 units network minimum。過長 email 仍可在 `小老鼠`之前、過長 explicit URL 仍可在 `冒號 斜線 斜線` 之後切開;像 `a@b.co`、`https://a.tw` 的某一側不足 8 units 且第一輪無解時, 才放寬該短 identifier 的 preferred cut。Identifier 內部 需要拆分時,Email 第一輪必須在 `小老鼠`、explicit URL 第一輪必須在 scheme 結尾的 grammar boundary 切開;第一輪 DP 不得跨越這兩類 mandatory cut。 一般 semantic boundary 保留原標點停頓,但下限為 250 ms,讓 waveform-only verifier 可以將已通過 local gate 的 chunk 穩定分開。具共享 component proof 的 `network_internal` seam 仍插入 400 ms zero silence,並使用 5 ms fade/crossfade。 Component proof 無法完整 重建、非 ASCII IRI 或單一不可拆 component 超限時會 fail closed。 Local fragment gate 只接受 range-bound exact proof,並檢查所有最佳 edit alignment;重複的普通文字 不能借用受保護片段的 proof。Joined、whole 與 final gate 不套用這項 local-only canonicalization。 這能降低模型把不常見 TLD 自動補成 `.com` 的風險;一般英文句子不會套用這個規則。 URL 與後續英文 prose 應以空白或中文標點分隔;未分隔的 RFC path punctuation 會視為 URL 本身的一部分並納入 exact gate。Quoted email local-part 暫不支援,輸入時會直接 fail closed。 URL identifier 目前只支援 ASCII;非 ASCII IRI 必須先轉成 ASCII/percent-encoded 形式, 否則會因字母讀音碰撞而在 normalization 前 fail closed。 為避免 Whisper 自動句末標點和 URL 內容不可判定,URL 不接受以 `. , ! ? ; : '` 結尾; 需要這些尾端符號時請使用 percent encoding。 Whisper 常見的 `15点30分`/`15點30分` 會視為同一讀法;裸寫 `15點30` 只有在 target 明確含同一個完整鐘點讀法時才作 ASR-only canonicalization。比分、秒數、重量、裸 `15點`、 非法時間與真實漏字仍會被拒絕。Whole-output ASR 只使用 waveform evidence:若所有符合 250 ms、兩側皆有 active speech 的 pause atoms 都能將音訊限制在 28 秒內,就保留 每一個 atom 獨立驗證(包含 semantic 250 ms 與 proof-bound 400 ms network seam)。只有 某個 atom 仍超過 28 秒時,才在相同 qualified pauses 上使用 coarse latest-pause fallback; 超過 12 個 atoms 或找不到安全 pause 都 fail closed,不使用最低能量或 target-dependent fallback。每段至少 1.25 秒、硬上限 30 秒,每次驗證最多 12 段, 並以最多 6 段的 microbatch 執行 ASR。 整句 intelligibility 仍使用 orthographic CER 上限 0.20。首尾、刪字與額外尾詞則使用固定 `pypinyin==0.55.0` 的 lexical-tone units 做完整 Levenshtein alignment:例如 `陶藝/陶逸` 可視為同一聲學讀法,`陶藝/陶一` 與 `改到/改造` 仍會被拒。ASCII network identifier 保持 逐字元比較,numeric/frequency contradiction 也仍 fail closed。這項 acoustic endpoint normalization 不會修改送入模型的文字;既有同音代詞 canonicalization 仍保留。 只有尚未 assemble 的 internal local chunk 允許首、尾 6 units 各最多一個 substitution,且 deletion 必須為 0;由原始 `GenerationChunkSpec` 證明的 request 第一段 prefix 與最後一段 suffix 仍必須 exact。Refill 依原始絕對 chunk provenance 套用相同角色。Joined turbo、 independent full-large-v3 與 final whole-waveform gate 全部維持 exact endpoint。 自然 stop 在目前 checkpoint 上仍可能過快或錯過句尾。Demo 不再用播放目標語速決定生成長度; 模型使用原生語速的安全上限,完成後才做保音高語速校正。缺少句末標點時只在模型輸入補上句號, 不超過 6 個 speech units 的短句會使用最低 CFG 3.0;stop threshold 只在預估進度 75% 後逐步降低,並於 95% 進度降至 0.05,以捕捉句末弱 stop 訊號而不影響前段內容。 一至兩個 speech units 的極短普通文字若原本會以 1.0 倍速輸出,會從未經處理的 model waveform 直接做一次保音高的 0.95 倍重繪,再接受完整 semantic、SQUIM、endpoint 與 PCM16 gate; 不會串接多次 time-scale modification,也不會放寬任何候選門檻。 最多 32 speech units 的一般 request 即使含標點也保持單一 generation chunk;33–48 units 才依自然標點切段,長文仍受 80-unit 上限約束。這避免常見約 6 秒請求因 24-unit cliff 不必要地變成兩次生成,同時不增加模型允許的最大 chunk。Release profile 鎖定 primary CFG 3.0 / alternate CFG 2.0 交錯候選排程、NFE 10、 後處理語速 1.0。 Online pace gate 與 1.5 秒 speaker eligibility 使用和獨立 hosted evaluator 相同的 active-frame interval union(25 ms frame、10 ms hop、peak -35 dB、absolute RMS floor 1e-4),因此句內或插入的 pause 不會稀釋 CPS。逐 active interval 的 stretcher 曾在固定 root ablation 造成長文 availability 回退,因此仍不切開 active intervals;production 會先在 model 的 raw completed waveform 上合併 total-duration 與 active-union 的 bounded rate,再量測一次 probe waveform。若 probe active pace 仍高於 3.95,會把 residual rate 乘回原 combined rate,並從 untouched model waveform 重新 render; 最後交給 verifier 與使用者的 samples 因此只經過一層 deterministic mono WSOLA。自動 stretch rate 的硬下限固定為 0.90,過去 long-text `0.76` escape 已移除;若 0.90 仍無法通過 4.85 CPS gate,必須改選候選或 fail closed,不得再拉長。公開的後處理語速固定為 1.0,避免第二次 time-scale modification 與隱性的 rate 相乘。 所有不少於 1.0 秒的 local、joined 與 final waveform 另通過 deterministic no-reference echo/smearing gate。Gate 聚合多個 active windows 的 24–180 ms 長 quefrency cepstrum,拒絕固定 delayed-copy comb peak 或 dense smearing signature;無效、非有限或缺少足夠 active-window evidence 都 fail closed。短於 1.0 秒時 estimator 不適用,仍由原有 strict semantic、endpoint 與 SQUIM gate 負責。任何仍過快、被 stretch 破壞或具有可量測回音的候選均不會回傳。 每個 local row 先完成 Turbo ASR 與同一 active-duration pace gate;若此時已確定為 semantic 或 pace hard failure,便保留該拒絕 evidence 並跳過 ECAPA、SQUIM、echo 與 F0。只有仍可能 回傳或進入 DP 的候選才支付完整 acoustic verification,最終 speaker/SQUIM/echo 門檻不變。 Median-F0 只對實際具備 DP coverage eligibility 的 multi-chunk rows 計算,single/K32、semantic reject、joined 與 final waveform 不再支付無用的 pYIN latency。 Generation work 以實際產生的 chunks 與其 speech units 累計:初始 trajectory 花費 `chunk count` 與全文 speech units;每次 refill 再加 1 chunk 與該段 units。任一 request 最多 32 generated TTS chunks、800 generated speech units,且 refill 只跑固定 NFE 10。 因此長文不再為了取得某一段替代候選而重生整篇,預算可直接集中在 ASR/speaker gate 顯示的 zero/low-coverage rows。若某 row 在預算內仍為 zero coverage,或 ragged graph 沒有 finite path, 服務立即 fail closed,不會放寬 semantic、speaker、pace 或 endpoint gate。 joined/final ECAPA 會從 1.48 秒 active duration 提前觸發,吸收 WAV round-trip 可能造成的 一至兩個 RMS frame 邊界位移;正式 release speaker eligibility 仍固定為 1.50 秒。 若所有候選仍有漏字、多說字、句尾截斷、語速過快或可量測的音色漂移,服務會拒絕輸出, 不會回傳任意 fallback。極短音檔的分段 ECAPA 不可靠,因此只採嚴格語意 gate;SQUIM hard gate 仍保留,但低於 1.50 秒不以 duration-biased preferred tier 強迫 refill 到 32-candidate 上限,並在 release audit 中獨立標記為 speaker-unverified;這不代表模型的短句音色已被根治。 目前部署路徑不使用 reference-style distribution score;固定 speaker centroid 只代表 identity, 不宣稱複製細緻情緒或韻律。Sequence DP 會使用相鄰候選的 ECAPA、active RMS 與 normalized median-F0 軟成本來降低跨句音色、響度與基準音高跳變;任何選中路徑仍必須重新通過整段 semantic、speaker、pace 與 endpoint hard gates。 請只使用已取得授權的參考音檔。合成語音僅供研究與評估展示,正式使用前請人工檢視。 程式碼與本機使用方式:[OpenFormosa/BlueMagpie-TTS](https://github.com/OpenFormosa/BlueMagpie-TTS)