Spaces:
Running on Zero
Download README.md from voidful/BlueMagpie-TTS-Demo: direct link, hf CLI and curl.
- Browser
- Download file 23.6 kB
-
https://huggingface.co/spaces/voidful/BlueMagpie-TTS-Demo/resolve/3352f9fefd25938525506fe5ebcdc60a4d029e8a/README.md
- Command line
-
hf download hf://spaces/voidful/BlueMagpie-TTS-Demo@3352f9fefd25938525506fe5ebcdc60a4d029e8a/README.md
-
curl -L -o README.md https://huggingface.co/spaces/voidful/BlueMagpie-TTS-Demo/resolve/3352f9fefd25938525506fe5ebcdc60a4d029e8a/README.md
title: BlueMagpie-TTS Demo
emoji: 🐦⬛
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: apache-2.0
short_description: 台灣華語與中英混合文字轉語音
models:
- OpenFormosa/BlueMagpie-TTS
BlueMagpie-TTS Demo
OpenFormosa/BlueMagpie-TTS 的線上展示,輸出 48 kHz 單聲道語音,支援台灣華語與中英混合文字。
模式
- 內建語者:使用模型發佈內附的 speaker centroid;預設為已重新稽核的
female_voice(介面中的內建語者 B)。 - 參考音色:從至少 3 秒的授權參考音檔抽取 ECAPA speaker embedding;參考內容不需要逐字稿。
- 穩定長文:使用和短句相同的 speaker embedding,切段後輸出一個完整音檔。
參考音色模式只把參考音檔轉成固定長度 speaker embedding,不把 raw reference-audio tokens 送入 TTS。Embedding 會在每個請求開始時抽取一次,並由該請求的所有文字切段共用。
Production Profile
Demo 預設採用目前通過長文穩定性評估的推論設定:
模型 snapshot 固定在 aaf1a0878e37875382bb0e5c8a3a2ba43be67297;
speaker-reference 模式使用的 ECAPA encoder 固定在
0f99f2d0ebe89ac095bcc5903c4dd8f72b367286;TTS runtime 以
ce384c8cc54efea1aaba7b9f1d7ded6c1c99aa9a 的已驗證原始碼隨 Space 保存,
Barbet 另固定在 6fcd7ce4aa37f2250a3242995bef0fbc3b026ba8,
避免 Space 重啟後在沒有程式版本變更的情況下改變模型或 speaker embedding 空間。
候選語意驗證固定使用 Whisper large-v3-turbo revision
41f01f3fe87f28c78e2fbf8b568835947dd65ed9;exact-assembled whole-output
語意驗證另固定使用 Whisper large-v3 revision
06f233fe06e710322aca913c1bc4249a0d71fce1,兩者必須同時通過。
| 設定 | 值 |
|---|---|
| CFG | fixed per-chunk schedule: row ordinal 0/even = 3.0, positive odd = 2.0; UI primary CFG is locked |
| NFE steps | fixed at the validated value 10 |
| Target pace | 4.0 speech units/sec |
| Initial whole-trajectory policy | base: 5.2 CJK / 4.6 ASCII-mixed units/sec + 1 latent step |
| Low-coverage chunk-refill policy | safe duration: 4.6 CJK / 4.0 ASCII-mixed units/sec + 1 latent step |
| Sparse completion-headroom policy | every fourth retry: 4.2 CJK / 3.6 ASCII-mixed units/sec + 1 latent step |
| Short-text guidance | minimum CFG 3.0 at no more than 6 speech units |
| Generation guidance | fixed mixed-CFG assignment within the bounded cascade; public NFE fixed at 10 |
| Email / URL frontend | email and generic URL fallback keep the hybrid explicit grammar; a bounded reversible HTTP(S) + lowercase short-path subset uses natural scheme/host/path cues with exact raw↔spoken atom-range proof |
| Stop policy | 0.50 → 0.05 from 75% to 95% predicted progress, 1 hit |
| Endpoint cue | append terminal punctuation for model input when missing, except very short text |
| Hard stop | native-pace target steps, independent of playback pace; URL/email chunks use a separate conservative endpoint-duration counter (ASCII alnum run ÷ 2) |
| Pace correction | derive total/active pace after completion; probe the corrected waveform and, when needed, re-render once from the untouched model waveform toward 3.95 active units/sec (ordinary 1536/384; URL/email code-switch 2048/512); no active-interval slicing |
| Pace / speaker eligibility duration | union of active 25 ms RMS frames at 10 ms hop; internal pauses excluded |
| Semantic verification | orthographic whole CER + tone-aware acoustic first/last 6 exact + no lexical tail; range-proven URL/email locals and every exact whole output require turbo/full-large-v3 hard intersection; fail closed |
| Local endpoint verification | internal chunk first/last 6 may contain at most one substitution and zero deletion; absolute request onset/ending remain exact |
| Acoustic quality gate | pinned CPU SQUIM objective SHA-256 2c54586fea83fb5eb5394d710038ee89f55cab7011a5bf730bebed4c8777e828; hard STOI ≥ 0.60 / PESQ ≥ 1.12; at ≥1.50 s, single-chunk early return additionally requires preferred STOI ≥ 0.72 / PESQ ≥ 1.20 and speaker 0.25 / 0.05; SI-SDR is bounded soft evidence only |
| Local speaker verification | frozen request centroid + bounded window begin/end directional drift |
| Local hard speaker gate | similarity ≥ 0.10 and directional drop ≤ 0.10 |
| Joined/final release-parity speaker guard | full-audio similarity ≥ 0.105; disjoint first/last-third drop ≤ 0.095 |
| Local long speaker evidence | one ECAPA batch over up to four 3-second windows + begin/end |
| Candidate cascade | exactly one same-seed whole trajectory, then deterministic single-chunk refills for zero/low-coverage rows only |
| Whole-trajectory qualification | assemble with production RMS/fade/pause/crossfade/speed, then whole-output gate |
| Sequence fallback | ragged DP with up to 3 culprit-diverse paths; boundary-only local rejects may enter through a frozen 0.15 cap, then every exact assembly must pass turbo + full large-v3 and the stricter 0.105/0.095 whole gate |
| Final output gate | re-verify joined/faded/RMS-matched/speed-adjusted whole waveform; fail closed |
| Runtime budget | at most 32 generated TTS chunks and 800 generated speech units per request; NFE fixed at 10 |
| Maximum chunk | ordinary request 80 speech units; every chunk in a URL/email request 36 units |
| Minimum chunk | ordinary text 12 speech units; URL/email-bearing chunks 8 units; genuine short non-network requests/natural short sentence boundaries are preserved |
| Crossfade / internal edge fade | semantic boundary 80 ms / 80 ms with at least 250 ms silence; proven URL/email internal boundary 5 ms / 5 ms with 400 ms zero-silence verifier seam |
| Chunk RMS adjustment | at most 4 dB |
| Generated continuation context | 0 sec |
| Generated-audio retry | disabled |
| Quality selection | enabled; no unverified fallback |
SQUIM objective runtime follows the official Torchaudio SQUIM documentation. The bundled DNS-Challenge weights are attributed under CC BY 4.0; the runtime additionally verifies the exact weight SHA-256 before deserialization.
每個請求會取得新的隨機 root seed。服務先且只先生成一次完整 trajectory;其中所有長文切段
共用 root seed,並使用 base policy(5.2 CJK / 4.6 ASCII)。如果這條 exact whole path 未通過,
服務會保留每一段的 local evidence,再依 (coverage, refill attempts, chunk index) 穩定順序,
只對 zero/low-coverage chunk 逐一生成 refill。每個 chunk 依自己的 refill ordinal 輪替 policy;大多數 refill 使用 safe-duration(4.6 CJK / 4.0 ASCII),
每第四個 retry 使用 completion-headroom(4.2 CJK / 3.6 ASCII)。Refill 使用
root seed + global generation offset,不會重新生成其他已覆蓋 chunks。三種 policy 都只改 native-duration endpoint estimate,
margin 固定 +1 latent step、min_len 固定為 2,不會為了播放目標語速硬撐 generation loop。
選中完整 trajectory 或 coverage-sequence-DP 混合路徑後,log 會記錄每個 chunk 的 candidate、seed、policy、實際 CFG,
並記錄每個已嘗試 row-local ordinal 的 schedule CFG。Ordinal 0 與正偶數使用 primary CFG 3.0,
正奇數使用 alternate CFG 2.0;含 URL/email component 的 chunk 仍有 CFG 3.0 的內容安全下限。
同一列 log 也記錄實際 generated chunks/text units 與各 row 的有效候選數。
每次生成前另外記錄 seed、local chunk index、duration policy、scheduled CFG,
endpoint-duration units/counter,以及 network/short-text floor 後的 effective CFG;因此即使最後
fail closed 也能重建嘗試軌跡。公開 pace、chunk planner 與 32/800 work budget 仍使用原本的
speech-unit counter;只有 URL/email-conditioned native endpoint estimate 對連續 ASCII alnum run
採用除以 2 的保守 counter。這個 counter 刻意同時決定 adaptive weak-stop 的
expected_steps progress 與 max_len hard cap;兩者屬於同一份 native endpoint plan,
但不會進入播放 pace、chunk planning 或 32/800 work accounting,min_len 仍固定為 2。
終端 outcome 另輸出 content-free canonical evidence schema v4,包含 original chunk indices、
row-local ordinal、network provenance、CFG contract、32/800 預算用量與 selected-path 交叉檢查,
每個 local row 的 independent-large-v3 attempted/pass/proof-count/result attestation,以及
local、joined、independent-large-v3 與 final gate 的 bounded SQUIM scalar,
不包含 target/transcript/audio/embedding。
為了讓 release contract 與正式量測一致,UI 不提供其他 primary CFG。
內部 hosted evaluator 可注入 [0, 2^31) 的固定 root seed 以重現結果;UI 不暴露這個參數,
一般請求仍只在未注入 seed 時使用系統亂數。
初始完整 same-seed trajectory 永遠優先;通過 turbo + full large-v3 exact-whole gate 就立即回傳,
不會執行 refill 或 DP。若失敗,之後只做單 chunk refill,不再生成第二條完整 trajectory。
Ragged DP 可混合已通過 hard gate 的 chunks,
以及下述唯一 boundary-only、drop 不超過 0.15 的 DP-only chunks。切段保留標點,並依逗號、分號或句末標點
插入不同長度的停頓。輸出最後會套用保守的 RMS floor 與 peak limit,通過所有 verifier
後才明確量化為 PCM16;不讓 Gradio 對 float ndarray 再做 peak normalization,以免額外放大底噪。
完成 RMS matching、edge fade、pause、crossfade 與使用者 speed 後,服務會再對最終整段 waveform
執行 normalized target 的 prefix/whole/suffix/tail、pace 與 speaker anchor/boundary gate,並要求
turbo 與 full large-v3 的語意 hard intersection;這個 post-join gate 失敗時不會回傳先前已通過的
chunk 音訊。除此之外,每個帶有 planner range proof 的 URL/email local chunk 在進入 coverage pool
以前,也必須用完全相同的 NetworkFragmentProof 同時通過 turbo 與 pinned full large-v3;
large-v3 reject 會成為不可被 boundary proxy 放寬的 semantic hard failure。一般 local chunk
維持 turbo gate;已被 turbo 的非 boundary hard gate 拒絕、原本就不可能進 pool 的 network row
不額外耗用 full large-v3。Production assembler 產生的 exact whole candidate 或 exact sequence path 仍全部
使用 dual-ASR。Initial whole rejection 會啟動低覆蓋 refill,sequence rejection 則只會前進到
下一條已排序的 exact path。
此外,初始 same-seed trajectory 只有在每個 local chunk 都通過後,才會先用同一個
_assemble_trajectory_audio 組成實際播放版本並跑整段 semantic/pace/speaker gate;joined gate
失敗會取消該 whole trajectory 的資格,但保留已通過的 local evidence 供 sequence DP 使用。
這個額外 dual-ASR whole gate 只對 local turbo 全數通過的候選執行,DP 選定後仍會保留上述最終
post-join gate。每個 request 只在 waveform float32 bytes、sample rate、normalized target 與 pinned
verifier profile 完全相同時重用 full large-v3 evidence;任何 sample 改變都會重新驗證。ASR loader
異常、OOM、空 transcript 或無效 evidence 都 fail closed,不會降級為 turbo-only。
初始 whole trajectory 未通過後,refill scheduler 會優先補 coverage=0 的 rows,再補不足 3 個
可用候選的低覆蓋 rows;失敗的 refill 也完整計入 generation budget,避免單一壞段無限重試。
完成 refill 或觸及任一硬預算後才建立 ragged DP,不要求每個 chunk 有相同候選數。
已通過 local hard gate 的 chunk 可直接進入 DP;唯一 rejection reason 是
boundary_speaker_drop、且 drop 不超過 0.15 的 chunk 也可作為 DP-only 候選。Single-chunk
request 可使用相同的 local proxy,但仍必須重新通過 exact whole-waveform 0.105/0.095 release guard。
這個 boundary-only relaxation 不適用於 semantic、pace、speaker similarity、缺少 evidence 或 whole-trajectory
accept,也不能在缺少整段 final verifier 時啟用。Local score 加上相鄰 ECAPA/RMS transition
cost 會產生最多 3 條 culprit-diverse 完整路徑;除了最低總成本,後續 exact checks 會優先選擇
在初始失敗 chunks 上使用不同候選的路徑,避免三次驗證都只改動已穩定段。每條都用同一個 production assembler 和整段
semantic/pace/speaker gate 驗證;整段仍必須通過 similarity 0.105 與 boundary drop 0.095,
第一條通過才回傳。三條都失敗、callback 異常或沒有完整
finite path 時一律 fail closed。
被拒的 joined/final log 只記 normalized CER、prefix/suffix CER、tail units 與 reasons,不記逐字
transcript。
候選排序與 DP 保留低成本的 windowed speaker evidence;joined candidate、k-best path 與
最終 waveform 會再跑一次與獨立發布評估定義相同的 ECAPA guard:原始 sample rate 上以
25/10 ms RMS 取 active 外緣,將整段與三個不重疊 thirds 分開 encode。這避免短音檔的
1.5 秒首尾視窗重疊而掩蓋句尾 speaker drift。
日期、24 小時制時間、百分比、常見單位與大寫 acronym/model code 會先轉成保守的
zh-TW spoken form,例如 2026/07/16、15:30、12.5% 與 10 km。ASR 會先統一
繁簡字形再評分,避免把正確的台灣華語輸出誤判為內容錯誤。
Email 與一般 URL fallback 會以混合式可辨識讀法展開:一般 local/domain/path label 保留
lexical ASCII(例如 tour、help),scheme、全大寫 atom 與短 suffix 使用明確 ASCII letter
tokens,數字逐位朗讀,分隔符也明確朗讀(例如 .tw 讀成「點、T、W」)。Email contract
不受下述 URL naturalization 影響。
另有一個刻意受限、可逆的 URL grammar:只接受 canonical lowercase http/https、
lowercase ASCII DNS host、至少一個能由固定 lexicon 唯一分詞的 compound label、無
userinfo/port/query/fragment,以及可選的單一 1–4 個 lowercase ASCII 字母 path;naturalized
spoken form 本身也必須不超過 36 public speech units。https 使用「安全網址是」,
http 使用不同的「網址是」;compound host 會自然分詞,dot 讀「點」,短 suffix 與 path
逐字母讀,最後明說「網址輸入完畢」。例如
https://coastwatch.example.tw/tide 會成為
安全網址是 coast watch 點 example 點 T W,路徑是 T I D E,網址輸入完畢。
大小寫、長 path、query、port、無唯一 compound 分詞或超限 URL 一律回到原本的 explicit
frontend,不作有損猜測。
Naturalized URL 每一個 scheme/host/dot/path/terminator atom 都保存原始 raw slice 與 emitted
spoken range;planner、Space runtime 與 local ASR 入口都會從 raw identifier 重新 render 並要求
proof 完全相等。若 nested proof 被移除、range/rule/reordering 被修改,或非 naturalized identifier
夾帶這組 provenance,會 fail closed。這個 naturalized phrase 在 generation planner 中是一個
不可切 component;H07 因此是 35 public units 的單一 network chunk,沒有 artificial
network_internal seam。兩個中文逗號只是模型 pause cue,Whisper 省略它們仍可通過;scheme cue、
host、path 與 terminator 仍是 exact lexical protected range。
完整 identifier 仍保留為 joined、full-large-v3 與 final ASR 的 exact protected target。完整
URL/email spoken proof 不超過 36 units 時,generation planner 會保持單一 atomic chunk,不在
:// 或 @ 中間製造人工 speaker boundary;URL 後的逗號也留在同一 chunk。只有 proof 超過
36 units 時,才在由原始 ASCII grammar 證明的 scheme/domain/path/query/email component 邊界切段,
並維持 32 units 目標、36 units 硬上限與 8 units network minimum。過長 email 仍可在 小老鼠
之前、過長 explicit URL 仍可在 冒號 斜線 斜線 之後切開;像 a@b.co、https://a.tw 的某一側
不足 8 units 且第一輪無解時,才放寬該短 identifier 的 preferred cut。Identifier 內部
需要拆分時,Email 第一輪必須在 小老鼠、explicit URL 第一輪必須在 scheme 結尾的
grammar boundary 切開;第一輪 DP 不得跨越這兩類 mandatory cut。
一般 semantic boundary 保留原標點停頓,但下限為 250 ms,讓 waveform-only verifier
可以將已通過 local gate 的 chunk 穩定分開。具共享 component proof 的
network_internal seam 仍插入 400 ms zero silence,並使用 5 ms fade/crossfade。
Component proof 無法完整
重建、非 ASCII IRI 或單一不可拆 component 超限時會 fail closed。
Local fragment gate 只接受 range-bound exact proof,並檢查所有最佳 edit alignment;重複的普通文字
不能借用受保護片段的 proof。Joined、whole 與 final gate 不套用這項 local-only canonicalization。
這能降低模型把不常見 TLD 自動補成 .com 的風險;一般英文句子不會套用這個規則。
URL 與後續英文 prose 應以空白或中文標點分隔;未分隔的 RFC path punctuation 會視為 URL
本身的一部分並納入 exact gate。Quoted email local-part 暫不支援,輸入時會直接 fail closed。
URL identifier 目前只支援 ASCII;非 ASCII IRI 必須先轉成 ASCII/percent-encoded 形式,
否則會因字母讀音碰撞而在 normalization 前 fail closed。
為避免 Whisper 自動句末標點和 URL 內容不可判定,URL 不接受以 . , ! ? ; : ' 結尾;
需要這些尾端符號時請使用 percent encoding。
Whisper 常見的 15点30分/15點30分 會視為同一讀法;裸寫 15點30 只有在 target
明確含同一個完整鐘點讀法時才作 ASR-only canonicalization。比分、秒數、重量、裸 15點、
非法時間與真實漏字仍會被拒絕。Whole-output ASR 只使用 waveform evidence:若所有符合
250 ms、兩側皆有 active speech 的 pause atoms 都能將音訊限制在 28 秒內,就保留
每一個 atom 獨立驗證(包含 semantic 250 ms 與 proof-bound 400 ms network seam)。只有
某個 atom 仍超過 28 秒時,才在相同 qualified pauses 上使用 coarse latest-pause fallback;
超過 12 個 atoms 或找不到安全 pause 都 fail closed,不使用最低能量或
target-dependent fallback。每段至少 1.25 秒、硬上限 30 秒,每次驗證最多 12 段,
並以最多 6 段的 microbatch 執行 ASR。
整句 intelligibility 仍使用 orthographic CER 上限 0.20。首尾、刪字與額外尾詞則使用固定
pypinyin==0.55.0 的 lexical-tone units 做完整 Levenshtein alignment:例如 陶藝/陶逸
可視為同一聲學讀法,陶藝/陶一 與 改到/改造 仍會被拒。ASCII network identifier 保持
逐字元比較,numeric/frequency contradiction 也仍 fail closed。這項 acoustic endpoint
normalization 不會修改送入模型的文字;既有同音代詞 canonicalization 仍保留。
只有尚未 assemble 的 internal local chunk 允許首、尾 6 units 各最多一個 substitution,且
deletion 必須為 0;由原始 GenerationChunkSpec 證明的 request 第一段 prefix 與最後一段
suffix 仍必須 exact。Refill 依原始絕對 chunk provenance 套用相同角色。Joined turbo、
independent full-large-v3 與 final whole-waveform gate 全部維持 exact endpoint。
自然 stop 在目前 checkpoint 上仍可能過快或錯過句尾。Demo 不再用播放目標語速決定生成長度;
模型使用原生語速的安全上限,完成後才做保音高語速校正。缺少句末標點時只在模型輸入補上句號,
不超過 6 個 speech units 的短句會使用最低 CFG 3.0;stop threshold 只在預估進度 75% 後逐步降低,並於 95% 進度降至 0.05,以捕捉句末弱 stop 訊號而不影響前段內容。
80 字內的 request 不再為了 onset 額外切段;長文才依自然標點與 80 字上限切段,避免人工
endpoint、孤立引號與不必要的局部 prefix/suffix gate。Release
profile 鎖定 primary CFG 3.0 / alternate CFG 2.0 交錯候選排程、NFE 10、
後處理語速 1.0。
Online pace gate 與 1.5 秒 speaker eligibility 使用和獨立 hosted evaluator 相同的 active-frame
interval union(25 ms frame、10 ms hop、peak -35 dB、absolute RMS floor 1e-4),因此句內或插入的
pause 不會稀釋 CPS。逐 active interval 的 stretcher 曾在固定 root ablation 造成長文 availability
回退,因此仍不切開 active intervals;production 會先在 model 的 raw completed waveform 上合併
total-duration 與 active-union 的 bounded rate,再量測一次 probe waveform。若 probe active pace
仍高於 3.95,會把 residual rate 乘回原 combined rate,並從 untouched model waveform 重新 render;
最後交給 verifier 與使用者的 samples 因此仍只經過一層 librosa phase vocoder。一般文字
固定 n_fft=1536、hop_length=384,URL/email code-switch chunk 使用 2048/512 保留較細的
network lexical distinction,總 stretch rate 不低於 0.80。這替最高 1.05 playback speed 與
chunk join 保留 4.3 CPS gate 的安全餘量;任何仍過快或被 stretch 破壞的候選照常由後續完整 gate
fail closed。Median-F0 只對實際具備 DP coverage eligibility 的 multi-chunk rows 計算,single/K32、
semantic reject、joined 與 final waveform 不再支付無用的 pYIN latency。
Generation work 以實際產生的 chunks 與其 speech units 累計:初始 trajectory 花費
chunk count 與全文 speech units;每次 refill 再加 1 chunk 與該段 units。任一 request 最多
32 generated TTS chunks、800 generated speech units,且 refill 只跑固定 NFE 10。
因此長文不再為了取得某一段替代候選而重生整篇,預算可直接集中在 ASR/speaker gate 顯示的
zero/low-coverage rows。若某 row 在預算內仍為 zero coverage,或 ragged graph 沒有 finite path,
服務立即 fail closed,不會放寬 semantic、speaker、pace 或 endpoint gate。
joined/final ECAPA 會從 1.48 秒 active duration 提前觸發,吸收 WAV round-trip 可能造成的
一至兩個 RMS frame 邊界位移;正式 release speaker eligibility 仍固定為 1.50 秒。
若所有候選仍有漏字、多說字、句尾截斷、語速過快或可量測的音色漂移,服務會拒絕輸出,
不會回傳任意 fallback。極短音檔的分段 ECAPA 不可靠,因此只採嚴格語意 gate;SQUIM hard gate
仍保留,但低於 1.50 秒不以 duration-biased preferred tier 強迫 refill 到 32-candidate 上限,並在
release audit 中獨立標記為 speaker-unverified;這不代表模型的短句音色已被根治。
目前部署路徑不使用 reference-style distribution score;固定 speaker centroid 只代表 identity, 不宣稱複製細緻情緒或韻律。Sequence DP 會使用相鄰候選的 ECAPA、active RMS 與 normalized median-F0 軟成本來降低跨句音色、響度與基準音高跳變;任何選中路徑仍必須重新通過整段 semantic、speaker、pace 與 endpoint hard gates。
請只使用已取得授權的參考音檔。合成語音僅供研究與評估展示,正式使用前請人工檢視。
程式碼與本機使用方式:OpenFormosa/BlueMagpie-TTS