Add adjustable cfg_value / inference_timesteps sliders (shared across tabs, defaults = model-card recommendation) a8231fb verified voidful commited on Jul 2
Voice clone: single reference-clip mode only (drop centroid/ECAPA stable mode + speechbrain) 355336d verified voidful commited on Jul 2
Wording: neutral phrasing for reference-clip support boundary dc2d2f8 verified voidful commited on Jul 2
step_0006000: fast transcript-free clone (reference_wav_path), new default params (cfg 2.0/steps 10) e66cd92 verified voidful commited on Jul 2
Add speaker selection (李宏毅 + 女聲) via multi-speaker centroids 7b111a1 verified voidful commited on Jun 28
Voice clone: read 10 sentences -> average ECAPA over all clips for a stabler centroid 9a401f5 verified voidful commited on Jun 23
Voice cloning: extract ECAPA speaker centroid from upload (no transcript); add speechbrain 5fc2aef verified voidful commited on Jun 23
Fix voice cloning: use prompt_wav+prompt_text (reference_wav_path produced gibberish); remove debug 5e31159 verified voidful commited on Jun 23
TEMP: add /tts_debug endpoint to diagnose voice-clone content failure 7afaa46 verified voidful commited on Jun 23
Long-text streaming: stream until GPU time budget (~100s) instead of fixed 12-sentence cap a858c6e verified voidful commited on Jun 23
Streaming tab: sentence-by-sentence streaming for long text (synth & play per sentence) 01153e3 verified voidful commited on Jun 23
Streaming tab: collect full audio then play smoothly (ZeroGPU is sub-realtime) dd9fd54 verified voidful commited on Jun 23
Recreate Space (zero-a10g) after region corruption; streaming included 927c1a5 verified voidful commited on Jun 23