--- title: OmniVoice Word-Control emoji: 🎛️ colorFrom: indigo colorTo: pink sdk: gradio sdk_version: 6.10.0 app_file: app.py pinned: false license: cc-by-nc-4.0 short_description: Voice cloning with word-level prosody control (WordVoice-5A) models: - multimodalart/omnivoice-word-control - k2-fsa/OmniVoice --- # OmniVoice Word-Control Zero-shot voice cloning with explicit **word-level control** over duration, boundary/pauses, energy, pitch, and tone contour — the [WordVoice](https://huggingface.co/papers/2607.06461) task realized on [OmniVoice](https://huggingface.co/k2-fsa/OmniVoice)'s masked-diffusion LM via inline control tokens, fine-tuned on [WordVoice-5A](https://huggingface.co/datasets/XXH333/WordVoice-5A) (English, ~2,138h). UI and inline-tag syntax adapted from [hugging-apps/wordvoice-tts](https://huggingface.co/spaces/hugging-apps/wordvoice-tts); generation flow adapted from the official [k2-fsa/OmniVoice](https://huggingface.co/spaces/k2-fsa/OmniVoice) Space. Example reference clip from the WordVoice-5A test split (CC-BY-4.0).