DoDuyTTS

DoDuyTTS is a massively multilingual zero-shot text-to-speech (TTS) model supporting 600+ languages, with voice cloning and voice design. Built on a diffusion language model-style architecture with fast inference (RTF as low as 0.025 on GPU).

Use it with the doduytts package:

from doduytts import DoDuyTTS
import soundfile as sf
import torch

model = DoDuyTTS.from_pretrained(
    "doduy1911/dodduyTTS",
    device_map="cuda:0",   # or "mps", "cpu"
    dtype=torch.float16,
)

# Voice cloning
audio = model.generate(
    text="Hello, this is a test of zero-shot voice cloning.",
    ref_audio="ref.wav",
    ref_text="Transcription of the reference audio.",
)

# Voice design (no reference audio)
audio = model.generate(
    text="Hello!",
    instruct="female, low pitch, british accent",
)

sf.write("out.wav", audio[0], 24000)

A streaming TTS server (incremental sentence chunking from LLM text streams + dynamic micro-batching) ships with the package: see doduytts-serve.

Attribution

This checkpoint is based on k2-fsa/OmniVoice by the Xiaomi AI Lab Next-gen Kaldi team, redistributed under the Apache License 2.0. All credit for the model architecture and training goes to the original authors.

Disclaimer

Users are strictly prohibited from using this model for unauthorized voice cloning, voice impersonation, fraud, scams, or any other illegal or unethical activities. All users shall ensure full compliance with applicable local laws, regulations, and ethical standards.

Downloads last month
10
Safetensors
Model size
0.6B params
Tensor type
I64
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support