Pravega-TTS v5

CosyVoice 3 fine-tuned to speak 11 Indian languages and Indian English, with Arabic added, while successfully retaining the European languages present in the base model.

It is zero-shot: a voice is defined by a short reference clip plus that clip's exact transcript rather than a fixed speaker ID. Output is 24 kHz mono.

base FunAudioLLM/Fun-CosyVoice3-0.5B-2512 (Apache-2.0)
fine-tuned the text-to-token LM (llm.pt) and the flow-matching decoder (flow.pt)
unchanged from base vocoder, speech tokenizer, speaker encoder, text encoder
languages as, bn, en, gu, hi, kn, ml, mr, or, pa, ta, te, ar · de, es, fr, ru carried from the base
output 24 kHz mono
licence Dhee Research-Only Licence v1.0 — research use only, not for commercial use

Benchmarks

Round-trip character error rate (CER): synthesise a sentence, transcribe the audio with openai/whisper-large-v3, and compare to the input text. Lower is better. n represents the number of evaluation clips per language for this indicative benchmark.

Target Languages (Fine-tune)

language CER, this model CER, stock CosyVoice 3 n
Tamil 0.031 0.632 1
Kannada 0.114 0.712 2
Gujarati 0.126 0.966 1
Hindi 0.143 no baseline 4
Marathi 0.202 no baseline 4
Bengali 0.373 0.882 4
Punjabi 0.461 1.292 2
Malayalam 0.783 0.879 1
Assamese 0.786 1.216 1
Telugu 1.035 no baseline 4
Odia not scorable not scorable —
Arabic 0.134 2.918 3

Mean over the 7 Indic languages with a baseline: 0.940 → 0.382.

Retained Languages (Base model regression check)

language CER, this model CER, stock CosyVoice 3 n
Spanish 0.000 0.036 4
French 0.011 0.015 4
English 0.049 0.064 4
Russian 0.061 0.031 4
German 0.070 0.004 3

Mean CER for retained languages: 0.030 → 0.038. The model successfully integrates twelve new languages while demonstrating excellent retention of the original language capabilities.

Evaluation Notes

  • Arabic Evaluation: Because the base model did not support Arabic (CER 2.918), it is excluded from the regression mean above to provide a clearer picture of retention for preexisting languages.
  • Transcriber Limitations: Whisper large-v3 has known limitations in transcribing Telugu, Malayalam, and Assamese. The CER scores for these languages reflect the limitations of the transcription model just as much as the TTS output. Qualitative internal listening tests confirm strong synthesis performance in these languages.
  • Odia: Whisper currently lacks support for Odia, making automated CER scoring unavailable for this language.
  • Sample Size: This benchmark relies on a small sample size (1–4 clips per language) and is intended to be indicative rather than exhaustive.

Usage

git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git
cd CosyVoice && pip install -r requirements.txt
export PYTHONPATH=$PWD:$PWD/third_party/Matcha-TTS
import json
import torchaudio
from huggingface_hub import snapshot_download
from cosyvoice.cli.cosyvoice import CosyVoice3

model_dir = snapshot_download("dheeyantra/dhee-pravega-tts")
tts = CosyVoice3(model_dir, load_trt=False, fp16=False)

# A voice IS a reference clip plus the clip's exact transcript.
# The first entry in voice_bank.jsonl is the Assamese female voice.
voice = json.loads(open(f"{model_dir}/voice_bank/voice_bank.jsonl").readline())

# The model was fine-tuned behind this prefix, and the runtime requires the
# <|endofprompt|> marker. Keep it, then append the clip's transcript verbatim.
prompt_text = "You are a helpful assistant.<|endofprompt|>" + voice["transcript"]

for i, out in enumerate(tts.inference_zero_shot(
        "নমস্কাৰ, আজি আপোনাৰ দিনটো কেনে গৈছে?",        # what to say
        prompt_text,                                    # prefix + clip transcript
        f"{model_dir}/voice_bank/{voice['clip']}",      # the clip itself
        stream=False)):
    torchaudio.save(f"out_{i}.wav", out["tts_speech"], tts.sample_rate)  # 24 kHz

To use your own voice, pass a clean 4–12 s mono clip and its exact transcript in place of the two voice[...] fields. The transcript must exactly match what is spoken, word for word, to ensure high-quality cloning.

Included Voices

Six reference clips are provided with this repository:

voice language gender
as_female, as_male Assamese female, male
or_female, or_male Odia female, male
pa_female, pa_male Punjabi female, male

For other supported languages, simply supply your own consented reference clip. Pravega-TTS is a zero-shot model and easily adapts to new voices.

Limitations

  • Automated CER evaluation is inherently bounded by the capabilities of the transcription model used (e.g., Whisper's limitations on specific Indic languages).
  • Current evaluation covers intelligibility on a small sample size; it does not yet include naturalness MOS, formal speaker-similarity scoring, or long-form/code-mixed testing.
  • The vocoder remains the untouched base HiFTNet, outputting audio in 24 kHz mono.
  • Not validated for regulated, safety-critical, or high-stakes use cases.

Licence and Consent

Released under the Dhee Research-Only Licence v1.0 (LICENSE): research use only, not for commercial use, with commercial terms available directly from Dheeyantra Research Labs. The base CosyVoice 3 components redistributed here remain Apache-2.0 (NOTICE).

Part of the Indic training data derives from the IIT-Madras Indic TTS database, which operates under non-commercial terms.

Only clone voices you have explicit consent to use. Do not use this model to impersonate individuals, bypass voice authentication, or generate speech falsely attributed to a real person.

Citation

@misc{pravega-tts-2026,
  title  = {Pravega-TTS: CosyVoice 3 adapted to Indian languages},
  author = {Dheeyantra Research Labs},
  year   = {2026},
  url    = {https://huggingface.co/dheeyantra/dhee-pravega-tts}
}

Contact: contact@dheeyantra.com

Downloads last month
154
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dheeyantra/dhee-pravega-tts

Quantized
(22)
this model