Kasanoma TTS v0.4 — Asante Twi + English + code-switch voice

CosyVoice3-0.5B (Apache-2.0) fine-tune for Ghanaian voice AI (Neriqlabs). v0.4 adds the WAXAL twi_tts studio split (google/WaxalNLP, CC-BY-4.0 — the cleanest Akan TTS audio that exists) to the clean-license training mix, which sharply improves English while keeping the Twi gains.

Eval — round-trip ASR-WER via TubaSTT v0.5 (lower=better) + UTMOSv2 naturalness (higher=better)

Category v0.1 WER v0.3 WER v0.4 WER v0.4 UTMOS
Asante Twi 26.7 20.6 18.5 3.02
English 23.7 26.6 18.9 3.18
Twi–English code-switch 57.2 46.1 47.8 3.06

Beats the original v0.1 on all three (Twi −8.2, English −4.8, code-switch −9.4) and the prior v0.3 decisively on English (−7.7 WER, +0.19 naturalness) while improving Twi. Round-trip WER is a relative intelligibility signal (it bakes in the ASR's own error); UTMOSv2 is English-trained so cross-lingual MOS is relative. Native MOS calibration pending.

Training data (all commercially clean-license)

BibleTTS Asante (CC-BY-SA) + Ashesi Financial-Inclusion (CC-BY) + Common Voice Twi (CC0) + WAXAL twi_tts studio (CC-BY-4.0). 36,926 utterances / 66 speakers. Only the llm (text→speech-token) stage is fine-tuned; flow-matching + HiFi-GAN vocoder transfer from the base. NFC-normalized, ɛ (U+025B) / ɔ (U+0254) preserved.

Usage

from cosyvoice.cli.cosyvoice import CosyVoice3
cv = CosyVoice3("kasanoma-tts-twi-v0.4", load_trt=False, fp16=False)
prompt_wav = "ref_voice.wav"      # 16 kHz reference for the voice to clone
text = "You are a helpful assistant.<|endofprompt|>me PIN no reset, please help me"
for out in cv.inference_cross_lingual(text, prompt_wav, stream=False, text_frontend=False):
    audio = out["tts_speech"]     # 24 kHz

License

Weights CC-BY-SA-4.0 (inherited from BibleTTS, most restrictive source). Backbone CosyVoice3 is Apache-2.0. Built by Neriqlabs (founder: Samson Nkrumah) for Ghanaian-language voice AI.

Downloads last month
48
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for neriqlabs/kasanoma-tts-twi-v0.4

Quantized
(22)
this model