Instructions to use neriqlabs/kasanoma-tts-twi-v0.4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- CosyVoice
How to use neriqlabs/kasanoma-tts-twi-v0.4 with CosyVoice:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Kasanoma TTS v0.4 — Asante Twi + English + code-switch voice
CosyVoice3-0.5B (Apache-2.0) fine-tune for Ghanaian voice AI (Neriqlabs). v0.4 adds the
WAXAL twi_tts studio split (google/WaxalNLP, CC-BY-4.0 — the cleanest Akan TTS audio that
exists) to the clean-license training mix, which sharply improves English while keeping the Twi gains.
Eval — round-trip ASR-WER via TubaSTT v0.5 (lower=better) + UTMOSv2 naturalness (higher=better)
| Category | v0.1 WER | v0.3 WER | v0.4 WER | v0.4 UTMOS |
|---|---|---|---|---|
| Asante Twi | 26.7 | 20.6 | 18.5 | 3.02 |
| English | 23.7 | 26.6 | 18.9 | 3.18 |
| Twi–English code-switch | 57.2 | 46.1 | 47.8 | 3.06 |
Beats the original v0.1 on all three (Twi −8.2, English −4.8, code-switch −9.4) and the prior v0.3 decisively on English (−7.7 WER, +0.19 naturalness) while improving Twi. Round-trip WER is a relative intelligibility signal (it bakes in the ASR's own error); UTMOSv2 is English-trained so cross-lingual MOS is relative. Native MOS calibration pending.
Training data (all commercially clean-license)
BibleTTS Asante (CC-BY-SA) + Ashesi Financial-Inclusion (CC-BY) + Common Voice Twi (CC0) +
WAXAL twi_tts studio (CC-BY-4.0). 36,926 utterances / 66 speakers. Only the llm
(text→speech-token) stage is fine-tuned; flow-matching + HiFi-GAN vocoder transfer from the base.
NFC-normalized, ɛ (U+025B) / ɔ (U+0254) preserved.
Usage
from cosyvoice.cli.cosyvoice import CosyVoice3
cv = CosyVoice3("kasanoma-tts-twi-v0.4", load_trt=False, fp16=False)
prompt_wav = "ref_voice.wav" # 16 kHz reference for the voice to clone
text = "You are a helpful assistant.<|endofprompt|>me PIN no reset, please help me"
for out in cv.inference_cross_lingual(text, prompt_wav, stream=False, text_frontend=False):
audio = out["tts_speech"] # 24 kHz
License
Weights CC-BY-SA-4.0 (inherited from BibleTTS, most restrictive source). Backbone CosyVoice3 is Apache-2.0. Built by Neriqlabs (founder: Samson Nkrumah) for Ghanaian-language voice AI.
- Downloads last month
- 48
Model tree for neriqlabs/kasanoma-tts-twi-v0.4
Base model
FunAudioLLM/Fun-CosyVoice3-0.5B-2512