- XTTS-v2 Hausa Conversational (Common Voice Domain-Adapted)
- Model Summary
- Conversational Adaptation Dashboard & Convergence
- Acoustic Improvements Over Stage 1
- Audio Comparison: Original vs. Stage 1 vs. Stage 2 Cloned Synthesis
- Sample 1 (Conversational Narrative): "Wani abu ya kama gaban jacob."
- Sample 2 (Community & Indigenes): "Na ci kwai da yawa a ranar bikin Easter Sunday."
- Sample 3 (Civic & Leadership): "Yunwa wani abu ne da jiki kan nuna yayin da yake buƙatar abinci."
- Sample 4 (Cultural & Heritage): "Ko zan iya fara canka."
- Sample 5 (Expressive Spoken Idiom): "Da gaske muke."
- 🔁 Stage 1 OpenBible Benchmark Samples: Foundation vs. Conversational Comparison
- Repository Files
- How to Use
- Limitations and Ethical Usage
- Author & Attribution
- Acknowledgments
- Model Summary
XTTS-v2 Hausa Conversational (Common Voice Domain-Adapted)
This repository provides fine-tuned weights, extended tokenizer vocabulary, and configuration files for Coqui XTTS-v2 adapted for conversational Hausa (ha).
This model represents Stage 2 in a two-stage training strategy:
- Stage 1 (Foundation): samuelolubukun/xtts-v2-hausa-openbible — 8-epoch foundation model trained on high-fidelity OpenBible audio to master Hausa orthography, hook letters (
ɓ, ɗ, ƙ, ƴ), and stable phonetics. - Stage 2 (Conversational Adaptation — This Model): Adapted from the Stage 1 checkpoint on benjaminogbonna/nigerian_common_voice_dataset (Hausa split) using conservative learning rates (
1e-6) to learn natural spoken cadence, informal pauses, and diverse speaker acoustic environments without catastrophic forgetting.
Model Summary
- Base Architecture: Coqui XTTS-v2 (Autoregressive GPT text/mel encoder + Discrete VAE spectrogram vocoder)
- Base Model Checkpoint: samuelolubukun/xtts-v2-hausa-openbible (31,360 steps)
- Parameters: 521.8M
- Target Language: Hausa (
ha) - Tokenizer: Extended Byte Pair Encoding (BPE) with special language tag
[ha](extended vocabulary: 8,320 tokens) - Adaptation Dataset: benjaminogbonna/nigerian_common_voice_dataset (Hausa split)
- Filtered Training Set: 3,795 clean conversational audio samples (filtered for length 1.5s–12s and valid text)
- Adaptation Steps: 7,592 optimization steps (4 epochs, lr=1.5e-6, grad_accum=4, batch_size=2)
- Hardware: NVIDIA A10G (24 GB VRAM)
Conversational Adaptation Dashboard & Convergence
The visualization below highlights the multidimensional shift from the Stage 1 OpenBible foundation model to this Stage 2 Common Voice conversational model across prosody, pitch dynamics, acoustic loss, and speech continuity:
Training Progression Summary
| Stage / Step | Global Step | Train Loss | Text CE Loss | Mel Spectrogram Loss | Notes |
|---|---|---|---|---|---|
| Adaptation Start | 0 | 0.9753 | 0.0358 | 3.8655 | Initialized directly from OpenBible 31,360-step foundation weights |
| Epoch 1 Complete | 1,898 | 0.8105 | 0.0368 | 3.2052 | Rapid acoustic convergence on crowdsourced Hausa voices |
| Epoch 2 Complete | 3,796 | 0.7707 | 0.0351 | 3.0477 | 50% adaptation milestone; significant decrease in scripted reading cadence |
| Epoch 3 Complete | 5,694 | 0.7424 | 0.0344 | 2.9402 | Pitch and breath boundaries align with natural colloquial speech |
| Epoch 4 Final | 7,592 | 0.7247 | 0.0344 | 2.8578 | Fully adapted conversational model (checkpoint_7500.pth) |
| Eval Epoch 4 | 7,592 | 2.8852 | 0.0317 | 2.8535 | Lowest validation spectrogram loss across all epochs |
Acoustic Improvements Over Stage 1
Objective acoustic and linguistic evaluation across unseen Common Voice test speakers demonstrates significant improvements:
| Metric | Stage 1 (OpenBible Foundation) | Stage 2 (Conversational Adapted - 4 Epochs) | Improvement / Impact |
|---|---|---|---|
| Silence / Pause Ratio | 24.2% | 18.2% | -6.0%: Eliminates stilted pauses; natural phrase flow |
| Phrasing Cohesion | 5 rigid micro-pauses | 2 natural breath boundaries | Seamless co-articulation across compound clauses |
| Pitch Dynamism (StdDev) | 50.6 Hz | 55.8 Hz | +5.2 Hz: Expressive tonal variations across Hausa questions & idioms |
| Vocal Timbre Fidelity | 65.0% | 94.0% | +44%: Zero-shot voice cloning matches target speaker without narrator bias |
Audio Comparison: Original vs. Stage 1 vs. Stage 2 Cloned Synthesis
Below is a side-by-side comparison between the unseen reference speakers and the speech synthesized by the Stage 1 foundation model versus this Stage 2 conversational adapted model.
All reference audios are sourced from benjaminogbonna/nigerian_common_voice_dataset (Hausa split) by Benjamin Ogbonna:
Sample 1 (Conversational Narrative): "Wani abu ya kama gaban jacob."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 2 (Community & Indigenes): "Na ci kwai da yawa a ranar bikin Easter Sunday."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 3 (Civic & Leadership): "Yunwa wani abu ne da jiki kan nuna yayin da yake buƙatar abinci."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 4 (Cultural & Heritage): "Ko zan iya fara canka."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 5 (Expressive Spoken Idiom): "Da gaske muke."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
🔁 Stage 1 OpenBible Benchmark Samples: Foundation vs. Conversational Comparison
These are the exact benchmark samples originally evaluated on the samuelolubukun/xtts-v2-hausa-openbible Stage 1 foundation model repository. Below, the same reference audio clips and prompt texts are synthesized side-by-side using the Stage 2 Conversational fine-tuned model to directly compare the change in cadence, natural pacing, and conversational tone:
Stage 1 Benchmark Sample 1: "Bude kofar. Na san kina ciki."
| Column 1: Original Reference Audio | Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) | Column 3: XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Stage 1 Benchmark Sample 2: "Na ga wani aikin dabba mai ban mamaki a circus"
| Column 1: Original Reference Audio | Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) | Column 3: XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Stage 1 Benchmark Sample 3: "Nagode da yadda ka ƙona mini riga ta da sigarin ka."
| Column 1: Original Reference Audio | Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) | Column 3: XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Repository Files
This repository contains all files necessary to run standalone synthesis or merge with standard Coqui pipelines:
| File | Description |
|---|---|
model.pth |
Conversational fine-tuned GPT decoder weights (checkpoint_7500.pth, 4 Epochs) |
config.json |
Updated model configuration with ha language enabled |
vocab.json |
Extended BPE tokenizer vocabulary (8,320 tokens) |
dvae.pth |
Discrete VAE acoustic spectrogram decoder |
mel_stats.pth |
Mel spectrogram normalization statistics |
training_curves.png |
High-resolution 4-panel conversational adaptation visual dashboard across 7,592 steps |
samples/ |
5-way zero-shot voice cloning audio demonstrations |
How to Use
Ensure coqui-tts and dependencies are installed:
pip install coqui-tts "transformers<4.45.0" torch torchaudio huggingface_hub
Load the model and synthesize speech directly from this Hugging Face repository:
import os
import re
import torch
import torchaudio
from huggingface_hub import snapshot_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts
# 1. Download model repository
model_dir = snapshot_download(repo_id="samuelolubukun/xtts-v2-hausa-openbible-conversational")
# 2. Load model config and checkpoint
config = XttsConfig()
config.load_json(os.path.join(model_dir, "config.json"))
model = Xtts.init_from_config(config)
model.load_checkpoint(
config,
checkpoint_path=os.path.join(model_dir, "model.pth"),
vocab_path=os.path.join(model_dir, "vocab.json"),
use_deepspeed=False
)
model.cuda() if torch.cuda.is_available() else model.cpu()
# 3. Text cleaner fallback for custom language support
orig_preprocess = model.tokenizer.preprocess_text
def safe_preprocess_text(txt, lang):
try:
return orig_preprocess(txt, lang)
except NotImplementedError:
txt = txt.lower()
return re.sub(r"\s+", " ", txt).strip()
model.tokenizer.preprocess_text = safe_preprocess_text
# 4. Compute speaker conditioning latents (using bundled reference audio or your own 3-10s WAV)
ref_audio = os.path.join(model_dir, "samples/sample1_ref.wav")
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
audio_path=ref_audio,
gpt_cond_len=30,
max_ref_length=30,
sound_norm_refs=False,
)
# 5. Synthesize speech
text = "A wannan lokacin banbancin tsakanin mutum da dabba ba daɗi ba ragi illa suttura."
out = model.inference(
text=text,
language="ha",
gpt_cond_latent=gpt_cond_latent,
speaker_embedding=speaker_embedding,
temperature=0.3,
length_penalty=1.0,
repetition_penalty=5.0,
top_k=30,
top_p=0.85,
)
# 6. Save output audio (24,000 Hz)
torchaudio.save("output_ha.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)
print("Saved synthesized speech to output_ha.wav")
Limitations and Ethical Usage
- Domain & Conversational Acoustic Artifacts: While Stage 2 removes scripture-reading bias and significantly improves informal prosody, Common Voice recordings contain diverse crowd-sourced microphones, varying gain levels, and background ambient noise. Very noisy speaker prompts may slightly affect synthesized clarity.
- Glottal and Hooked Letter Sensitivity: Hausa relies heavily on hooked characters (
ɓ,ɗ,ƙ,ƴ) and apostrophes for glottal stops. Text fed into the model should preserve proper standard Hausa orthography for optimal pronunciation. - Ethical Voice Cloning: This model is intended for research, linguistic preservation, and accessibility in African languages. Explicit consent should be obtained prior to cloning recognizable individual voices.
- Licensing: Inherits the Coqui Public Model License / MRCUL.
Author & Attribution
Developed and fine-tuned by Samuel Olubukun:
- Website: samuelolubukun.com
- GitHub: @samolubukun
- Hugging Face: @samuelolubukun
Citation
If you use this model or checkpoints in your research, applications, or benchmarks, please cite:
@misc{olubukun2026xttshausaconversational,
author = {Samuel Olubukun},
title = {XTTS-v2 Hausa Conversational: Domain Adaptation on Nigerian Common Voice},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/samuelolubukun/xtts-v2-hausa-conversational}}
}
Acknowledgments
- Base Model: Coqui XTTS-v2
- Foundation Model: samuelolubukun/xtts-v2-hausa-openbible
- Dataset: Benjamin Ogbonna for curating and releasing benjaminogbonna/nigerian_common_voice_dataset, and the Mozilla Common Voice contributors.
- Downloads last month
- 111
