- XTTS-v2 Igbo Conversational (Common Voice Domain-Adapted)
- Model Summary
- Conversational Adaptation Dashboard & Convergence
- Acoustic Improvements Over Stage 1
- Audio Comparison: Original vs. Stage 1 vs. Stage 2 Cloned Synthesis
- Sample 1 (Cultural Festival): "kọwàrà mmemme iri ji ọhụrụ dịka oke omenala n'ala Igbo."
- Sample 2 (National Recognition): "ma mee ya ka ọ gbagotekwuo n'ogoogo ịbụ emume a mà àmá n'ala Nigeria."
- Sample 3 (Expressive & Reflective): "Ọ kpọpụtàsịrị ihe dị iche iche nke ahụ pụrụ ibute."
- Sample 4 (Everyday Descriptive): "Nwata nwagbọghọ a mara mma ma nwekwaa ezigbo agwa."
- Sample 5 (Civic & Education): "A kọwaala ọba akwụkwọ dịka ụlọọrụ dị mkpà n'ịkwàlite mmepe obodo."
- 🔁 Stage 1 OpenBible Benchmark Samples: Foundation vs. Conversational Comparison
- Repository Files
- How to Use
- Limitations and Ethical Usage
- Author & Attribution
- Acknowledgments
- Model Summary
XTTS-v2 Igbo Conversational (Common Voice Domain-Adapted)
This repository provides fine-tuned weights, extended tokenizer vocabulary, and configuration files for Coqui XTTS-v2 adapted for conversational Igbo (ig).
This model represents Stage 2 in a two-stage training strategy:
- Stage 1 (Foundation): samuelolubukun/xtts-v2-igbo-openbible — 8-epoch foundation model trained on high-fidelity OpenBible audio to master Igbo orthography, sub-dots (
ị, ọ, ụ), nasalñ, and stable phonetics. - Stage 2 (Conversational Adaptation — This Model): Adapted from the Stage 1 checkpoint on benjaminogbonna/nigerian_common_voice_dataset (Igbo split) across 4 full epochs using a conservative learning rate (
1.5e-6) to break the rigid, ecclesiastical narration register and learn natural conversational speech rates, downstep pitch contours, and diverse speaker timbres without catastrophic forgetting.
Model Summary
- Base Architecture: Coqui XTTS-v2 (Autoregressive GPT text/mel encoder + Discrete VAE spectrogram vocoder)
- Base Model Checkpoint: samuelolubukun/xtts-v2-igbo-openbible (31,360 steps)
- Parameters: 521.8M
- Target Language: Igbo (
ig) - Tokenizer: Extended Byte Pair Encoding (BPE) with special language tag
[ig](extended vocabulary: 11,533 tokens) - Adaptation Dataset: benjaminogbonna/nigerian_common_voice_dataset (Igbo split)
- Filtered Training Set: 2,809 clean conversational audio samples (filtered for length 1.5s–12s and valid text)
- Adaptation Steps: 7,488 optimization steps (4 epochs, lr=1.5e-6, grad_accum=4, batch_size=2)
- Hardware: NVIDIA A10G (24 GB VRAM)
Conversational Adaptation Dashboard & Convergence
The visualization below highlights the multidimensional shift from the Stage 1 OpenBible foundation model to this Stage 2 Common Voice conversational model across speaking rate alignment, pitch tracking, acoustic loss, and timbre fidelity:
Training Progression Summary
| Stage / Step | Global Step | Train Loss | Text CE Loss | Mel Spectrogram Loss | Notes |
|---|---|---|---|---|---|
| Adaptation Start | 0 | 0.9812 | 0.0425 | 3.8210 | Initialized directly from OpenBible 31,360-step foundation weights |
| Epoch 2 Midpoint | 2,100 | 0.8760 | 0.0421 | 3.4620 | Mel loss drops steadily as conversational acoustics take hold |
| Epoch 3 Midpoint | 4,200 | 0.8607 | 0.0411 | 3.4019 | Consistent descent; text loss stable with zero forgetting |
| Adaptation Final | 7,488 | 0.8487 | 0.0384 | 3.3547 | Mastered human conversational pace and downstep contours |
Acoustic Improvements Over Stage 1
Objective acoustic and linguistic evaluation across unseen Common Voice test speakers demonstrates significant improvements:
| Metric | Stage 1 (OpenBible Foundation) | Stage 2 (Conversational Adapted - This Model) | Improvement / Impact |
|---|---|---|---|
| Human Speaking Rate Match | Rushed / abrupt on compounds (4.78s) | 6.36s (Matches Human Target: 6.59s) | Natural conversational phrasing matching human breath groups |
| Female Speaker Timbre Match | 197.8 Hz (biased to male narrator) | 233.7 Hz (Matches Female Target: 251.5 Hz) | +35.9 Hz: Clones real female pitch without deep voice bleeding |
| Pitch Dynamic Range (StdDev) | 40.5 Hz | 67.6 Hz | +66.9%: Dynamic tonal expressiveness on Igbo downstep contours |
| Vocal Timbre Match Score | 68.0% | 93.0% | +36% improvement in zero-shot voice cloning fidelity |
Audio Comparison: Original vs. Stage 1 vs. Stage 2 Cloned Synthesis
Below is a side-by-side comparison between the unseen reference speakers and the speech synthesized by the Stage 1 foundation model versus this Stage 2 conversational adapted model.
All reference audios are sourced from benjaminogbonna/nigerian_common_voice_dataset (Igbo split) by Benjamin Ogbonna:
Sample 1 (Cultural Festival): "kọwàrà mmemme iri ji ọhụrụ dịka oke omenala n'ala Igbo."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 2 (National Recognition): "ma mee ya ka ọ gbagotekwuo n'ogoogo ịbụ emume a mà àmá n'ala Nigeria."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 3 (Expressive & Reflective): "Ọ kpọpụtàsịrị ihe dị iche iche nke ahụ pụrụ ibute."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 4 (Everyday Descriptive): "Nwata nwagbọghọ a mara mma ma nwekwaa ezigbo agwa."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Sample 5 (Civic & Education): "A kọwaala ọba akwụkwọ dịka ụlọọrụ dị mkpà n'ịkwàlite mmepe obodo."
| Original Reference (Common Voice) | XTTS-v2 Stage 1 (OpenBible Foundation) | XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
🔁 Stage 1 OpenBible Benchmark Samples: Foundation vs. Conversational Comparison
These are the exact benchmark samples originally evaluated on the samuelolubukun/xtts-v2-igbo-openbible Stage 1 foundation model repository. Below, the same reference audio clips and prompt texts are synthesized side-by-side using the Stage 2 Conversational fine-tuned model to directly compare the change in cadence, natural pacing, and conversational tone:
Stage 1 Benchmark Sample 1: "mbize kwesiri ka e lebàra ya anya mgwamgwa ugbua"
| Column 1: Original Reference Audio | Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) | Column 3: XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Stage 1 Benchmark Sample 2: "Amụrụ ya na Jos, steetị Naijiria."
| Column 1: Original Reference Audio | Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) | Column 3: XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Stage 1 Benchmark Sample 3: "nke onye nọchitere anya ya bụ onyeisi ọrụ oyibo na steeti ahụ"
| Column 1: Original Reference Audio | Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) | Column 3: XTTS-v2 Stage 2 (Conversational Adapted) |
|---|---|---|
Repository Files
This repository contains all files necessary to run standalone synthesis or merge with standard Coqui pipelines:
| File | Description |
|---|---|
model.pth |
Conversational fine-tuned GPT decoder weights (checkpoint_7200.pth) |
config.json |
Updated model configuration with ig language enabled |
vocab.json |
Extended BPE tokenizer vocabulary (11,533 tokens) |
dvae.pth |
Discrete VAE acoustic spectrogram decoder |
mel_stats.pth |
Mel spectrogram normalization statistics |
training_curves.png |
High-resolution 4-panel conversational adaptation dashboard |
samples/ |
5-way zero-shot voice cloning audio demonstrations |
How to Use
Ensure coqui-tts and dependencies are installed:
pip install coqui-tts "transformers<4.45.0" torch torchaudio huggingface_hub
Load the model and synthesize speech directly from this Hugging Face repository:
import os
import re
import torch
import torchaudio
from huggingface_hub import snapshot_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts
# 1. Download model repository
model_dir = snapshot_download(repo_id="samuelolubukun/xtts-v2-igbo-openbible-conversational")
# 2. Load model config and checkpoint
config = XttsConfig()
config.load_json(os.path.join(model_dir, "config.json"))
model = Xtts.init_from_config(config)
model.load_checkpoint(
config,
checkpoint_path=os.path.join(model_dir, "model.pth"),
vocab_path=os.path.join(model_dir, "vocab.json"),
use_deepspeed=False
)
model.cuda() if torch.cuda.is_available() else model.cpu()
# 3. Text cleaner fallback for custom language support
orig_preprocess = model.tokenizer.preprocess_text
def safe_preprocess_text(txt, lang):
try:
return orig_preprocess(txt, lang)
except NotImplementedError:
txt = txt.lower()
return re.sub(r"\s+", " ", txt).strip()
model.tokenizer.preprocess_text = safe_preprocess_text
# 4. Compute speaker conditioning latents (using bundled reference audio or your own 3-10s WAV)
ref_audio = os.path.join(model_dir, "samples/sample1_ref.wav")
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
audio_path=ref_audio,
gpt_cond_len=30,
max_ref_length=30,
sound_norm_refs=False,
)
# 5. Synthesize speech
text = "Nnọọ, kedu ka ị mere? Anyị na-asụ asụsụ Igbo ugbu a n'ụzọ dị nro."
out = model.inference(
text=text,
language="ig",
gpt_cond_latent=gpt_cond_latent,
speaker_embedding=speaker_embedding,
temperature=0.3,
length_penalty=1.0,
repetition_penalty=5.0,
top_k=30,
top_p=0.85,
)
# 6. Save output audio (24,000 Hz)
torchaudio.save("output_ig.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)
print("Saved synthesized speech to output_ig.wav")
Limitations and Ethical Usage
- Conversational Acoustics vs. Studio Purity: While Stage 2 eliminates the solemn scripture register and unlocks flexible speaker timbres, Common Voice recordings contain diverse crowd-sourced microphones and acoustic environments. Reference audios with heavy background chatter or clipping should be avoided.
- Sub-dot and Diacritic Sensitivity: Igbo distinguishes vowel harmony sets using under-dots (
ị, ọ, ụ) and nasal consonants (ñ). Text fed into the model should preserve proper standard Igbo orthography for optimal pronunciation. - Ethical Voice Cloning: This model is intended for research, linguistic preservation, and accessibility in African languages. Explicit consent should be obtained prior to cloning recognizable individual voices.
- Licensing: Inherits the Coqui Public Model License / MRCUL.
Author & Attribution
Developed and fine-tuned by Samuel Olubukun:
- Website: samuelolubukun.com
- GitHub: @samolubukun
- Hugging Face: @samuelolubukun
Citation
If you use this model or checkpoints in your research, applications, or benchmarks, please cite:
@misc{olubukun2026xttsigboconversational,
author = {Samuel Olubukun},
title = {XTTS-v2 Igbo Conversational: 4-Epoch Domain Adaptation on Nigerian Common Voice},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/samuelolubukun/xtts-v2-igbo-conversational}}
}
Acknowledgments
- Base Model: Coqui XTTS-v2
- Foundation Model: samuelolubukun/xtts-v2-igbo-openbible
- Dataset: Benjamin Ogbonna for curating and releasing benjaminogbonna/nigerian_common_voice_dataset, and the Mozilla Common Voice contributors.
- Downloads last month
- 184
