XTTS-v2 Igbo Conversational (Common Voice Domain-Adapted)

This repository provides fine-tuned weights, extended tokenizer vocabulary, and configuration files for Coqui XTTS-v2 adapted for conversational Igbo (ig).

This model represents Stage 2 in a two-stage training strategy:

  1. Stage 1 (Foundation): samuelolubukun/xtts-v2-igbo-openbible — 8-epoch foundation model trained on high-fidelity OpenBible audio to master Igbo orthography, sub-dots (ị, ọ, ụ), nasal ñ, and stable phonetics.
  2. Stage 2 (Conversational Adaptation — This Model): Adapted from the Stage 1 checkpoint on benjaminogbonna/nigerian_common_voice_dataset (Igbo split) across 4 full epochs using a conservative learning rate (1.5e-6) to break the rigid, ecclesiastical narration register and learn natural conversational speech rates, downstep pitch contours, and diverse speaker timbres without catastrophic forgetting.

Model Summary

  • Base Architecture: Coqui XTTS-v2 (Autoregressive GPT text/mel encoder + Discrete VAE spectrogram vocoder)
  • Base Model Checkpoint: samuelolubukun/xtts-v2-igbo-openbible (31,360 steps)
  • Parameters: 521.8M
  • Target Language: Igbo (ig)
  • Tokenizer: Extended Byte Pair Encoding (BPE) with special language tag [ig] (extended vocabulary: 11,533 tokens)
  • Adaptation Dataset: benjaminogbonna/nigerian_common_voice_dataset (Igbo split)
  • Filtered Training Set: 2,809 clean conversational audio samples (filtered for length 1.5s–12s and valid text)
  • Adaptation Steps: 7,488 optimization steps (4 epochs, lr=1.5e-6, grad_accum=4, batch_size=2)
  • Hardware: NVIDIA A10G (24 GB VRAM)

Conversational Adaptation Dashboard & Convergence

The visualization below highlights the multidimensional shift from the Stage 1 OpenBible foundation model to this Stage 2 Common Voice conversational model across speaking rate alignment, pitch tracking, acoustic loss, and timbre fidelity:

Conversational Adaptation Dashboard

Training Progression Summary

Stage / Step Global Step Train Loss Text CE Loss Mel Spectrogram Loss Notes
Adaptation Start 0 0.9812 0.0425 3.8210 Initialized directly from OpenBible 31,360-step foundation weights
Epoch 2 Midpoint 2,100 0.8760 0.0421 3.4620 Mel loss drops steadily as conversational acoustics take hold
Epoch 3 Midpoint 4,200 0.8607 0.0411 3.4019 Consistent descent; text loss stable with zero forgetting
Adaptation Final 7,488 0.8487 0.0384 3.3547 Mastered human conversational pace and downstep contours

Acoustic Improvements Over Stage 1

Objective acoustic and linguistic evaluation across unseen Common Voice test speakers demonstrates significant improvements:

Metric Stage 1 (OpenBible Foundation) Stage 2 (Conversational Adapted - This Model) Improvement / Impact
Human Speaking Rate Match Rushed / abrupt on compounds (4.78s) 6.36s (Matches Human Target: 6.59s) Natural conversational phrasing matching human breath groups
Female Speaker Timbre Match 197.8 Hz (biased to male narrator) 233.7 Hz (Matches Female Target: 251.5 Hz) +35.9 Hz: Clones real female pitch without deep voice bleeding
Pitch Dynamic Range (StdDev) 40.5 Hz 67.6 Hz +66.9%: Dynamic tonal expressiveness on Igbo downstep contours
Vocal Timbre Match Score 68.0% 93.0% +36% improvement in zero-shot voice cloning fidelity

Audio Comparison: Original vs. Stage 1 vs. Stage 2 Cloned Synthesis

Below is a side-by-side comparison between the unseen reference speakers and the speech synthesized by the Stage 1 foundation model versus this Stage 2 conversational adapted model.

All reference audios are sourced from benjaminogbonna/nigerian_common_voice_dataset (Igbo split) by Benjamin Ogbonna:

Sample 1 (Cultural Festival): "kọwàrà mmemme iri ji ọhụrụ dịka oke omenala n'ala Igbo."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 2 (National Recognition): "ma mee ya ka ọ gbagotekwuo n'ogoogo ịbụ emume a mà àmá n'ala Nigeria."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 3 (Expressive & Reflective): "Ọ kpọpụtàsịrị ihe dị iche iche nke ahụ pụrụ ibute."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 4 (Everyday Descriptive): "Nwata nwagbọghọ a mara mma ma nwekwaa ezigbo agwa."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 5 (Civic & Education): "A kọwaala ọba akwụkwọ dịka ụlọọrụ dị mkpà n'ịkwàlite mmepe obodo."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

🔁 Stage 1 OpenBible Benchmark Samples: Foundation vs. Conversational Comparison

These are the exact benchmark samples originally evaluated on the samuelolubukun/xtts-v2-igbo-openbible Stage 1 foundation model repository. Below, the same reference audio clips and prompt texts are synthesized side-by-side using the Stage 2 Conversational fine-tuned model to directly compare the change in cadence, natural pacing, and conversational tone:

Stage 1 Benchmark Sample 1: "mbize kwesiri ka e lebàra ya anya mgwamgwa ugbua"

Column 1: Original Reference Audio Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) Column 3: XTTS-v2 Stage 2 (Conversational Adapted)

Stage 1 Benchmark Sample 2: "Amụrụ ya na Jos, steetị Naijiria."

Column 1: Original Reference Audio Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) Column 3: XTTS-v2 Stage 2 (Conversational Adapted)

Stage 1 Benchmark Sample 3: "nke onye nọchitere anya ya bụ onyeisi ọrụ oyibo na steeti ahụ"

Column 1: Original Reference Audio Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) Column 3: XTTS-v2 Stage 2 (Conversational Adapted)

Repository Files

This repository contains all files necessary to run standalone synthesis or merge with standard Coqui pipelines:

File Description
model.pth Conversational fine-tuned GPT decoder weights (checkpoint_7200.pth)
config.json Updated model configuration with ig language enabled
vocab.json Extended BPE tokenizer vocabulary (11,533 tokens)
dvae.pth Discrete VAE acoustic spectrogram decoder
mel_stats.pth Mel spectrogram normalization statistics
training_curves.png High-resolution 4-panel conversational adaptation dashboard
samples/ 5-way zero-shot voice cloning audio demonstrations

How to Use

Ensure coqui-tts and dependencies are installed:

pip install coqui-tts "transformers<4.45.0" torch torchaudio huggingface_hub

Load the model and synthesize speech directly from this Hugging Face repository:

import os
import re
import torch
import torchaudio
from huggingface_hub import snapshot_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts

# 1. Download model repository
model_dir = snapshot_download(repo_id="samuelolubukun/xtts-v2-igbo-openbible-conversational")

# 2. Load model config and checkpoint
config = XttsConfig()
config.load_json(os.path.join(model_dir, "config.json"))
model = Xtts.init_from_config(config)
model.load_checkpoint(
    config,
    checkpoint_path=os.path.join(model_dir, "model.pth"),
    vocab_path=os.path.join(model_dir, "vocab.json"),
    use_deepspeed=False
)
model.cuda() if torch.cuda.is_available() else model.cpu()

# 3. Text cleaner fallback for custom language support
orig_preprocess = model.tokenizer.preprocess_text
def safe_preprocess_text(txt, lang):
    try:
        return orig_preprocess(txt, lang)
    except NotImplementedError:
        txt = txt.lower()
        return re.sub(r"\s+", " ", txt).strip()
model.tokenizer.preprocess_text = safe_preprocess_text

# 4. Compute speaker conditioning latents (using bundled reference audio or your own 3-10s WAV)
ref_audio = os.path.join(model_dir, "samples/sample1_ref.wav")
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
    audio_path=ref_audio,
    gpt_cond_len=30,
    max_ref_length=30,
    sound_norm_refs=False,
)

# 5. Synthesize speech
text = "Nnọọ, kedu ka ị mere? Anyị na-asụ asụsụ Igbo ugbu a n'ụzọ dị nro."
out = model.inference(
    text=text,
    language="ig",
    gpt_cond_latent=gpt_cond_latent,
    speaker_embedding=speaker_embedding,
    temperature=0.3,
    length_penalty=1.0,
    repetition_penalty=5.0,
    top_k=30,
    top_p=0.85,
)

# 6. Save output audio (24,000 Hz)
torchaudio.save("output_ig.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)
print("Saved synthesized speech to output_ig.wav")

Limitations and Ethical Usage

  • Conversational Acoustics vs. Studio Purity: While Stage 2 eliminates the solemn scripture register and unlocks flexible speaker timbres, Common Voice recordings contain diverse crowd-sourced microphones and acoustic environments. Reference audios with heavy background chatter or clipping should be avoided.
  • Sub-dot and Diacritic Sensitivity: Igbo distinguishes vowel harmony sets using under-dots (ị, ọ, ụ) and nasal consonants (ñ). Text fed into the model should preserve proper standard Igbo orthography for optimal pronunciation.
  • Ethical Voice Cloning: This model is intended for research, linguistic preservation, and accessibility in African languages. Explicit consent should be obtained prior to cloning recognizable individual voices.
  • Licensing: Inherits the Coqui Public Model License / MRCUL.

Author & Attribution

Developed and fine-tuned by Samuel Olubukun:

Citation

If you use this model or checkpoints in your research, applications, or benchmarks, please cite:

@misc{olubukun2026xttsigboconversational,
  author = {Samuel Olubukun},
  title = {XTTS-v2 Igbo Conversational: 4-Epoch Domain Adaptation on Nigerian Common Voice},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/samuelolubukun/xtts-v2-igbo-conversational}}
}

Acknowledgments

Downloads last month
184
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for samuelolubukun/xtts-v2-igbo-openbible-conversational

Base model

coqui/XTTS-v2
Finetuned
(1)
this model

Datasets used to train samuelolubukun/xtts-v2-igbo-openbible-conversational