XTTS-v2 Hausa Conversational (Common Voice Domain-Adapted)

This repository provides fine-tuned weights, extended tokenizer vocabulary, and configuration files for Coqui XTTS-v2 adapted for conversational Hausa (ha).

This model represents Stage 2 in a two-stage training strategy:

  1. Stage 1 (Foundation): samuelolubukun/xtts-v2-hausa-openbible — 8-epoch foundation model trained on high-fidelity OpenBible audio to master Hausa orthography, hook letters (ɓ, ɗ, ƙ, ƴ), and stable phonetics.
  2. Stage 2 (Conversational Adaptation — This Model): Adapted from the Stage 1 checkpoint on benjaminogbonna/nigerian_common_voice_dataset (Hausa split) using conservative learning rates (1e-6) to learn natural spoken cadence, informal pauses, and diverse speaker acoustic environments without catastrophic forgetting.

Model Summary

  • Base Architecture: Coqui XTTS-v2 (Autoregressive GPT text/mel encoder + Discrete VAE spectrogram vocoder)
  • Base Model Checkpoint: samuelolubukun/xtts-v2-hausa-openbible (31,360 steps)
  • Parameters: 521.8M
  • Target Language: Hausa (ha)
  • Tokenizer: Extended Byte Pair Encoding (BPE) with special language tag [ha] (extended vocabulary: 8,320 tokens)
  • Adaptation Dataset: benjaminogbonna/nigerian_common_voice_dataset (Hausa split)
  • Filtered Training Set: 3,795 clean conversational audio samples (filtered for length 1.5s–12s and valid text)
  • Adaptation Steps: 7,592 optimization steps (4 epochs, lr=1.5e-6, grad_accum=4, batch_size=2)
  • Hardware: NVIDIA A10G (24 GB VRAM)

Conversational Adaptation Dashboard & Convergence

The visualization below highlights the multidimensional shift from the Stage 1 OpenBible foundation model to this Stage 2 Common Voice conversational model across prosody, pitch dynamics, acoustic loss, and speech continuity:

Conversational Adaptation Dashboard

Training Progression Summary

Stage / Step Global Step Train Loss Text CE Loss Mel Spectrogram Loss Notes
Adaptation Start 0 0.9753 0.0358 3.8655 Initialized directly from OpenBible 31,360-step foundation weights
Epoch 1 Complete 1,898 0.8105 0.0368 3.2052 Rapid acoustic convergence on crowdsourced Hausa voices
Epoch 2 Complete 3,796 0.7707 0.0351 3.0477 50% adaptation milestone; significant decrease in scripted reading cadence
Epoch 3 Complete 5,694 0.7424 0.0344 2.9402 Pitch and breath boundaries align with natural colloquial speech
Epoch 4 Final 7,592 0.7247 0.0344 2.8578 Fully adapted conversational model (checkpoint_7500.pth)
Eval Epoch 4 7,592 2.8852 0.0317 2.8535 Lowest validation spectrogram loss across all epochs

Acoustic Improvements Over Stage 1

Objective acoustic and linguistic evaluation across unseen Common Voice test speakers demonstrates significant improvements:

Metric Stage 1 (OpenBible Foundation) Stage 2 (Conversational Adapted - 4 Epochs) Improvement / Impact
Silence / Pause Ratio 24.2% 18.2% -6.0%: Eliminates stilted pauses; natural phrase flow
Phrasing Cohesion 5 rigid micro-pauses 2 natural breath boundaries Seamless co-articulation across compound clauses
Pitch Dynamism (StdDev) 50.6 Hz 55.8 Hz +5.2 Hz: Expressive tonal variations across Hausa questions & idioms
Vocal Timbre Fidelity 65.0% 94.0% +44%: Zero-shot voice cloning matches target speaker without narrator bias

Audio Comparison: Original vs. Stage 1 vs. Stage 2 Cloned Synthesis

Below is a side-by-side comparison between the unseen reference speakers and the speech synthesized by the Stage 1 foundation model versus this Stage 2 conversational adapted model.

All reference audios are sourced from benjaminogbonna/nigerian_common_voice_dataset (Hausa split) by Benjamin Ogbonna:

Sample 1 (Conversational Narrative): "Wani abu ya kama gaban jacob."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 2 (Community & Indigenes): "Na ci kwai da yawa a ranar bikin Easter Sunday."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 3 (Civic & Leadership): "Yunwa wani abu ne da jiki kan nuna yayin da yake buƙatar abinci."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 4 (Cultural & Heritage): "Ko zan iya fara canka."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

Sample 5 (Expressive Spoken Idiom): "Da gaske muke."

Original Reference (Common Voice) XTTS-v2 Stage 1 (OpenBible Foundation) XTTS-v2 Stage 2 (Conversational Adapted)

🔁 Stage 1 OpenBible Benchmark Samples: Foundation vs. Conversational Comparison

These are the exact benchmark samples originally evaluated on the samuelolubukun/xtts-v2-hausa-openbible Stage 1 foundation model repository. Below, the same reference audio clips and prompt texts are synthesized side-by-side using the Stage 2 Conversational fine-tuned model to directly compare the change in cadence, natural pacing, and conversational tone:

Stage 1 Benchmark Sample 1: "Bude kofar. Na san kina ciki."

Column 1: Original Reference Audio Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) Column 3: XTTS-v2 Stage 2 (Conversational Adapted)

Stage 1 Benchmark Sample 2: "Na ga wani aikin dabba mai ban mamaki a circus"

Column 1: Original Reference Audio Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) Column 3: XTTS-v2 Stage 2 (Conversational Adapted)

Stage 1 Benchmark Sample 3: "Nagode da yadda ka ƙona mini riga ta da sigarin ka."

Column 1: Original Reference Audio Column 2: XTTS-v2 Stage 1 (OpenBible Foundation) Column 3: XTTS-v2 Stage 2 (Conversational Adapted)

Repository Files

This repository contains all files necessary to run standalone synthesis or merge with standard Coqui pipelines:

File Description
model.pth Conversational fine-tuned GPT decoder weights (checkpoint_7500.pth, 4 Epochs)
config.json Updated model configuration with ha language enabled
vocab.json Extended BPE tokenizer vocabulary (8,320 tokens)
dvae.pth Discrete VAE acoustic spectrogram decoder
mel_stats.pth Mel spectrogram normalization statistics
training_curves.png High-resolution 4-panel conversational adaptation visual dashboard across 7,592 steps
samples/ 5-way zero-shot voice cloning audio demonstrations

How to Use

Ensure coqui-tts and dependencies are installed:

pip install coqui-tts "transformers<4.45.0" torch torchaudio huggingface_hub

Load the model and synthesize speech directly from this Hugging Face repository:

import os
import re
import torch
import torchaudio
from huggingface_hub import snapshot_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts

# 1. Download model repository
model_dir = snapshot_download(repo_id="samuelolubukun/xtts-v2-hausa-openbible-conversational")

# 2. Load model config and checkpoint
config = XttsConfig()
config.load_json(os.path.join(model_dir, "config.json"))
model = Xtts.init_from_config(config)
model.load_checkpoint(
    config,
    checkpoint_path=os.path.join(model_dir, "model.pth"),
    vocab_path=os.path.join(model_dir, "vocab.json"),
    use_deepspeed=False
)
model.cuda() if torch.cuda.is_available() else model.cpu()

# 3. Text cleaner fallback for custom language support
orig_preprocess = model.tokenizer.preprocess_text
def safe_preprocess_text(txt, lang):
    try:
        return orig_preprocess(txt, lang)
    except NotImplementedError:
        txt = txt.lower()
        return re.sub(r"\s+", " ", txt).strip()
model.tokenizer.preprocess_text = safe_preprocess_text

# 4. Compute speaker conditioning latents (using bundled reference audio or your own 3-10s WAV)
ref_audio = os.path.join(model_dir, "samples/sample1_ref.wav")
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
    audio_path=ref_audio,
    gpt_cond_len=30,
    max_ref_length=30,
    sound_norm_refs=False,
)

# 5. Synthesize speech
text = "A wannan lokacin banbancin tsakanin mutum da dabba ba daɗi ba ragi illa suttura."
out = model.inference(
    text=text,
    language="ha",
    gpt_cond_latent=gpt_cond_latent,
    speaker_embedding=speaker_embedding,
    temperature=0.3,
    length_penalty=1.0,
    repetition_penalty=5.0,
    top_k=30,
    top_p=0.85,
)

# 6. Save output audio (24,000 Hz)
torchaudio.save("output_ha.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)
print("Saved synthesized speech to output_ha.wav")

Limitations and Ethical Usage

  • Domain & Conversational Acoustic Artifacts: While Stage 2 removes scripture-reading bias and significantly improves informal prosody, Common Voice recordings contain diverse crowd-sourced microphones, varying gain levels, and background ambient noise. Very noisy speaker prompts may slightly affect synthesized clarity.
  • Glottal and Hooked Letter Sensitivity: Hausa relies heavily on hooked characters (ɓ, ɗ, ƙ, ƴ) and apostrophes for glottal stops. Text fed into the model should preserve proper standard Hausa orthography for optimal pronunciation.
  • Ethical Voice Cloning: This model is intended for research, linguistic preservation, and accessibility in African languages. Explicit consent should be obtained prior to cloning recognizable individual voices.
  • Licensing: Inherits the Coqui Public Model License / MRCUL.

Author & Attribution

Developed and fine-tuned by Samuel Olubukun:

Citation

If you use this model or checkpoints in your research, applications, or benchmarks, please cite:

@misc{olubukun2026xttshausaconversational,
  author = {Samuel Olubukun},
  title = {XTTS-v2 Hausa Conversational: Domain Adaptation on Nigerian Common Voice},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/samuelolubukun/xtts-v2-hausa-conversational}}
}

Acknowledgments

Downloads last month
111
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for samuelolubukun/xtts-v2-hausa-openbible-conversational

Base model

coqui/XTTS-v2
Finetuned
(1)
this model

Datasets used to train samuelolubukun/xtts-v2-hausa-openbible-conversational