BIA-SPARKTTS-V2 / README.md
sawadogosalif's picture
Update README.md
e522617 verified
|
Raw History Blame Contribute Delete
7.9 kB
---
license: apache-2.0
language:
- mos
- fr
tags:
- text-to-speech
- tts
- moore
- burkina-faso
- african-languages
- voice-cloning
- speech-synthesis
pipeline_tag: text-to-speech
---
# BIA-SPARKTTS-V2
## Model Description
BIA-SPARKTTS-V2 is a fine-tuned version of SparkTTS specifically optimized for Moore language text-to-speech synthesis. This model enables high-quality speech generation in Moore with voice cloning capabilities, supporting the linguistic needs of Burkina Faso and the wider Moore-speaking community.
Our model is built upon the powerful SparkTTS architecture, leveraging state-of-the-art neural vocoding and voice cloning techniques. It has been fine-tuned on curated Moore speech data to ensure natural and contextually appropriate speech synthesis for Moore speakers.
## Language Information
- Primary Language: Moore (Mooré)
- Language Code: `moor_Latn`
- Speech Synthesis: Text-to-Speech (TTS)
- Region: Burkina Faso, West Africa
- Speakers: ~5 million native speakers
- Language Family: Niger-Congo → Gur → Oti-Volta → Moore
## Model Performance
The model achieves the following results on the evaluation set:
- Step 1000: Training Loss: 4.158300 | Validation Loss: 4.068496
- Step 2000: Training Loss: 3.816300 | Validation Loss: 3.789840
- Step 3000: Training Loss: 3.655900 | Validation Loss: 3.688933
- Step 4000: Training Loss: 3.609400 | Validation Loss: 3.628450
- Step 5000: Training Loss: 3.528000 | Validation Loss: 3.593397
- Step 6000: Training Loss: 3.463600 | Validation Loss: 3.563343
- Step 7000: Training Loss: 3.404400 | Validation Loss: 3.546153
- Step 8000: Training Loss: 3.395500 | Validation Loss: 3.528718
- Step 9000: Training Loss: 3.367900 | Validation Loss: 3.524420
- Step 10000: Training Loss: 3.359900 | Validation Loss: 3.522905
- Parameters: 0.5B (500M)
- Total Training Steps: 10,680 (8 epochs)
- Sample Rate: 16kHz
- Audio Quality: Natural, clear pronunciation
## Features
- **Voice Creation**: Generate speech with customizable gender, pitch, and speed
- **Voice Cloning**: Clone voices from audio samples (male/female)
- **Custom Voice Cloning**: Upload your own reference audio for personalized voice cloning
- **Flexible Generation**: Adjustable temperature, top-k, and top-p sampling parameters
## Usage
### Installation
```bash
pip install torch transformers gradio soundfile
pip install git+https://github.com/SparkAudio/SparkTTS.git
```
### Basic Usage
```python
import torch
from SparkTTS import SparkTTS
# Initialize model
device = "cuda:0" if torch.cuda.is_available() else "cpu"
model = SparkTTS("burkimbia/BIA-SPARKTTS-V1", device=device)
# Generate speech with voice creation
text = "Yʋʋmd fãa, a ɛlkana rag n yita a tẽngẽ wã"
wav = model.inference(
text,
gender="male",
pitch="moderate",
speed="moderate",
do_sample=True,
temperature=0.8,
top_k=50,
top_p=0.95
)
# Save audio
import soundfile as sf
sf.write("output.wav", wav, 16000)
```
### Voice Cloning with Audio Sample
```python
# Clone voice from reference audio
reference_audio = "path/to/reference/audio.wav"
wav = model.inference(
text,
prompt_speech_path=reference_audio,
prompt_text=None, # Optional: transcript of reference audio
do_sample=True,
temperature=0.8,
top_k=50,
top_p=0.95
)
```
### Advanced Parameters
```python
# Fine-tune generation with custom parameters
wav = model.inference(
text,
gender="female",
pitch="high", # very_low | low | moderate | high | very_high
speed="moderate", # very_low | low | moderate | high | very_high
do_sample=True, # Enable sampling for more natural variation
temperature=0.8, # Higher = more creative (0.1-1.5)
top_k=50, # Number of tokens to consider (1-100)
top_p=0.95, # Nucleus sampling threshold (0.1-1.0)
)
```
## Model Details
### Training Configuration
- Base Model: SparkTTS-0.5B
- Model Size: 500M parameters
- Training Steps: 10,680 steps (8 epochs)
- Batch Size: 32 per device
- Gradient Accumulation Steps: 2
- Effective Batch Size: 64
- Learning Rate: 5e-5
- LR Scheduler: Cosine with 500 warmup steps
- Optimizer: AdamW 8-bit with 0.01 weight decay
- Sample Rate: 16kHz
- Max Tokens: 3000
- Hardware: NVIDIA A100-SXM4-40GB (40GB VRAM)
- Framework: PyTorch (FP32)
- Seed: 2025
### Training Progress
The model showed consistent improvement during training:
- Rapid initial improvement (4.07 → 3.79 validation loss in first 2000 steps)
- Steady convergence from 3.69 to 3.52 over 10,000 steps
- Best model checkpoint saved at step 10,000 (validation loss: 3.522905)
- Evaluation frequency: Every 1000 steps
- Total training time: ~12-14 hours on A100
### Architecture Components
- **Audio Tokenizer**: BiCodec for high-quality audio encoding
- **Language Model**: Causal LM for sequential token generation
- **Vocoder**: Neural vocoder for waveform synthesis
- **Voice Encoder**: For voice cloning and style transfer
## Generation Parameters
### do_sample
- `True`: Enables sampling for natural variation (recommended)
- `False`: Greedy decoding for deterministic output
### temperature
- Range: 0.1 - 1.5
- Default: 0.8
- Lower values (0.3-0.5): More deterministic, consistent pronunciation
- Higher values (0.8-1.2): More natural variation and expressiveness
### top_k
- Range: 1 - 100
- Default: 50
- Number of highest probability tokens to consider during sampling
### top_p (nucleus sampling)
- Range: 0.1 - 1.0
- Default: 0.95
- Cumulative probability threshold for token selection
## Limitations
- The model performs best on standard Moore text with proper orthography
- Extremely long texts may require chunking
- Voice quality depends on reference audio quality for cloning
- Generated speech may vary slightly between runs when sampling is enabled
- Specialized technical terms may require phonetic adjustments
## Use Cases
- **Education**: Moore language learning and pronunciation
- **Accessibility**: Text-to-speech for visually impaired Moore speakers
- **Content Creation**: Audiobooks, podcasts, and media in Moore
- **Communication**: Voice assistants and chatbots for Moore speakers
- **Preservation**: Digital archiving of Moore language resources
## Gradio Demo
A web interface is available for easy testing:
```python
import gradio as gr
from SparkTTS import SparkTTS
model = SparkTTS("burkimbia/BIA-SPARKTTS-V1")
def generate_speech(text, gender, pitch, speed):
wav = model.inference(text, gender=gender, pitch=pitch, speed=speed)
return (16000, wav)
demo = gr.Interface(
fn=generate_speech,
inputs=[
gr.Textbox(label="Moore Text"),
gr.Radio(["male", "female"], label="Gender"),
gr.Slider(1, 5, step=1, value=3, label="Pitch"),
gr.Slider(1, 5, step=1, value=3, label="Speed"),
],
outputs=gr.Audio(label="Generated Speech")
)
demo.launch()
```
## Feedback
We're continuously striving to improve our model's performance and usability. If you have any feedback, suggestions, or encounter any issues, please don't hesitate to reach out to us at [BurkimbIA](https://huggingface.co/burkimbia).
## Citation
If you use this model in your research, please cite:
```bibtex
@misc{bia-sparktts-v1,
title={BIA-SPARKTTS-V2: A Fine-tuned SparkTTS Model for Moore Language Speech Synthesis},
author={BurkimbIA},
year={2025},
howpublished={\url{https://huggingface.co/burkimbia/BIA-SPARKTTS-V1}}
}
```
## Acknowledgments
- Coqui tts team for the SparkTTS architecture and pretrained models
- Speech data contributors who provided high-quality recordings
- BurkimbIA team for data curation and model development
## Framework Versions
- Transformers: 4.57.1
- PyTorch: 2.2.1+cu121
- Python: 3.12+
- Gradio: 5.0+
Developed by [BurkimbIA](https://huggingface.co/burkimbia) - Advancing African Language Technologies