Instructions to use ehabnegm/masri-higgs-v3-egyptian-tts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ehabnegm/masri-higgs-v3-egyptian-tts with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="ehabnegm/masri-higgs-v3-egyptian-tts")# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("ehabnegm/masri-higgs-v3-egyptian-tts", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Masri Higgs v3 — Egyptian Arabic TTS (production, merged)
Drop-in 4 B TTS model for Egyptian Arabic. This is ehabnegm/higgs-tts-3-4b-egyptian-v3 with a
LoRA adapter merged in, fine-tuned on 98 hours of denoised, consensus-verified Egyptian speech.
No adapter loading required — it works with the existing serving stack unchanged.
Why this checkpoint
Trained for 3 epochs; step 600 was selected, not the final step. Validation loss kept falling through step 1134, but ASR round-trip word error rate started rising after step 600 — the signature of mild overfitting to the training distribution. We shipped the checkpoint that generalizes best, not the one with the prettiest loss curve.
| Checkpoint | Val loss | CER mean | CER median | WER mean |
|---|---|---|---|---|
base (higgs-tts-3-4b-egyptian-v3) |
4.2168 | 0.0300 | 0.0177 | 0.1146 |
| step 600 — this model | 4.0995 | 0.0301 | 0.0145 | 0.1012 |
| step 1134 (final) | 4.0875 | 0.0331 | 0.0198 | 0.1219 |
Measured on 40 held-out clips, video-disjoint from training. CER/WER come from transcribing the
generated audio back with oddadmix/whisper-large-v3-turbo-arabic-dialectal-v2 and comparing to the
input text. WER improved 12% relative to base; CER mean is unchanged.
Quickstart
pip install torch transformers safetensors soundfile
import sys, torch, soundfile as sf, numpy as np
sys.path.insert(0, "serving")
from transformers import AutoTokenizer, AutoModel
from higgs_v3_model import HiggsV3ForTTS
tok = AutoTokenizer.from_pretrained(".")
model, _ = HiggsV3ForTTS.from_checkpoint(".", dtype=torch.bfloat16); model.cuda().eval()
codec = AutoModel.from_pretrained("bosonai/higgs-audio-v2-tokenizer",
trust_remote_code=True).eval().cuda()
Generation (reference-conditioned, 25 Hz frames, 8 codebooks with delay pattern) is implemented in
serving/serve_openai.py — gen_raw() and synth_pcm().
OpenAI-compatible server
pip install -r serving/requirements.txt
MODEL_DIR=. REFS_DIR=./refs uvicorn serving.serve_openai:app --host 0.0.0.0 --port 8000
curl -X POST localhost:8000/v1/audio/speech \
-H 'content-type: application/json' \
-d '{"input":"إزيك يا صاحبي؟ عامل إيه النهاردة؟","voice":"masri"}' --output out.wav
You must supply a voice reference in refs/*.pt — a dict with codes [T, 8] and text_ids.
Any 6–10 s clip of the target speaker works.
Control tags
Inherited from Higgs v3 — see PROMPTING.md:
<|emotion:elation|>أهلا بيكم، مبسوطين جدا إنكم معانا النهاردة!
كلام أول <|prosody:pause|> وبعدين كلام تاني
<|sfx:laughter|>هاها، تمام كده
Training
| Base | ehabnegm/higgs-tts-3-4b-egyptian-v3 (4.043 B) |
| Data | ehabnegm/masri-100h-egyptian-tts-enhanced — 98.2 h, 24,445 clips |
| Method | LoRA r=64, α=128, dropout 0.05 on q,k,v,o,gate,up,down |
| Trainable | 132.1 M / 4.176 B (3.16%) |
| Schedule | 3 epochs, 1,134 steps, batch 32 × grad-accum 2, OneCycle LR 1e-4 |
| Precision | bf16 + gradient checkpointing |
| Hardware | 1× A100-SXM4-80 GB, ~4 h |
| Audio codec | bosonai/higgs-audio-v2-tokenizer, 25 Hz, 8 codebooks |
Full training and evaluation scripts: ehabnegm/masri-higgs-v3-egyptian-lora
Out-of-domain behaviour
Tested on 15 sentences that appear nowhere in the training data:
| Category | Mean CER | Note |
|---|---|---|
| Long paragraphs (~22 s) | 0.013 | no drift or collapse — the key production risk, and it's clean |
| Difficult letters (ض ظ ذ ث ق غ) | 0.058 | solid |
| Prosody (questions, exclamations) | 0.026 | solid |
| Numbers | 0.571 | ⚠ scoring artifact — model says «تمانية وأربعين», reference says «48» |
| Code-switch (AR↔EN) | 0.558 | ⚠ scoring artifact — Arabic ASR transliterates English |
The numbers and code-switch scores are measurement artifacts, not failures. Judge those by ear.
Limitations
- Single narrator, single domain. Popular-science / history narration. Zero-shot cloning of other voices still works (inherited from Higgs v3) but has not been re-benchmarked after fine-tuning.
- Transcripts were machine-generated. The training text came from a 3-way ASR consensus, not human transcription. Residual systematic errors can be learned.
- CER did not improve, only WER. Gains are at the word level, not the character level.
- Not evaluated for speaker similarity. We measured intelligibility, not voice likeness — the metric that matters most for voice products was not quantified.
- Inherits all Higgs Audio v3 licence restrictions.
Ethics
The voice derives from publicly available YouTube content by a real, identifiable person. Do not publish, commercialize, or impersonate without the speaker's consent. Disclose synthetic audio.
Citation
@misc{negm_masri_higgs_v3_2026,
title = {Masri Higgs v3: Egyptian Arabic TTS fine-tuned on consensus-verified speech},
author = {Negm, Ehab},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/ehabnegm/masri-higgs-v3-egyptian-tts}}
}
- Downloads last month
- 46
Model tree for ehabnegm/masri-higgs-v3-egyptian-tts
Base model
ehabnegm/higgs-tts-3-4b-egyptian-v3