Masri Higgs v3 — Egyptian Arabic TTS (production, merged)

Drop-in 4 B TTS model for Egyptian Arabic. This is ehabnegm/higgs-tts-3-4b-egyptian-v3 with a LoRA adapter merged in, fine-tuned on 98 hours of denoised, consensus-verified Egyptian speech. No adapter loading required — it works with the existing serving stack unchanged.

Why this checkpoint

Trained for 3 epochs; step 600 was selected, not the final step. Validation loss kept falling through step 1134, but ASR round-trip word error rate started rising after step 600 — the signature of mild overfitting to the training distribution. We shipped the checkpoint that generalizes best, not the one with the prettiest loss curve.

Checkpoint Val loss CER mean CER median WER mean
base (higgs-tts-3-4b-egyptian-v3) 4.2168 0.0300 0.0177 0.1146
step 600 — this model 4.0995 0.0301 0.0145 0.1012
step 1134 (final) 4.0875 0.0331 0.0198 0.1219

Measured on 40 held-out clips, video-disjoint from training. CER/WER come from transcribing the generated audio back with oddadmix/whisper-large-v3-turbo-arabic-dialectal-v2 and comparing to the input text. WER improved 12% relative to base; CER mean is unchanged.

Quickstart

pip install torch transformers safetensors soundfile
import sys, torch, soundfile as sf, numpy as np
sys.path.insert(0, "serving")
from transformers import AutoTokenizer, AutoModel
from higgs_v3_model import HiggsV3ForTTS

tok   = AutoTokenizer.from_pretrained(".")
model, _ = HiggsV3ForTTS.from_checkpoint(".", dtype=torch.bfloat16); model.cuda().eval()
codec = AutoModel.from_pretrained("bosonai/higgs-audio-v2-tokenizer",
                                  trust_remote_code=True).eval().cuda()

Generation (reference-conditioned, 25 Hz frames, 8 codebooks with delay pattern) is implemented in serving/serve_openai.py — gen_raw() and synth_pcm().

OpenAI-compatible server

pip install -r serving/requirements.txt
MODEL_DIR=. REFS_DIR=./refs uvicorn serving.serve_openai:app --host 0.0.0.0 --port 8000
curl -X POST localhost:8000/v1/audio/speech \
  -H 'content-type: application/json' \
  -d '{"input":"إزيك يا صاحبي؟ عامل إيه النهاردة؟","voice":"masri"}' --output out.wav

You must supply a voice reference in refs/*.pt — a dict with codes [T, 8] and text_ids. Any 6–10 s clip of the target speaker works.

Control tags

Inherited from Higgs v3 — see PROMPTING.md:

<|emotion:elation|>أهلا بيكم، مبسوطين جدا إنكم معانا النهاردة!
كلام أول <|prosody:pause|> وبعدين كلام تاني
<|sfx:laughter|>هاها، تمام كده

Training

Base ehabnegm/higgs-tts-3-4b-egyptian-v3 (4.043 B)
Data ehabnegm/masri-100h-egyptian-tts-enhanced — 98.2 h, 24,445 clips
Method LoRA r=64, α=128, dropout 0.05 on q,k,v,o,gate,up,down
Trainable 132.1 M / 4.176 B (3.16%)
Schedule 3 epochs, 1,134 steps, batch 32 × grad-accum 2, OneCycle LR 1e-4
Precision bf16 + gradient checkpointing
Hardware 1× A100-SXM4-80 GB, ~4 h
Audio codec bosonai/higgs-audio-v2-tokenizer, 25 Hz, 8 codebooks

Full training and evaluation scripts: ehabnegm/masri-higgs-v3-egyptian-lora

Out-of-domain behaviour

Tested on 15 sentences that appear nowhere in the training data:

Category Mean CER Note
Long paragraphs (~22 s) 0.013 no drift or collapse — the key production risk, and it's clean
Difficult letters (ض ظ ذ ث ق غ) 0.058 solid
Prosody (questions, exclamations) 0.026 solid
Numbers 0.571 ⚠ scoring artifact — model says «تمانية وأربعين», reference says «48»
Code-switch (AR↔EN) 0.558 ⚠ scoring artifact — Arabic ASR transliterates English

The numbers and code-switch scores are measurement artifacts, not failures. Judge those by ear.

Limitations

  • Single narrator, single domain. Popular-science / history narration. Zero-shot cloning of other voices still works (inherited from Higgs v3) but has not been re-benchmarked after fine-tuning.
  • Transcripts were machine-generated. The training text came from a 3-way ASR consensus, not human transcription. Residual systematic errors can be learned.
  • CER did not improve, only WER. Gains are at the word level, not the character level.
  • Not evaluated for speaker similarity. We measured intelligibility, not voice likeness — the metric that matters most for voice products was not quantified.
  • Inherits all Higgs Audio v3 licence restrictions.

Ethics

The voice derives from publicly available YouTube content by a real, identifiable person. Do not publish, commercialize, or impersonate without the speaker's consent. Disclose synthetic audio.

Citation

@misc{negm_masri_higgs_v3_2026,
  title  = {Masri Higgs v3: Egyptian Arabic TTS fine-tuned on consensus-verified speech},
  author = {Negm, Ehab},
  year   = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/ehabnegm/masri-higgs-v3-egyptian-tts}}
}
Downloads last month
46
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ehabnegm/masri-higgs-v3-egyptian-tts

Finetuned
(1)
this model