Higgs-TTS-3-4B โ€” Egyptian Arabic (v3) + OpenAI-compatible streaming server

Egyptian-Arabic TTS fine-tune of bosonai/higgs-tts-3-4b (v3 = noselleel + moustafa + quality-filtered mosaifside, pronunciation lexicon at inference). Ships a self-contained OpenAI-compatible server with low-latency streaming.

  • Voices: noselleel, moustafa, mosaifside ยท 24 kHz
  • Model id: higgs-egyptian-v3
  • Endpoints: POST /v1/audio/speech (OpenAI-compatible), GET /health|/v1/models|/v1/voices|/metrics

Performance (RTX 4090)

  • torch.compile decode: ~16 ms/frame โ†’ RTF ~0.47 (2.4x faster than real-time)
  • Streaming time-to-first-chunk ~0.7 s (windowed decode: left ctx + right lookahead, bit-exact vs full decode)
  • Single continuous generation per turn (no sentence-boundary stalls); startup warm-up

Run

pip install -r serving/requirements.txt
MODEL_DIR=. CODEC=bosonai/higgs-audio-v2-tokenizer \
  uvicorn serving.serve_openai:app --host 0.0.0.0 --port 8000
# first start compiles (~90s) then warms; needs a CUDA GPU

Use (OpenAI SDK)

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
c.audio.speech.create(model="higgs-egyptian-v3", voice="mosaifside",
                      input="ุงุฒูŠูƒ ุนุงู…ู„ ุงูŠู‡ ุงู„ู†ู‡ุงุฑุฏู‡ุŸ").stream_to_file("out.wav")

response_format: pcm (streaming s16le 24 kHz) or wav. See serving/SERVE.md.

Downloads last month
24
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ehabnegm/higgs-tts-3-4b-egyptian-v3-serve

Finetuned
(16)
this model