--- license: other language: [te, en] pipeline_tag: text-to-speech tags: [text-to-speech, telugu, higgs-audio, voice-avatar, code-switch] base_model: bosonai/higgs-audio-v3-tts-4b --- # Higgs Audio v3 — Telugu single-speaker voice (ISO-romanized fine-tune) Expressive single-speaker **Telugu** (code-switch) TTS, fine-tuned from Higgs Audio v3 (4B). Generates this speaker's voice **from text with no reference clip**. **Frontend matters:** Telugu text must be ISO-15919 romanized with the *same* `frontend.py` used in training (Telugu tokenizes to byte-fragments otherwise). `generate_speech` handles the rest. ```python import torch, torchaudio from transformers import AutoModelForCausalLM, AutoTokenizer from frontend import romanize # shipped in this repo repo = "BNarayanaReddy/higgs-telugu-e5-lora-epoch_6" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).to("cuda").eval() text = "హలో, ఈ రోజు ఎలా ఉన్నారు?" wav = model.generate_speech(romanize(text, "iso"), tok, temperature=0.7, top_p=0.95) torchaudio.save("out.wav", wav.unsqueeze(0), model.config.sample_rate) ``` Run #1 of an ISO-romanized single-speaker adaptation. Research use.