File size: 1,332 Bytes
6f0964d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
---
license: other
language: [te, en]
pipeline_tag: text-to-speech
tags: [text-to-speech, telugu, higgs-audio, voice-avatar, code-switch]
base_model: bosonai/higgs-audio-v3-tts-4b
---

# Higgs Audio v3 — Telugu single-speaker voice (ISO-romanized fine-tune)

Expressive single-speaker **Telugu** (code-switch) TTS, fine-tuned from Higgs
Audio v3 (4B). Generates this speaker's voice **from text with no reference clip**.

**Frontend matters:** Telugu text must be ISO-15919 romanized with the *same*
`frontend.py` used in training (Telugu tokenizes to byte-fragments otherwise).
`generate_speech` handles the rest.

```python
import torch, torchaudio
from transformers import AutoModelForCausalLM, AutoTokenizer
from frontend import romanize          # shipped in this repo

repo = "BNarayanaReddy/higgs-telugu-e5-lora-epoch_6"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True,
                                             dtype=torch.bfloat16).to("cuda").eval()

text = "హలో, ఈ రోజు ఎలా ఉన్నారు?"
wav = model.generate_speech(romanize(text, "iso"), tok, temperature=0.7, top_p=0.95)
torchaudio.save("out.wav", wav.unsqueeze(0), model.config.sample_rate)
```

Run #1 of an ISO-romanized single-speaker adaptation. Research use.