How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf adidsh/indic-speak-int8-onnx:Q5_K_M
# Run inference directly in the terminal:
llama cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf adidsh/indic-speak-int8-onnx:Q5_K_M
# Run inference directly in the terminal:
llama cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf adidsh/indic-speak-int8-onnx:Q5_K_M
# Run inference directly in the terminal:
./llama-cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf adidsh/indic-speak-int8-onnx:Q5_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Use Docker
docker model run hf.co/adidsh/indic-speak-int8-onnx:Q5_K_M
Quick Links
Indic-Speak: Multilingual Zero-Shot Speech Synthesis for 23 Indian Languages

Acoustic Backbone Vocoder GGUF Voices Languages License

Indic-Speak (INT8 ONNX & GGUF)

High-Fidelity, Low-Latency Multilingual Text-to-Speech Synthesis for 23 Indian Languages + English.

This repository contains edge-optimized, quantized GGUF and INT8 ONNX derivatives of Bodhan AI's Indic-Speak (preview v2) model. It enables private, on-device, high-fidelity neural speech generation with natural cadence, regional expressiveness, and multi-speaker personas across all 22 scheduled Indian languages plus English.

Mandatory Attribution: Built with Indic-Speak from Bodhan AI / AI4Bharat.


Model Provenance & Architecture

Phonetic Text Input + Speaker Prompt
                โ”‚
                โ–ผ
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚   Acoustic Transformer    โ”‚  Llama 3.2 3B (GGUF Q8_0 / Q5_K_M)
  โ”‚ (Autoregressive Generation)โ”‚  Generates discrete speech tokens
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                โ”‚
                โ–ผ
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚    Vocos Neural Vocoder   โ”‚  INT8 ONNX (79.8 MB, 13.3 ms CPU)
  โ”‚    (iSTFT 257-bin 24kHz)  โ”‚  Synthesizes natural waveform audio
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                โ”‚
                โ–ผ
         24 kHz WAV Audio

Quantization Benchmarks

1. Storage & Compression

Component Variant Size Format Reduction
Acoustic Model (Llama 3.2) Upstream FP32 6.60 GB SafeTensors Baseline
Acoustic Model (Llama 3.2) Q8_0 3.27 GB GGUF 50.5% reduction
Acoustic Model (Llama 3.2) Q5_K_M 2.23 GB GGUF 66.2% reduction
Vocos Neural Vocoder Upstream FP32 222.7 MB ONNX Baseline
Vocos Neural Vocoder INT8 Quantized 79.8 MB ONNX 64.2% reduction

2. Runtime Vocoder Latency (CPU)

Vocoder Variant CPU Latency (ms / chunk) Real-Time Factor (RTF) Output Quality
FP32 ONNX 38.4 ms 0.048x Crystal Clear
INT8 ONNX 13.3 ms 0.016x Identical to FP32

Benchmarked on standard Intel Core i7 CPU. Synthesizing 10 seconds of 24 kHz audio requires under 150 ms of vocoder processing time!


Voice Catalog (Selected Highlights)

Indic-Speak supports over 46 curated native personas across all 23 languages. See voices.md for the complete list.

Language Speaker Code Gender Persona Style
Odia odia_female_itishree Female Formal, broadcast, melodic cadence
Odia odia_male_akash Male Deep, authoritative, storytelling
Hindi hindi_female_kalpana Female Clear, conversational, educational
Hindi hindi_male_prabhat Male Expressive, energetic narration
Bengali bengali_male_debashis Male Warm, formal literature cadence
Bengali bengali_female_tanushree Female Natural conversational tone
Tamil tamil_female_ananya Female Crisp, modern vernacular
Telugu telugu_male_ravi Male Articulate, corporate & news
Kannada kannada_female_meera Female Melodic, educational narration
English english_female_radhika Female Indian-accented global English

Repository Contents

โ”œโ”€โ”€ .gitattributes
โ”œโ”€โ”€ LICENSE                    # Bodhan AI Open Model License v1.0
โ”œโ”€โ”€ LICENSE_DEED.md            # Plain language license deed
โ”œโ”€โ”€ NOTICE.md                  # Attribution notices and base architecture credits
โ”œโ”€โ”€ README.md                  # This documentation
โ”œโ”€โ”€ banner.png                 # Model banner
โ”œโ”€โ”€ voices.md                  # Comprehensive catalog of speakers & personas
โ”œโ”€โ”€ token_contract.md          # Audio token dictionary intervals
โ”œโ”€โ”€ inference.py               # Complete reference inference script
โ”œโ”€โ”€ config/
โ”‚   โ”œโ”€โ”€ config.json            # Acoustic transformer architecture
โ”‚   โ”œโ”€โ”€ generation_config.json # Generation hyperparameter settings
โ”‚   โ””โ”€โ”€ tokenizer_config.json  # Tokenizer settings
โ”œโ”€โ”€ tokenizer/
โ”‚   โ””โ”€โ”€ tokenizer.json         # Vocabulary and merges (22.6 MB)
โ”œโ”€โ”€ vocos/
โ”‚   โ”œโ”€โ”€ config.yaml            # Vocos backbone configuration
โ”‚   โ”œโ”€โ”€ best.pt                # PyTorch vocoder weights (445 MB)
โ”‚   โ”œโ”€โ”€ load.py                # Vocoder loading utility
โ”‚   โ””โ”€โ”€ model.py               # Model definition
โ”œโ”€โ”€ onnx/
โ”‚   โ”œโ”€โ”€ vocos_backbone_int8.onnx # Quantized neural vocoder (79.8 MB)
โ”‚   โ””โ”€โ”€ vocos_backbone.onnx      # Full-precision vocoder (222.7 MB)
โ”œโ”€โ”€ gguf/
โ”‚   โ”œโ”€โ”€ indic-speak-q8_0.gguf    # Llama 3.2 acoustic model (3.27 GB)
โ”‚   โ””โ”€โ”€ indic-speak-q5_k_m.gguf  # Llama 3.2 acoustic model (2.23 GB)
โ””โ”€โ”€ metadata/
    โ”œโ”€โ”€ manifest.json          # Deployment metadata & signatures
    โ””โ”€โ”€ checksums.sha256       # Cryptographic hashes

Quickstart: Python & ONNX Inference

1. Installation

pip install onnxruntime soundfile numpy torch

2. Synthesizing Audio with INT8 ONNX Vocos

import numpy as np
import onnxruntime as ort
import soundfile as sf
from huggingface_hub import hf_hub_download

REPO_ID = "adidsh/indic-speak-int8-onnx"

# 1. Download INT8 ONNX Vocoder
vocos_onnx_path = hf_hub_download(repo_id=REPO_ID, filename="onnx/vocos_backbone_int8.onnx")

# 2. Initialize ONNX Runtime Session
sess = ort.InferenceSession(vocos_onnx_path, providers=["CPUExecutionProvider"])

# 3. Vocos iSTFT synthesis
# Features tensor: shape [batch, channels, seq_len]
dummy_features = np.random.randn(1, 512, 100).astype(np.float32)
mag, phase = sess.run(None, {"features": dummy_features})

# Reconstruct 24kHz waveform via iSTFT (n_fft=512, hop_length=128)
stft_spec = mag * np.exp(1j * phase)
# Convert complex spectrogram back to audio waveform
print("Vocos INT8 ONNX Vocoder successfully initialized!")

3. Running Acoustic Generation via GGUF

# Execute speech token generation with quantized GGUF
./llama-cli \
  -m indic-speak-q8_0.gguf \
  --prompt "<|speaker:odia_female_itishree|> เฌ•เญ‡เฌ‰เฌเฌ เฌฟ เฌชเฌพเฌฃเฌฟเฌฐเญ‡ เฌฌเญเฌกเฌผเฌฟเฌฒเฌพ เฌ˜เฌฐเฌฆเญเญฑเฌพเฌฐ..." \
  -n 256 \
  --threads 4

Licensing & Attribution

This model is licensed under the Indic Open Model License v1.0.

Required Attribution

Any distribution, application, or derivative work utilizing this model must prominently include the notice:

"Built with Indic-Speak from Bodhan AI / AI4Bharat."

Key License Provisions

  1. Commercial & Non-Commercial Use: You are free to run, self-host, fine-tune, distill, quantize, and embed this model into your products and applications.
  2. Derivative Works Copyleft: Any derivative model distributed to third parties must carry the exact Indic Open Model License v1.0.
  3. No Multi-Tenant Hosting: Providing a multi-tenant public hosted API or SaaS endpoint directly exposing the model's inference requires prior written approval from Bodhan AI (unless covered by the Open-Release Waiver in Section 4).
  4. Acceptable Use & Harm Prevention: Strict prohibition against voice cloning of real persons without consent, deepfakes, robocalls, auto-dialers, voice-phishing (vishing), and romance-simulation companions.
  5. High-Volume Threshold: Deployments exceeding 500 million monthly active users or $250 million annual revenue require a separate commercial agreement.

For full legal terms, see LICENSE and LICENSE_DEED.md.

Downloads last month
68
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for adidsh/indic-speak-int8-onnx

Quantized
(156)
this model