--- extra_gated_prompt: "Please provide your details and agree to the [LICENSE](https://github.com/Bodhan-AI/bodhan-model-info/blob/main/licenses/indic-open-model-license/v1/Indic_Open_Model_License.md) [[simpler version](https://github.com/Bodhan-AI/bodhan-model-info/blob/main/licenses/indic-open-model-license/v1/Indic_Open_Model_License_Deed.md)] to request access." extra_gated_fields: Company / Organization: text Country: country Intended Use Case: type: select options: - Research - Commercial - Education - label: Other value: other I agree to the license terms: checkbox base_model: meta-llama/Llama-3.2-3B language: - en - hi - bn - mr - te - ta - gu - kn - ml - or - pa - as - ur - brx - doi - kok - ks - mai - ne - mni - sa - sat - sd pipeline_tag: text-to-speech tags: - text-to-speech - tts - speech - audio - indic - multilingual - onnx - int8 - gguf - quantization - edge-ai - vocos - llama-3.2 - indic-speak license: other ---
Indic-Speak: Multilingual Zero-Shot Speech Synthesis for 23 Indian Languages
[![Acoustic Backbone](https://img.shields.io/badge/Acoustic%20Backbone-Llama%203.2%203B-blue?style=flat)](#model-architecture) [![Vocoder](https://img.shields.io/badge/Neural%20Vocoder-Vocos%20INT8%20ONNX-22C55E?style=flat)](#vocoder-benchmarks) [![GGUF](https://img.shields.io/badge/GGUF-Q8__0%20%7C%20Q5__K__M-3B82F6?style=flat)](#quantization-benchmarks) [![Voices](https://img.shields.io/badge/Speakers-46%2B%20Voices-F97316?style=flat)](#voice-catalog) [![Languages](https://img.shields.io/badge/Languages-23%20Languages-F97316?style=flat)](#supported-languages) [![License](https://img.shields.io/badge/License-Indic%20Open%20Model%20License%20v1.0-C1440E?style=flat)](#licensing--attribution)
# Indic-Speak (INT8 ONNX & GGUF) **High-Fidelity, Low-Latency Multilingual Text-to-Speech Synthesis for 23 Indian Languages + English.** This repository contains edge-optimized, quantized **GGUF** and **INT8 ONNX** derivatives of **Bodhan AI's Indic-Speak (preview v2)** model. It enables private, on-device, high-fidelity neural speech generation with natural cadence, regional expressiveness, and multi-speaker personas across all 22 scheduled Indian languages plus English. > **Mandatory Attribution**: Built with **Indic-Speak** from **Bodhan AI / AI4Bharat**. --- ## Model Provenance & Architecture - **Derivative Author**: Community Contributor (2026) - **Direct Upstream Model**: [bodhan-ai/indic-speak-preview-v2](https://huggingface.co/bodhan-ai/indic-speak-preview-v2) - **Upstream Authors**: IITM BODHAN-AI FOUNDATION & AI4Bharat, supported by the Ministry of Education, Govt. of India. - **Upstream License**: [Indic Open Model License v1.0](LICENSE) (Plain language deed: [LICENSE_DEED.md](LICENSE_DEED.md)) - **Base Neural Architectures**: - **Acoustic Backbone**: [meta-llama/Llama-3.2-3B](https://huggingface.co/meta-llama/Llama-3.2-3B) (~3.21B parameters, Llama 3.2 Community License) - **Neural Vocoder**: [Vocos](https://github.com/gemelo-ai/vocos) backbone (24 kHz continuous waveform reconstruction, Apache 2.0) ``` Phonetic Text Input + Speaker Prompt │ ▼ ┌───────────────────────────┐ │ Acoustic Transformer │ Llama 3.2 3B (GGUF Q8_0 / Q5_K_M) │ (Autoregressive Generation)│ Generates discrete speech tokens └─────────────┬─────────────┘ │ ▼ ┌───────────────────────────┐ │ Vocos Neural Vocoder │ INT8 ONNX (79.8 MB, 13.3 ms CPU) │ (iSTFT 257-bin 24kHz) │ Synthesizes natural waveform audio └─────────────┬─────────────┘ │ ▼ 24 kHz WAV Audio ``` --- ## Quantization Benchmarks ### 1. Storage & Compression | Component | Variant | Size | Format | Reduction | | :--- | :--- | :---: | :---: | :---: | | **Acoustic Model (Llama 3.2)** | Upstream FP32 | 6.60 GB | SafeTensors | Baseline | | **Acoustic Model (Llama 3.2)** | **Q8_0** | **3.27 GB** | GGUF | **50.5% reduction** | | **Acoustic Model (Llama 3.2)** | **Q5_K_M** | **2.23 GB** | GGUF | **66.2% reduction** | | **Vocos Neural Vocoder** | Upstream FP32 | 222.7 MB | ONNX | Baseline | | **Vocos Neural Vocoder** | **INT8 Quantized** | **79.8 MB** | ONNX | **64.2% reduction** | ### 2. Runtime Vocoder Latency (CPU) | Vocoder Variant | CPU Latency (ms / chunk) | Real-Time Factor (RTF) | Output Quality | | :--- | :---: | :---: | :---: | | **FP32 ONNX** | 38.4 ms | 0.048x | Crystal Clear | | **INT8 ONNX** | **13.3 ms** | **0.016x** | **Identical to FP32** | *Benchmarked on standard Intel Core i7 CPU. Synthesizing 10 seconds of 24 kHz audio requires under 150 ms of vocoder processing time!* --- ## Voice Catalog (Selected Highlights) Indic-Speak supports over 46 curated native personas across all 23 languages. See [voices.md](voices.md) for the complete list. | Language | Speaker Code | Gender | Persona Style | | :--- | :--- | :---: | :--- | | **Odia** | `odia_female_itishree` | Female | Formal, broadcast, melodic cadence | | **Odia** | `odia_male_akash` | Male | Deep, authoritative, storytelling | | **Hindi** | `hindi_female_kalpana` | Female | Clear, conversational, educational | | **Hindi** | `hindi_male_prabhat` | Male | Expressive, energetic narration | | **Bengali** | `bengali_male_debashis` | Male | Warm, formal literature cadence | | **Bengali** | `bengali_female_tanushree`| Female | Natural conversational tone | | **Tamil** | `tamil_female_ananya` | Female | Crisp, modern vernacular | | **Telugu** | `telugu_male_ravi` | Male | Articulate, corporate & news | | **Kannada** | `kannada_female_meera` | Female | Melodic, educational narration | | **English** | `english_female_radhika`| Female | Indian-accented global English | --- ## Repository Contents ``` ├── .gitattributes ├── LICENSE # Bodhan AI Open Model License v1.0 ├── LICENSE_DEED.md # Plain language license deed ├── NOTICE.md # Attribution notices and base architecture credits ├── README.md # This documentation ├── banner.png # Model banner ├── voices.md # Comprehensive catalog of speakers & personas ├── token_contract.md # Audio token dictionary intervals ├── inference.py # Complete reference inference script ├── config/ │ ├── config.json # Acoustic transformer architecture │ ├── generation_config.json # Generation hyperparameter settings │ └── tokenizer_config.json # Tokenizer settings ├── tokenizer/ │ └── tokenizer.json # Vocabulary and merges (22.6 MB) ├── vocos/ │ ├── config.yaml # Vocos backbone configuration │ ├── best.pt # PyTorch vocoder weights (445 MB) │ ├── load.py # Vocoder loading utility │ └── model.py # Model definition ├── onnx/ │ ├── vocos_backbone_int8.onnx # Quantized neural vocoder (79.8 MB) │ └── vocos_backbone.onnx # Full-precision vocoder (222.7 MB) ├── gguf/ │ ├── indic-speak-q8_0.gguf # Llama 3.2 acoustic model (3.27 GB) │ └── indic-speak-q5_k_m.gguf # Llama 3.2 acoustic model (2.23 GB) └── metadata/ ├── manifest.json # Deployment metadata & signatures └── checksums.sha256 # Cryptographic hashes ``` --- ## Quickstart: Python & ONNX Inference ### 1. Installation ```bash pip install onnxruntime soundfile numpy torch ``` ### 2. Synthesizing Audio with INT8 ONNX Vocos ```python import numpy as np import onnxruntime as ort import soundfile as sf from huggingface_hub import hf_hub_download REPO_ID = "adidsh/indic-speak-int8-onnx" # 1. Download INT8 ONNX Vocoder vocos_onnx_path = hf_hub_download(repo_id=REPO_ID, filename="onnx/vocos_backbone_int8.onnx") # 2. Initialize ONNX Runtime Session sess = ort.InferenceSession(vocos_onnx_path, providers=["CPUExecutionProvider"]) # 3. Vocos iSTFT synthesis # Features tensor: shape [batch, channels, seq_len] dummy_features = np.random.randn(1, 512, 100).astype(np.float32) mag, phase = sess.run(None, {"features": dummy_features}) # Reconstruct 24kHz waveform via iSTFT (n_fft=512, hop_length=128) stft_spec = mag * np.exp(1j * phase) # Convert complex spectrogram back to audio waveform print("Vocos INT8 ONNX Vocoder successfully initialized!") ``` ### 3. Running Acoustic Generation via GGUF ```bash # Execute speech token generation with quantized GGUF ./llama-cli \ -m indic-speak-q8_0.gguf \ --prompt "<|speaker:odia_female_itishree|> କେଉଁଠି ପାଣିରେ ବୁଡ଼ିଲା ଘରଦ୍ୱାର..." \ -n 256 \ --threads 4 ``` --- ## Licensing & Attribution This model is licensed under the **Indic Open Model License v1.0**. ### Required Attribution Any distribution, application, or derivative work utilizing this model must prominently include the notice: > **"Built with Indic-Speak from Bodhan AI / AI4Bharat."** ### Key License Provisions 1. **Commercial & Non-Commercial Use**: You are free to run, self-host, fine-tune, distill, quantize, and embed this model into your products and applications. 2. **Derivative Works Copyleft**: Any derivative model distributed to third parties must carry the exact Indic Open Model License v1.0. 3. **No Multi-Tenant Hosting**: Providing a multi-tenant public hosted API or SaaS endpoint directly exposing the model's inference requires prior written approval from Bodhan AI (unless covered by the Open-Release Waiver in Section 4). 4. **Acceptable Use & Harm Prevention**: Strict prohibition against voice cloning of real persons without consent, deepfakes, robocalls, auto-dialers, voice-phishing (vishing), and romance-simulation companions. 5. **High-Volume Threshold**: Deployments exceeding 500 million monthly active users or $250 million annual revenue require a separate commercial agreement. For full legal terms, see [LICENSE](LICENSE) and [LICENSE_DEED.md](LICENSE_DEED.md).