# Notices & Attribution ## 1. Bodhan AI / AI4Bharat Indic-Speak This repository provides quantized (INT8 ONNX and GGUF) derivatives of **Indic-Speak (preview v2)**, developed by **Bodhan AI** and **AI4Bharat** at the Center of Excellence for Artificial Intelligence in Education (supported by the Ministry of Education, Govt. of India). - Upstream Model: [bodhan-ai/indic-speak-preview-v2](https://huggingface.co/bodhan-ai/indic-speak-preview-v2) - Upstream Commit: `c36d2c49c18d361860a4f5f540b686e066a362bf` - License: Indic Open Model License v1.0 (see [LICENSE](LICENSE) and [LICENSE_DEED.md](LICENSE_DEED.md)) - Scope: Zero-shot, highly expressive multilingual and cross-lingual text-to-speech synthesis supporting 23 Indian languages + English, across 46+ curated speaker voices. - Mandatory Attribution Notice: **"Built with Indic-Speak from Bodhan AI / AI4Bharat."** ## 2. Base Architecture Credits 1. **Autoregressive Acoustic Transformer**: Based on **Meta Llama 3.2 3B** (~3.21B parameters), fine-tuned for generating discrete speech tokens from phonetic text input. - Base Model: [meta-llama/Llama-3.2-3B](https://huggingface.co/meta-llama/Llama-3.2-3B) - Base Model License: [Llama 3.2 Community License Agreement](https://llama.meta.com/llama3/license/) 2. **Neural Vocoder**: Based on **Vocos** architecture, designed by Charactr Inc. - Upstream Vocos License: Apache 2.0 - Function: Reconstructs 24 kHz high-fidelity continuous waveform audio from acoustic token embeddings via inverse Short-Time Fourier Transform (iSTFT) with 257 frequency bins (n_fft=512, hop_length=128). ## 3. Derivative Modifications - **GGUF Quantization of Llama 3.2 Backbone**: - `indic-speak-q8_0.gguf` (3.27 GB) and `indic-speak-q5_k_m.gguf` (2.23 GB), reducing memory footprint by up to 66.2% from the 6.6 GB FP32 checkpoint while retaining near-lossless acoustic token probabilities. - **ONNX INT8 Quantization of Vocos Neural Vocoder**: - `vocos_backbone_int8.onnx`: Quantized from 222.7 MB FP32 down to **79.8 MB INT8**, achieving ultra-fast **13.3 ms CPU latency** per chunk on standard edge processors. - Native export: Clean graph without operator warnings, compatible with ONNX Runtime CPU and DirectML. - **Target Distribution**: Client-side local inference for private voice assistants, reading tools, on-device screen readers, and low-latency interactive audio systems. No multi-tenant hosted API services are provided under this distribution. - **Derivative Release**: Packaged and published under the Indic Open Model License v1.0 by Community Contributor (2026).