adidsh commited on
Commit
ae16006
·
verified ·
1 Parent(s): 4b4c76f

Upload NOTICE.md

Browse files
Files changed (1) hide show
  1. NOTICE.md +26 -0
NOTICE.md ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Notices & Attribution
2
+
3
+ ## 1. Bodhan AI / AI4Bharat Indic-Speak
4
+ This repository provides quantized (INT8 ONNX and GGUF) derivatives of **Indic-Speak (preview v2)**, developed by **Bodhan AI** and **AI4Bharat** at the Center of Excellence for Artificial Intelligence in Education (supported by the Ministry of Education, Govt. of India).
5
+ - Upstream Model: [bodhan-ai/indic-speak-preview-v2](https://huggingface.co/bodhan-ai/indic-speak-preview-v2)
6
+ - Upstream Commit: `c36d2c49c18d361860a4f5f540b686e066a362bf`
7
+ - License: Indic Open Model License v1.0 (see [LICENSE](LICENSE) and [LICENSE_DEED.md](LICENSE_DEED.md))
8
+ - Scope: Zero-shot, highly expressive multilingual and cross-lingual text-to-speech synthesis supporting 23 Indian languages + English, across 46+ curated speaker voices.
9
+ - Mandatory Attribution Notice: **"Built with Indic-Speak from Bodhan AI / AI4Bharat."**
10
+
11
+ ## 2. Base Architecture Credits
12
+ 1. **Autoregressive Acoustic Transformer**: Based on **Meta Llama 3.2 3B** (~3.21B parameters), fine-tuned for generating discrete speech tokens from phonetic text input.
13
+ - Base Model: [meta-llama/Llama-3.2-3B](https://huggingface.co/meta-llama/Llama-3.2-3B)
14
+ - Base Model License: [Llama 3.2 Community License Agreement](https://llama.meta.com/llama3/license/)
15
+ 2. **Neural Vocoder**: Based on **Vocos** architecture, designed by Charactr Inc.
16
+ - Upstream Vocos License: Apache 2.0
17
+ - Function: Reconstructs 24 kHz high-fidelity continuous waveform audio from acoustic token embeddings via inverse Short-Time Fourier Transform (iSTFT) with 257 frequency bins (n_fft=512, hop_length=128).
18
+
19
+ ## 3. Derivative Modifications
20
+ - **GGUF Quantization of Llama 3.2 Backbone**:
21
+ - `indic-speak-q8_0.gguf` (3.27 GB) and `indic-speak-q5_k_m.gguf` (2.23 GB), reducing memory footprint by up to 66.2% from the 6.6 GB FP32 checkpoint while retaining near-lossless acoustic token probabilities.
22
+ - **ONNX INT8 Quantization of Vocos Neural Vocoder**:
23
+ - `vocos_backbone_int8.onnx`: Quantized from 222.7 MB FP32 down to **79.8 MB INT8**, achieving ultra-fast **13.3 ms CPU latency** per chunk on standard edge processors.
24
+ - Native export: Clean graph without operator warnings, compatible with ONNX Runtime CPU and DirectML.
25
+ - **Target Distribution**: Client-side local inference for private voice assistants, reading tools, on-device screen readers, and low-latency interactive audio systems. No multi-tenant hosted API services are provided under this distribution.
26
+ - **Derivative Release**: Packaged and published under the Indic Open Model License v1.0 by Community Contributor (2026).