Text-to-Speech
ONNX
GGUF
tts
speech
audio
indic
multilingual
int8
quantization
edge-ai
vocos
llama-3.2
indic-speak
Instructions to use adidsh/indic-speak-int8-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use adidsh/indic-speak-int8-onnx with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf adidsh/indic-speak-int8-onnx:Q5_K_M # Run inference directly in the terminal: llama cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf adidsh/indic-speak-int8-onnx:Q5_K_M # Run inference directly in the terminal: llama cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf adidsh/indic-speak-int8-onnx:Q5_K_M # Run inference directly in the terminal: ./llama-cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf adidsh/indic-speak-int8-onnx:Q5_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf adidsh/indic-speak-int8-onnx:Q5_K_M
Use Docker
docker model run hf.co/adidsh/indic-speak-int8-onnx:Q5_K_M
- LM Studio
- Jan
- Ollama
How to use adidsh/indic-speak-int8-onnx with Ollama:
ollama run hf.co/adidsh/indic-speak-int8-onnx:Q5_K_M
- Unsloth Desktop
- Docker Model Runner
How to use adidsh/indic-speak-int8-onnx with Docker Model Runner:
docker model run hf.co/adidsh/indic-speak-int8-onnx:Q5_K_M
- Lemonade
How to use adidsh/indic-speak-int8-onnx with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull adidsh/indic-speak-int8-onnx:Q5_K_M
Run and chat with the model
lemonade run user.indic-speak-int8-onnx-Q5_K_M
List all available models
lemonade list
- Atomic Chat
Upload NOTICE.md
Browse files
NOTICE.md
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Notices & Attribution
|
| 2 |
+
|
| 3 |
+
## 1. Bodhan AI / AI4Bharat Indic-Speak
|
| 4 |
+
This repository provides quantized (INT8 ONNX and GGUF) derivatives of **Indic-Speak (preview v2)**, developed by **Bodhan AI** and **AI4Bharat** at the Center of Excellence for Artificial Intelligence in Education (supported by the Ministry of Education, Govt. of India).
|
| 5 |
+
- Upstream Model: [bodhan-ai/indic-speak-preview-v2](https://huggingface.co/bodhan-ai/indic-speak-preview-v2)
|
| 6 |
+
- Upstream Commit: `c36d2c49c18d361860a4f5f540b686e066a362bf`
|
| 7 |
+
- License: Indic Open Model License v1.0 (see [LICENSE](LICENSE) and [LICENSE_DEED.md](LICENSE_DEED.md))
|
| 8 |
+
- Scope: Zero-shot, highly expressive multilingual and cross-lingual text-to-speech synthesis supporting 23 Indian languages + English, across 46+ curated speaker voices.
|
| 9 |
+
- Mandatory Attribution Notice: **"Built with Indic-Speak from Bodhan AI / AI4Bharat."**
|
| 10 |
+
|
| 11 |
+
## 2. Base Architecture Credits
|
| 12 |
+
1. **Autoregressive Acoustic Transformer**: Based on **Meta Llama 3.2 3B** (~3.21B parameters), fine-tuned for generating discrete speech tokens from phonetic text input.
|
| 13 |
+
- Base Model: [meta-llama/Llama-3.2-3B](https://huggingface.co/meta-llama/Llama-3.2-3B)
|
| 14 |
+
- Base Model License: [Llama 3.2 Community License Agreement](https://llama.meta.com/llama3/license/)
|
| 15 |
+
2. **Neural Vocoder**: Based on **Vocos** architecture, designed by Charactr Inc.
|
| 16 |
+
- Upstream Vocos License: Apache 2.0
|
| 17 |
+
- Function: Reconstructs 24 kHz high-fidelity continuous waveform audio from acoustic token embeddings via inverse Short-Time Fourier Transform (iSTFT) with 257 frequency bins (n_fft=512, hop_length=128).
|
| 18 |
+
|
| 19 |
+
## 3. Derivative Modifications
|
| 20 |
+
- **GGUF Quantization of Llama 3.2 Backbone**:
|
| 21 |
+
- `indic-speak-q8_0.gguf` (3.27 GB) and `indic-speak-q5_k_m.gguf` (2.23 GB), reducing memory footprint by up to 66.2% from the 6.6 GB FP32 checkpoint while retaining near-lossless acoustic token probabilities.
|
| 22 |
+
- **ONNX INT8 Quantization of Vocos Neural Vocoder**:
|
| 23 |
+
- `vocos_backbone_int8.onnx`: Quantized from 222.7 MB FP32 down to **79.8 MB INT8**, achieving ultra-fast **13.3 ms CPU latency** per chunk on standard edge processors.
|
| 24 |
+
- Native export: Clean graph without operator warnings, compatible with ONNX Runtime CPU and DirectML.
|
| 25 |
+
- **Target Distribution**: Client-side local inference for private voice assistants, reading tools, on-device screen readers, and low-latency interactive audio systems. No multi-tenant hosted API services are provided under this distribution.
|
| 26 |
+
- **Derivative Release**: Packaged and published under the Indic Open Model License v1.0 by Community Contributor (2026).
|