Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| gguf | 3 items | ||
| onnx | 5 items | ||
| voices | 33 items | ||
| .gitattributes | 386 Bytes xet | 9e830801 | |
| .gitignore | 11 Bytes xet | d60fd134 | |
| README.md | 13.4 kB xet | 2aded34f | |
| banner.png | 248 kB xet | 1890a381 | |
| config.json | 378 Bytes xet | 5f288a8f | |
| null_voice_emb.npy | 30.8 kB xet | 56667b6b | |
| silence_frame.npy | 256 Bytes xet | 55d2ecbf | |
| tokenizer.json | 192 kB xet | 4a7ee138 |
ZeroTTS — GGUF
The ZeroTTS weights packaged as GGUF, for the ggml runtime instead of onnxruntime. Same model, ~4× smaller download, and it runs in the browser under WebAssembly.
This is a repackaging, not a different model. The weights are the same ones
in zeroweight-ai/ZeroTTS; the
f32 file here is verified to produce bit-identical frame codes to the ONNX
graphs for the same text, voice and random draws. If you want the ONNX runtime
or the zerotts pip package, use that repo — this one is for the ggml/WASM
runtime in cpp/.
The most accurate open Vietnamese TTS we know of — 4× fewer word errors than the next best model, and it runs faster than real time on a laptop CPU.
🎯 Ultra-natural — 2.91 UTMOS above every other open Vietnamese system, with near-zero dead air (0.029 s).
🗣️ Zero-shot voice cloning — a voice is a small latent array; drop it in and the model speaks in it, cloned from as little as 3 seconds of reference audio (up to 30 seconds). No fine-tuning, no per-speaker training.
⚡ Real-time on CPU, streaming — ~2× faster than real time (RTF 0.5×), first audio chunk in ~70 ms. No GPU required.
🇻🇳 Built for Vietnamese — tones, code-switched English, and reads
31/12/2025andZeroTTSwithout text normalizer.Code, examples, browser demo: https://github.com/zeroweight-ai/ZeroTTS
Benchmark dataset: https://huggingface.co/datasets/zeroweight-ai/ZeroBench-TTS
Blogpost: https://zeroweight.ai/blog/zero-tts
Samples
Two-speaker conversation
Long-form narration
News read, code-switched English
Cross-lingual
Reference audio (Vietnamese)
Output (English)
Files
| path | what it is |
|---|---|
gguf/zerotts-q8_0.gguf |
the default — 206 MB, the text encoder, global decoder and depth decoder |
gguf/zerotts-q4_0.gguf |
124 MB, when the download budget is tight |
gguf/zerotts-f32.gguf |
772 MB, unquantized; the reference the others are measured against |
onnx/codec/ |
the MOSS audio-tokenizer decoder, still ONNX (Apache-2.0, ~45 MB) — it is not on the per-frame hot path |
voices/ |
the voice latents; a voice is a 10×768 float32 array |
tokenizer.json, config.json, null_voice_emb.npy, silence_frame.npy |
runtime side files, identical to the ONNX repo |
You need one .gguf plus onnx/codec/, voices/ and the side files.
Which one to take
Measured in Chrome on a 14-core Apple-silicon laptop, one 3.2 s utterance,
frame generation only. "drift" is the share of sampled codes that come out
differently from f32 when both are driven along the same code sequence — the
<eoa> stop decision was identical in every build.
| file | size | drift vs f32 | 1 thread | 4 threads | 8 threads |
|---|---|---|---|---|---|
f32 |
772 MB | — | 2.40× | 6.1× | 6.2× |
q8_0 |
206 MB | 4.9% | 1.06× | 3.6× | 5.3× |
q4_0 |
124 MB | 36.5% | 1.27× | 4.1× | 5.8× |
For reference, the ONNX graphs in
zeroweight-ai/ZeroTTS run at
4.2× on the same machine, at the 4 threads onnxruntime-web caps itself to.
Quantization here buys download size, not speed. WebAssembly's SIMD has no
integer dot-product instruction, so a q8_0 block costs roughly 20 SIMD ops per
32 multiply-accumulates where f32 costs 8 fused multiply-adds — on a machine
with memory bandwidth to spare the extra instructions cost more than the saved
bytes. Take q8_0 unless you have a reason not to: it is 4× smaller than the
ONNX graphs and faster than them at the thread counts each actually uses (5.3×
on 8 threads against 4.2× on 4). Take f32 if you want bit-exactness, the
highest speed at any thread count, or you are running on very few cores — it is
the fastest build here, just the largest. q4_0 changes a third of the drawn
codes, which is audible as a different take rather than a broken one, but it is
the wrong trade unless 124 MB is the binding constraint.
Numbers on hardware that is genuinely short of bandwidth rather than compute should look different; the benchmark page in the repo re-runs this table in one click.
Usage
The runtime is cpp/
in the main repository — a ggml implementation of the generation loop, building
natively and to WebAssembly.
git clone --recurse-submodules https://github.com/zeroweight-ai/ZeroTTS
cd ZeroTTS/cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release . && cmake --build build -j
./build/zerotts-bench /path/to/zerotts-q8_0.gguf voice.bin text_ids.txt
In the browser, the demo in js/
uses this repo by default:
cd ZeroTTS/js && npm install && npm run dev
voice.bin is the raw little-endian float32 blob under voices/<name>/ —
voice.npz is the same latents for numpy callers.
For the Python package, the ONNX graphs and pip install zerotts, use
zeroweight-ai/ZeroTTS instead.
Benchmarks
These are the unquantized model's quality numbers, measured through the
PyTorch runtime. f32 here is that model exactly; see the drift column above
for what the quantized files change.
Measured on ZeroBench-TTS
Every system reads raw text — dates, numbers and acronyms verbatim, exactly as they appear in the wild, with no text frontend in front of the model.
| ZeroTTS | OmniVoice | XTTS-v2-vietnamse | viXTTS | |
|---|---|---|---|---|
| WER ↓ | 1.03 % | 4.13 % | 16.42 % | 18.40 % |
| Naturalness (UTMOS) ↑ | 2.91 | 2.76 | 2.43 | 2.35 |
| Voice similarity (SSIM) ↑ | 0.936 | 0.950 | 0.940 | 0.935 |
| Dead air (excess silence) ↓ | 0.029 s | 0.340 s | 0.532 s | 0.233 s |
| RTF, CPU ↓ | 0.50× | 6.12× | 0.71× | 0.73× |
| Time to first audio, CPU ↓ | ~70 ms | ~34 s | ~6.1 s | ~5.1 s |
| Parameters ↓ | 202 M | 775 M | 467 M | 467 M |
4× fewer word errors than the next-best system, and the fastest of the four on CPU. The gap is much wider in latency than in throughput: the two XTTS fine-tunes also beat real time (0.71×) but need seconds to emit their first sample, while OmniVoice is 6× slower than real time. All three are sized and tuned for a GPU, and it shows.
Full comparison tables, per-subset breakdowns, and CPU speed methodology: docs/BENCHMARKS.md
Speed — CPU
RTF (realtime factor, wall-clock synthesis time ÷ output audio duration — lower is faster; below 1× is faster than real time) and time-to-first-audio, all measured on CPU, single request, 8 inference threads pinned to a dedicated core pool (no other synthesis running concurrently). Three Vietnamese samples — short (26 chars), medium (77 chars), long (227 chars) — each run 6 times with the first 2 (cold-cache) discarded; figures below are the mean of the remaining 4.
| ZeroTTS | OmniVoice | XTTS-v2-vietnamse | viXTTS | |
|---|---|---|---|---|
| RTF — short | 0.51× | 10.87× | 0.70× | 0.71× |
| RTF — medium | 0.47× | 4.82× | 0.70× | 0.70× |
| RTF — long | 0.53× | 2.67× | 0.71× | 0.78× |
| TTFA — short | 53 ms | 21.7 s | 4.02 s | 2.45 s |
| TTFA — medium | 66 ms | 28.9 s | 4.02 s | 3.72 s |
| TTFA — long | 89 ms | 52.3 s | 10.3 s | 9.22 s |
ZeroTTS's time-to-first-audio comes from its real streaming path — first audio frame, not first full utterance. The three baselines have no working CPU streaming path, so their TTFA is the time to the complete utterance.
Voices, and voice cloning
A voice is a small array of speaker latents, (1, n_voice_queries, d_model),
shipped as a .npz under voices/. That array is the entire speaker
conditioning — no reference transcript, no audio prompt.
Voice cloning is not available in this release. Those latents come from a voice encoder that reads a reference clip, and that encoder is not published. This repository ships ready-to-use voices; it cannot create new ones from audio.
To get latents for your own speaker, see zeroweight.ai or get in touch.
Because a voice is just an array, latents obtained that way drop into
voices/<name>/voice.npz and work with no code change.
Intended use and limitations
Built for Vietnamese. It handles English words embedded in Vietnamese text
(code_switch), but it is not an English TTS system and is not evaluated as one.
Do not use it to impersonate a real person, to generate speech attributed to someone without their consent, or to produce audio intended to deceive. The shipped voices are for evaluation and demos.
Synthetic speech should be disclosed as synthetic wherever a listener might reasonably assume otherwise.
Credits
Speech codec: MOSS-Audio-Tokenizer-Nano by the OpenMOSS team, Apache-2.0.
Its ONNX decoder graphs are redistributed under onnx/codec/ so ZeroTTS has
no external runtime dependency; the encoder is not included. See
onnx/codec/LICENSE-Apache-2.0.txt.
@misc{gong2026mossaudiotokenizerscalingaudiotokenizers,
title={MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models},
author={Yitian Gong and Kuangwei Chen and Zhaoye Fei and Xiaogui Yang and Ke Chen
and Yang Wang and Kexin Huang and Mingshu Chen and Ruixiao Li
and Qingyuan Cheng and Shimin Li and Xipeng Qiu},
year={2026}, eprint={2602.10934}, archivePrefix={arXiv}, primaryClass={cs.SD}
}
License
ZeroTTS weights and code: MIT.
The ZeroBench-TTS dataset is CC-BY-NC-4.0 because it redistributes reference audio from VIVOS, viVoice, phoaudiobook and Emilia. That license applies to the benchmark dataset only — not to these weights.
- Total size
- 1.21 GB
- Files
- 49
- Last updated
- Sep 15
- Pre-warmed CDN
- US EU US EU