1.21 GB
49 files
Updated 16 days ago
Name
Size
gguf
onnx
voices
.gitattributes386 Bytes
xet
.gitignore11 Bytes
xet
README.md13.4 kB
xet
banner.png248 kB
xet
config.json378 Bytes
xet
null_voice_emb.npy30.8 kB
xet
silence_frame.npy256 Bytes
xet
tokenizer.json192 kB
xet
README.md
ZeroTTS — Vietnamese zero-shot text-to-speech

ZeroTTS — GGUF

The ZeroTTS weights packaged as GGUF, for the ggml runtime instead of onnxruntime. Same model, ~4× smaller download, and it runs in the browser under WebAssembly.

This is a repackaging, not a different model. The weights are the same ones in zeroweight-ai/ZeroTTS; the f32 file here is verified to produce bit-identical frame codes to the ONNX graphs for the same text, voice and random draws. If you want the ONNX runtime or the zerotts pip package, use that repo — this one is for the ggml/WASM runtime in cpp/.

The most accurate open Vietnamese TTS we know of — 4× fewer word errors than the next best model, and it runs faster than real time on a laptop CPU.

  • 🎯 Ultra-natural — 2.91 UTMOS above every other open Vietnamese system, with near-zero dead air (0.029 s).

  • 🗣️ Zero-shot voice cloning — a voice is a small latent array; drop it in and the model speaks in it, cloned from as little as 3 seconds of reference audio (up to 30 seconds). No fine-tuning, no per-speaker training.

  • ⚡ Real-time on CPU, streaming — ~2× faster than real time (RTF 0.5×), first audio chunk in ~70 ms. No GPU required.

  • 🇻🇳 Built for Vietnamese — tones, code-switched English, and reads 31/12/2025 and ZeroTTS without text normalizer.

  • Code, examples, browser demo: https://github.com/zeroweight-ai/ZeroTTS

  • Benchmark dataset: https://huggingface.co/datasets/zeroweight-ai/ZeroBench-TTS

  • Blogpost: https://zeroweight.ai/blog/zero-tts

Samples

Two-speaker conversation

Long-form narration

News read, code-switched English

Cross-lingual

Reference audio (Vietnamese)

Output (English)

Files

path what it is
gguf/zerotts-q8_0.gguf the default — 206 MB, the text encoder, global decoder and depth decoder
gguf/zerotts-q4_0.gguf 124 MB, when the download budget is tight
gguf/zerotts-f32.gguf 772 MB, unquantized; the reference the others are measured against
onnx/codec/ the MOSS audio-tokenizer decoder, still ONNX (Apache-2.0, ~45 MB) — it is not on the per-frame hot path
voices/ the voice latents; a voice is a 10×768 float32 array
tokenizer.json, config.json, null_voice_emb.npy, silence_frame.npy runtime side files, identical to the ONNX repo

You need one .gguf plus onnx/codec/, voices/ and the side files.

Which one to take

Measured in Chrome on a 14-core Apple-silicon laptop, one 3.2 s utterance, frame generation only. "drift" is the share of sampled codes that come out differently from f32 when both are driven along the same code sequence — the <eoa> stop decision was identical in every build.

file size drift vs f32 1 thread 4 threads 8 threads
f32 772 MB — 2.40× 6.1× 6.2×
q8_0 206 MB 4.9% 1.06× 3.6× 5.3×
q4_0 124 MB 36.5% 1.27× 4.1× 5.8×

For reference, the ONNX graphs in zeroweight-ai/ZeroTTS run at 4.2× on the same machine, at the 4 threads onnxruntime-web caps itself to.

Quantization here buys download size, not speed. WebAssembly's SIMD has no integer dot-product instruction, so a q8_0 block costs roughly 20 SIMD ops per 32 multiply-accumulates where f32 costs 8 fused multiply-adds — on a machine with memory bandwidth to spare the extra instructions cost more than the saved bytes. Take q8_0 unless you have a reason not to: it is 4× smaller than the ONNX graphs and faster than them at the thread counts each actually uses (5.3× on 8 threads against 4.2× on 4). Take f32 if you want bit-exactness, the highest speed at any thread count, or you are running on very few cores — it is the fastest build here, just the largest. q4_0 changes a third of the drawn codes, which is audible as a different take rather than a broken one, but it is the wrong trade unless 124 MB is the binding constraint.

Numbers on hardware that is genuinely short of bandwidth rather than compute should look different; the benchmark page in the repo re-runs this table in one click.

Usage

The runtime is cpp/ in the main repository — a ggml implementation of the generation loop, building natively and to WebAssembly.

git clone --recurse-submodules https://github.com/zeroweight-ai/ZeroTTS
cd ZeroTTS/cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release . && cmake --build build -j
./build/zerotts-bench /path/to/zerotts-q8_0.gguf voice.bin text_ids.txt

In the browser, the demo in js/ uses this repo by default:

cd ZeroTTS/js && npm install && npm run dev

voice.bin is the raw little-endian float32 blob under voices/<name>/ — voice.npz is the same latents for numpy callers.

For the Python package, the ONNX graphs and pip install zerotts, use zeroweight-ai/ZeroTTS instead.

Benchmarks

These are the unquantized model's quality numbers, measured through the PyTorch runtime. f32 here is that model exactly; see the drift column above for what the quantized files change.

Measured on ZeroBench-TTS

Every system reads raw text — dates, numbers and acronyms verbatim, exactly as they appear in the wild, with no text frontend in front of the model.

ZeroTTS OmniVoice XTTS-v2-vietnamse viXTTS
WER ↓ 1.03 % 4.13 % 16.42 % 18.40 %
Naturalness (UTMOS) ↑ 2.91 2.76 2.43 2.35
Voice similarity (SSIM) ↑ 0.936 0.950 0.940 0.935
Dead air (excess silence) ↓ 0.029 s 0.340 s 0.532 s 0.233 s
RTF, CPU ↓ 0.50× 6.12× 0.71× 0.73×
Time to first audio, CPU ↓ ~70 ms ~34 s ~6.1 s ~5.1 s
Parameters ↓ 202 M 775 M 467 M 467 M

4× fewer word errors than the next-best system, and the fastest of the four on CPU. The gap is much wider in latency than in throughput: the two XTTS fine-tunes also beat real time (0.71×) but need seconds to emit their first sample, while OmniVoice is 6× slower than real time. All three are sized and tuned for a GPU, and it shows.

Full comparison tables, per-subset breakdowns, and CPU speed methodology: docs/BENCHMARKS.md

Speed — CPU

RTF (realtime factor, wall-clock synthesis time ÷ output audio duration — lower is faster; below 1× is faster than real time) and time-to-first-audio, all measured on CPU, single request, 8 inference threads pinned to a dedicated core pool (no other synthesis running concurrently). Three Vietnamese samples — short (26 chars), medium (77 chars), long (227 chars) — each run 6 times with the first 2 (cold-cache) discarded; figures below are the mean of the remaining 4.

ZeroTTS OmniVoice XTTS-v2-vietnamse viXTTS
RTF — short 0.51× 10.87× 0.70× 0.71×
RTF — medium 0.47× 4.82× 0.70× 0.70×
RTF — long 0.53× 2.67× 0.71× 0.78×
TTFA — short 53 ms 21.7 s 4.02 s 2.45 s
TTFA — medium 66 ms 28.9 s 4.02 s 3.72 s
TTFA — long 89 ms 52.3 s 10.3 s 9.22 s

ZeroTTS's time-to-first-audio comes from its real streaming path — first audio frame, not first full utterance. The three baselines have no working CPU streaming path, so their TTFA is the time to the complete utterance.

Voices, and voice cloning

A voice is a small array of speaker latents, (1, n_voice_queries, d_model), shipped as a .npz under voices/. That array is the entire speaker conditioning — no reference transcript, no audio prompt.

Voice cloning is not available in this release. Those latents come from a voice encoder that reads a reference clip, and that encoder is not published. This repository ships ready-to-use voices; it cannot create new ones from audio.

To get latents for your own speaker, see zeroweight.ai or get in touch.

Because a voice is just an array, latents obtained that way drop into voices/<name>/voice.npz and work with no code change.

Intended use and limitations

Built for Vietnamese. It handles English words embedded in Vietnamese text (code_switch), but it is not an English TTS system and is not evaluated as one.

Do not use it to impersonate a real person, to generate speech attributed to someone without their consent, or to produce audio intended to deceive. The shipped voices are for evaluation and demos.

Synthetic speech should be disclosed as synthetic wherever a listener might reasonably assume otherwise.

Credits

Speech codec: MOSS-Audio-Tokenizer-Nano by the OpenMOSS team, Apache-2.0. Its ONNX decoder graphs are redistributed under onnx/codec/ so ZeroTTS has no external runtime dependency; the encoder is not included. See onnx/codec/LICENSE-Apache-2.0.txt.

@misc{gong2026mossaudiotokenizerscalingaudiotokenizers,
  title={MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models},
  author={Yitian Gong and Kuangwei Chen and Zhaoye Fei and Xiaogui Yang and Ke Chen
          and Yang Wang and Kexin Huang and Mingshu Chen and Ruixiao Li
          and Qingyuan Cheng and Shimin Li and Xipeng Qiu},
  year={2026}, eprint={2602.10934}, archivePrefix={arXiv}, primaryClass={cs.SD}
}

License

ZeroTTS weights and code: MIT.

The ZeroBench-TTS dataset is CC-BY-NC-4.0 because it redistributes reference audio from VIVOS, viVoice, phoaudiobook and Emilia. That license applies to the benchmark dataset only — not to these weights.

Total size
1.21 GB
Files
49
Last updated
Sep 15
Pre-warmed CDN
US EU US EU

Contributors