GGUF
tts
qwen3-tts

Qwen3-TTS-12Hz-0.6B-Base GGUF (llama.cpp official conversion)

Converted using the official conversion/qwen3tts.py in ggml-org/llama.cpp (mtmd Qwen3-TTS PR #26254), from Qwen/Qwen3-TTS-12Hz-0.6B-Base.

Files

  • qwen-talker-0.6b-base-Q8_0.gguf โ€” backbone (talker), load with llama_model_load_from_file
  • qwen-tokenizer-12hz-f16.gguf โ€” tokenizer/codec (mmproj), load with mtmd_init_from_file

Usage

llama-tts -m qwen-talker-0.6b-base-Q8_0.gguf -mm qwen-tokenizer-12hz-f16.gguf -p "Hello world" --output out.wav

Other GGUF conversions of this model floating around (e.g. koboldcpp-oriented exports) use non-standard metadata key names (qwen3-tts.* with hyphen, talker. subsections, general.file_type stored as string) that are incompatible with the official llama.cpp qwen3tts architecture loader. This repo uses the stock conversion script, so the metadata matches what llama.cpp expects out of the box.

Downloads last month
163
GGUF
Model size
0.6B params
Architecture
qwen3tts
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Mouserat/qwen3-tts-0.6b-base-gguf

Quantized
(34)
this model