--- license: apache-2.0 library_name: transformers pipeline_tag: text-to-speech language: - en - zh base_model: - Qwen/Qwen3-TTS-12Hz-0.6B-Base base_model_relation: finetune tags: - zen3 - zen3-tts - zenlm - qwen3-tts - speech-synthesis - edge - low-latency - 12hz --- # Zen3 TTS 0.6B Compact ~0.6B Zen3 text-to-speech model at 12 Hz, sized for edge and latency-sensitive synthesis. Derived by fine-tuning [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base) (Alibaba Cloud, Apache-2.0). - Architecture: `Qwen3TTSForConditionalGeneration` (`qwen3_tts`) - Parameters: ~0.6B - Base model: [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base) ## Weights This repository contains the model weights: `model.safetensors` (talker) plus a `speech_tokenizer/` module (12 Hz codec), config and tokenizer files. The model uses the `qwen3_tts` architecture and loads with `transformers` (>= 4.57). It is API-compatible with the upstream base — follow the inference recipe on the base model card [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base). ## Provenance Fine-tuned from [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base) (Apache-2.0). See `NOTICE` for full attribution.