Qwen3-TTS 12Hz 0.6B Base

Qwen's voice-cloning TTS talker, exported for loom.cpp. Family 10: text and a reference voice in, neural-codec tokens out -- pair it with qwen3-tts-tokenizer-12hz-loom for audio.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from Qwen/Qwen3-TTS-12Hz-0.6B-Base. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

apache-2.0, inherited from the base model above.

Language(s)

zh, en, ja, ko, de, fr, ru, pt, es, it

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom
import librosa

model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")

# The voice is a REFERENCE CLIP, not a speaker id -- this checkpoint carries no speaker table. A few
# clear seconds is enough. 24 kHz is what the speaker encoder expects, so resample on the way in.
reference, _ = librosa.load("reference.wav", sr=24000)

# What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
# `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
codes = model.text2codes.infer(
    "The quick brown fox jumps over the lazy dog.",
    waveform=reference.tolist(),
    language_id=2050,          # English; see "Known limitations" for the rest
    max_new_tokens=200, seed=1234,
)
print(len(codes), "frames x", len(codes[0]), "codebooks")

# The other half of the pair, in a repo of its own: the codec serves every size and variant of this
# talker, and the codes are worth having on their own -- cache them, edit them, decode them elsewhere.
codec = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-tokenizer-12hz-loom")
audio = codec.codes2speech.infer(codes)
print(len(audio), "samples at", audio.sample_rate, "Hz =", round(audio.duration, 2), "s")
audio.save("out.wav")

# Nothing goes between those two calls. Both files declare the width of a frame, so a pair that does
# not fit says so instead of producing audio of the wrong duration:
print(model.hparam("codec.n_codebooks"), "==", codec.hparam("codec.n_codebooks"))

# This model SAMPLES by default, at its own generation config's settings. `seed=` is what makes a
# take reproducible; pass temperature=0 for greedy, which reproduces `transformers` exactly.
print(model.hparam("sampling.temperature", "f32"),
      model.hparam("sampling.repetition_penalty", "f32"))

Cloning with the transcript (ICL)

Upstream's higher-fidelity mode: the clip's transcript and the clip's own codec tokens go into the prompt, as if the model had just said it. This file draws those tokens itself -- the codec's encoder rides inside it -- so what you add is the clip again as ref_audio= and what it says as ref_tokens=.

# What reference.wav says, word for word (this one is JFK's inaugural address, public domain).
reference_text = ("And so, my fellow Americans, ask not what your country can do for you; "
                  "ask what you can do for your country.")
codes = model.text2codes.infer(
    "The quick brown fox jumps over the lazy dog.",
    waveform=reference.tolist(),                  # the speaker's x-vector, as above
    ref_audio=reference.tolist(),                 # the same clip, encoded to codes by this file
    ref_tokens=model.tokenize(reference_text),    # and what it says
    language_id=2050, max_new_tokens=200, seed=1234,
)
audio = codec.codes2speech.infer(codes)
audio.save("icl.wav")

A few clear seconds is enough. The clip is used in whole 80 ms codec frames, and those frames count against the prompt at 12.5 per second.

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

This model does not produce audio. It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- qwen3-tts-tokenizer-12hz-loom, which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.

The voice comes from a reference clip, and there is no speaker table. This checkpoint's spk_id is empty, so cloning is the only mode: pass waveform= and the speaker encoder extracts an x-vector from it. Pass x_vector= instead to reuse one you already have -- it is 1024 floats and it is the whole of what the voice contributes. Copies of this file downloaded before 2026-09-29 abort the process on waveform=; download it again.

Pass the sentence as it is. The talker was trained on <|im_start|>assistant\n{text}<|im_end|>\n<|im_start|>assistant\n, and this file wraps whatever it is given in that template. Ids you have already wrapped yourself -- tokens= that open with <|im_start|>assistant\n -- are passed through as they are, and so is a reference transcript in ICL mode below.

The reference TEXT is used if you pass it -- that is ICL mode. Upstream's higher-fidelity clone mode conditions on a transcript of the reference clip plus the clip's own codec tokens, replayed as if the model had just spoken it. Pass ref_audio= (the same clip, 24 kHz) and ref_tokens= (its transcript, tokenized with this model's own tokenizer) and this file draws those codes itself: the codec's ENCODE half rides inside this GGUF, since the prompt is its only caller. Pass ref_code= instead if you already hold them, frame-major and sixteen wide. Omit both and you get the x-vector mode above, which is what the rest of this card describes.

A reference clip is used in whole codec frames. 1920 samples at 24 kHz, 80 ms: the driver trims to a multiple of that before encoding, so up to 79 ms of the clip's tail is not heard. Every convolution in the encode stack pads by a length-derived amount that would not otherwise trace, and on a frame boundary that amount is provably zero.

language_id is a raw number, because that is what the checkpoint declares -- English 2050, Chinese 2055, Spanish 2054, German 2053, Japanese 2058, French 2061, Korean 2064, Russian 2069, Portuguese 2071, Italian 2070. Omit it and English is assumed.

It samples by default. The export declares this checkpoint's own decoding -- temperature 0.9, top_k 50, repetition_penalty 1.05, and a second set for the code predictor -- so two runs of a sentence give two takes; seed= is what pins one. temperature=0 decodes greedily and reproduces transformers exactly: verified at 672 codes over 42 frames, every one identical, and the pair's audio transcribes back to the sentence it was given.

The repetition penalty is not optional here. transformers applies it as a processor rather than a warper, so it moves a greedy argmax too, and without it this model never emits its end token -- it runs to max_new_tokens. Passing repetition_penalty=1.0 turns it off and is a good way to see that.

max_new_tokens counts audio frames at 12.5 per second, not decoder steps -- one frame is sixteen transformer passes here, since a code predictor emits fifteen of the sixteen codebooks from the talker's hidden state.

Generation is slower than the parameter count suggests, for two reasons that are this export's rather than the model's. The code predictor runs without a KV cache (it re-reads its own sixteen-position prefix each step, which is what lets it share a file with a cached talker), and the attention is materialised as MHA rather than GQA -- k_proj/v_proj are duplicated so the key/value head count matches the query's, +69 M parameters and a doubled cache, because the grouped form does not survive conversion.

It is a big download: 3.9 GB, F32, like the rest of this collection. Nearly a third of it is the 151936 x 2048 text embedding table. loom-export --quantize Q8_0 on the upstream checkpoint packs the eligible weights to about 1.1 GB if you would rather have that.

Files

  • qwen3-tts-12hz-0.6b.gguf -- the model, exported with loom-exporter.
Downloads last month
434
GGUF
Model size
1B params
Architecture
loom-qwen3_tts
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/qwen3-tts-12hz-0.6b-loom

Quantized
(37)
this model

Collection including loom-ai-org/qwen3-tts-12hz-0.6b-loom