Re-export qwen3-tts-12hz-0.6b (loom-exporter f1c2bbc)
Browse files- .gitattributes +1 -0
- README.md +121 -0
- qwen3-tts-12hz-0.6b.gguf +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
qwen3-tts-12hz-0.6b.gguf filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- zh
|
| 5 |
+
- en
|
| 6 |
+
- ja
|
| 7 |
+
- ko
|
| 8 |
+
- de
|
| 9 |
+
- fr
|
| 10 |
+
- ru
|
| 11 |
+
- pt
|
| 12 |
+
- es
|
| 13 |
+
- it
|
| 14 |
+
base_model:
|
| 15 |
+
- Qwen/Qwen3-TTS-12Hz-0.6B-Base
|
| 16 |
+
pipeline_tag: text-to-speech
|
| 17 |
+
tags:
|
| 18 |
+
- loom
|
| 19 |
+
- text-to-codes
|
| 20 |
+
library_name: loom-py-rt
|
| 21 |
+
---
|
| 22 |
+
# Qwen3-TTS 12Hz 0.6B Base
|
| 23 |
+
|
| 24 |
+
Qwen's voice-cloning TTS talker, exported for loom.cpp. Family 10: text and a reference voice in, neural-codec tokens out -- pair it with `qwen3-tts-tokenizer-12hz-loom` for audio.
|
| 25 |
+
|
| 26 |
+
This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
|
| 27 |
+
that carries its own graph topologies, tokenizer (if any) and driver script, produced by
|
| 28 |
+
[loom-exporter](https://github.com/loom-ai-org/loom-exporter).
|
| 29 |
+
|
| 30 |
+
## Original model
|
| 31 |
+
|
| 32 |
+
Exported from [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base). Weights are unmodified; this repo packages the same parameters into
|
| 33 |
+
loom.cpp's GGUF format.
|
| 34 |
+
|
| 35 |
+
## License
|
| 36 |
+
|
| 37 |
+
`apache-2.0`, inherited from the base model above.
|
| 38 |
+
|
| 39 |
+
## Language(s)
|
| 40 |
+
|
| 41 |
+
`zh`, `en`, `ja`, `ko`, `de`, `fr`, `ru`, `pt`, `es`, `it`
|
| 42 |
+
|
| 43 |
+
## Usage
|
| 44 |
+
|
| 45 |
+
Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:
|
| 46 |
+
|
| 47 |
+
```sh
|
| 48 |
+
pip install -U "loom-py-rt[hub]"
|
| 49 |
+
```
|
| 50 |
+
|
| 51 |
+
```python
|
| 52 |
+
import loom
|
| 53 |
+
import librosa
|
| 54 |
+
|
| 55 |
+
model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")
|
| 56 |
+
|
| 57 |
+
# The voice is a REFERENCE CLIP, not a speaker id -- this checkpoint carries no speaker table. A few
|
| 58 |
+
# clear seconds is enough. 24 kHz is what the speaker encoder expects, so resample on the way in.
|
| 59 |
+
reference, _ = librosa.load("reference.wav", sr=24000)
|
| 60 |
+
|
| 61 |
+
# What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
|
| 62 |
+
# `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
|
| 63 |
+
codes = model.text2codes.infer(
|
| 64 |
+
"The quick brown fox jumps over the lazy dog.",
|
| 65 |
+
waveform=reference.tolist(),
|
| 66 |
+
language_id=2050, # English; see "Known limitations" for the rest
|
| 67 |
+
max_new_tokens=200, seed=1234,
|
| 68 |
+
)
|
| 69 |
+
print(len(codes), "frames x", len(codes[0]), "codebooks")
|
| 70 |
+
|
| 71 |
+
# The other half of the pair, in a repo of its own: the codec serves every size and variant of this
|
| 72 |
+
# talker, and the codes are worth having on their own -- cache them, edit them, decode them elsewhere.
|
| 73 |
+
codec = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-tokenizer-12hz-loom")
|
| 74 |
+
audio = codec.codes2speech.infer(codes)
|
| 75 |
+
print(len(audio), "samples at", audio.sample_rate, "Hz =", round(audio.duration, 2), "s")
|
| 76 |
+
audio.save("out.wav")
|
| 77 |
+
|
| 78 |
+
# Nothing goes between those two calls. Both files declare the width of a frame, so a pair that does
|
| 79 |
+
# not fit says so instead of producing audio of the wrong duration:
|
| 80 |
+
print(model.hparam("codec.n_codebooks"), "==", codec.hparam("codec.n_codebooks"))
|
| 81 |
+
|
| 82 |
+
# This model SAMPLES by default, at its own generation config's settings. `seed=` is what makes a
|
| 83 |
+
# take reproducible; pass temperature=0 for greedy, which reproduces `transformers` exactly.
|
| 84 |
+
print(model.hparam("sampling.temperature", "f32"),
|
| 85 |
+
model.hparam("sampling.repetition_penalty", "f32"))
|
| 86 |
+
```
|
| 87 |
+
|
| 88 |
+
### The layer underneath
|
| 89 |
+
|
| 90 |
+
The call above is the high-level door: one per task, named for the modality pair it maps between, with
|
| 91 |
+
the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
|
| 92 |
+
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
|
| 93 |
+
door does not name.
|
| 94 |
+
|
| 95 |
+
`model.driver_source` prints that driver, including a header comment documenting every argument it
|
| 96 |
+
accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
|
| 97 |
+
[loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.
|
| 98 |
+
|
| 99 |
+
## Known limitations
|
| 100 |
+
|
| 101 |
+
**This model does not produce audio.** It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- [`qwen3-tts-tokenizer-12hz-loom`](https://huggingface.co/loom-ai-org/qwen3-tts-tokenizer-12hz-loom), which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.
|
| 102 |
+
|
| 103 |
+
**The voice comes from a reference clip, and there is no speaker table.** This checkpoint's `spk_id` is empty, so cloning is the only mode: pass `waveform=` and the speaker encoder extracts an x-vector from it. Pass `x_vector=` instead to reuse one you already have -- it is 1024 floats and it is the whole of what the voice contributes.
|
| 104 |
+
|
| 105 |
+
**The reference TEXT is not used.** Upstream offers a second, higher-fidelity clone mode that conditions on a transcript of the reference clip plus its codec tokens; it needs the codec's ENCODE half, which this collection does not export, so this file implements the x-vector mode only.
|
| 106 |
+
|
| 107 |
+
**`language_id` is a raw number**, because that is what the checkpoint declares -- English 2050, Chinese 2055, Spanish 2054, German 2053, Japanese 2058, French 2061, Korean 2064, Russian 2069, Portuguese 2071, Italian 2070. Omit it and English is assumed.
|
| 108 |
+
|
| 109 |
+
**It samples by default.** The export declares this checkpoint's own decoding -- `temperature 0.9`, `top_k 50`, `repetition_penalty 1.05`, and a second set for the code predictor -- so two runs of a sentence give two takes; `seed=` is what pins one. `temperature=0` decodes greedily and reproduces `transformers` **exactly**: verified at 672 codes over 42 frames, every one identical, and the pair's audio transcribes back to the sentence it was given.
|
| 110 |
+
|
| 111 |
+
**The repetition penalty is not optional here.** `transformers` applies it as a processor rather than a warper, so it moves a greedy argmax too, and without it this model never emits its end token -- it runs to `max_new_tokens`. Passing `repetition_penalty=1.0` turns it off and is a good way to see that.
|
| 112 |
+
|
| 113 |
+
**`max_new_tokens` counts audio frames at 12.5 per second**, not decoder steps -- one frame is sixteen transformer passes here, since a code predictor emits fifteen of the sixteen codebooks from the talker's hidden state.
|
| 114 |
+
|
| 115 |
+
**Generation is slower than the parameter count suggests**, for two reasons that are this export's rather than the model's. The code predictor runs without a KV cache (it re-reads its own sixteen-position prefix each step, which is what lets it share a file with a cached talker), and the attention is materialised as MHA rather than GQA -- `k_proj`/`v_proj` are duplicated so the key/value head count matches the query's, +69 M parameters and a doubled cache, because the grouped form does not survive conversion.
|
| 116 |
+
|
| 117 |
+
**It is a big download**: 3.9 GB, F32, like the rest of this collection. Nearly a third of it is the 151936 x 2048 text embedding table. `loom-export --quantize Q8_0` on the upstream checkpoint packs the eligible weights to about 1.1 GB if you would rather have that.
|
| 118 |
+
|
| 119 |
+
## Files
|
| 120 |
+
|
| 121 |
+
- `qwen3-tts-12hz-0.6b.gguf` -- the model, exported with loom-exporter.
|
qwen3-tts-12hz-0.6b.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:090979e436535e0ab7b0dbf0ffa67d9bb708945f269763f97747435e03747bf4
|
| 3 |
+
size 3942136000
|