fdemelo commited on
Commit
356ac12
·
verified ·
1 Parent(s): 8287196

Re-export qwen3-tts-12hz-0.6b (loom-exporter f1c2bbc)

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +121 -0
  3. qwen3-tts-12hz-0.6b.gguf +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ qwen3-tts-12hz-0.6b.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - zh
5
+ - en
6
+ - ja
7
+ - ko
8
+ - de
9
+ - fr
10
+ - ru
11
+ - pt
12
+ - es
13
+ - it
14
+ base_model:
15
+ - Qwen/Qwen3-TTS-12Hz-0.6B-Base
16
+ pipeline_tag: text-to-speech
17
+ tags:
18
+ - loom
19
+ - text-to-codes
20
+ library_name: loom-py-rt
21
+ ---
22
+ # Qwen3-TTS 12Hz 0.6B Base
23
+
24
+ Qwen's voice-cloning TTS talker, exported for loom.cpp. Family 10: text and a reference voice in, neural-codec tokens out -- pair it with `qwen3-tts-tokenizer-12hz-loom` for audio.
25
+
26
+ This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
27
+ that carries its own graph topologies, tokenizer (if any) and driver script, produced by
28
+ [loom-exporter](https://github.com/loom-ai-org/loom-exporter).
29
+
30
+ ## Original model
31
+
32
+ Exported from [`Qwen/Qwen3-TTS-12Hz-0.6B-Base`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base). Weights are unmodified; this repo packages the same parameters into
33
+ loom.cpp's GGUF format.
34
+
35
+ ## License
36
+
37
+ `apache-2.0`, inherited from the base model above.
38
+
39
+ ## Language(s)
40
+
41
+ `zh`, `en`, `ja`, `ko`, `de`, `fr`, `ru`, `pt`, `es`, `it`
42
+
43
+ ## Usage
44
+
45
+ Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:
46
+
47
+ ```sh
48
+ pip install -U "loom-py-rt[hub]"
49
+ ```
50
+
51
+ ```python
52
+ import loom
53
+ import librosa
54
+
55
+ model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")
56
+
57
+ # The voice is a REFERENCE CLIP, not a speaker id -- this checkpoint carries no speaker table. A few
58
+ # clear seconds is enough. 24 kHz is what the speaker encoder expects, so resample on the way in.
59
+ reference, _ = librosa.load("reference.wav", sr=24000)
60
+
61
+ # What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
62
+ # `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
63
+ codes = model.text2codes.infer(
64
+ "The quick brown fox jumps over the lazy dog.",
65
+ waveform=reference.tolist(),
66
+ language_id=2050, # English; see "Known limitations" for the rest
67
+ max_new_tokens=200, seed=1234,
68
+ )
69
+ print(len(codes), "frames x", len(codes[0]), "codebooks")
70
+
71
+ # The other half of the pair, in a repo of its own: the codec serves every size and variant of this
72
+ # talker, and the codes are worth having on their own -- cache them, edit them, decode them elsewhere.
73
+ codec = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-tokenizer-12hz-loom")
74
+ audio = codec.codes2speech.infer(codes)
75
+ print(len(audio), "samples at", audio.sample_rate, "Hz =", round(audio.duration, 2), "s")
76
+ audio.save("out.wav")
77
+
78
+ # Nothing goes between those two calls. Both files declare the width of a frame, so a pair that does
79
+ # not fit says so instead of producing audio of the wrong duration:
80
+ print(model.hparam("codec.n_codebooks"), "==", codec.hparam("codec.n_codebooks"))
81
+
82
+ # This model SAMPLES by default, at its own generation config's settings. `seed=` is what makes a
83
+ # take reproducible; pass temperature=0 for greedy, which reproduces `transformers` exactly.
84
+ print(model.hparam("sampling.temperature", "f32"),
85
+ model.hparam("sampling.repetition_penalty", "f32"))
86
+ ```
87
+
88
+ ### The layer underneath
89
+
90
+ The call above is the high-level door: one per task, named for the modality pair it maps between, with
91
+ the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
92
+ passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
93
+ door does not name.
94
+
95
+ `model.driver_source` prints that driver, including a header comment documenting every argument it
96
+ accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
97
+ [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.
98
+
99
+ ## Known limitations
100
+
101
+ **This model does not produce audio.** It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- [`qwen3-tts-tokenizer-12hz-loom`](https://huggingface.co/loom-ai-org/qwen3-tts-tokenizer-12hz-loom), which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.
102
+
103
+ **The voice comes from a reference clip, and there is no speaker table.** This checkpoint's `spk_id` is empty, so cloning is the only mode: pass `waveform=` and the speaker encoder extracts an x-vector from it. Pass `x_vector=` instead to reuse one you already have -- it is 1024 floats and it is the whole of what the voice contributes.
104
+
105
+ **The reference TEXT is not used.** Upstream offers a second, higher-fidelity clone mode that conditions on a transcript of the reference clip plus its codec tokens; it needs the codec's ENCODE half, which this collection does not export, so this file implements the x-vector mode only.
106
+
107
+ **`language_id` is a raw number**, because that is what the checkpoint declares -- English 2050, Chinese 2055, Spanish 2054, German 2053, Japanese 2058, French 2061, Korean 2064, Russian 2069, Portuguese 2071, Italian 2070. Omit it and English is assumed.
108
+
109
+ **It samples by default.** The export declares this checkpoint's own decoding -- `temperature 0.9`, `top_k 50`, `repetition_penalty 1.05`, and a second set for the code predictor -- so two runs of a sentence give two takes; `seed=` is what pins one. `temperature=0` decodes greedily and reproduces `transformers` **exactly**: verified at 672 codes over 42 frames, every one identical, and the pair's audio transcribes back to the sentence it was given.
110
+
111
+ **The repetition penalty is not optional here.** `transformers` applies it as a processor rather than a warper, so it moves a greedy argmax too, and without it this model never emits its end token -- it runs to `max_new_tokens`. Passing `repetition_penalty=1.0` turns it off and is a good way to see that.
112
+
113
+ **`max_new_tokens` counts audio frames at 12.5 per second**, not decoder steps -- one frame is sixteen transformer passes here, since a code predictor emits fifteen of the sixteen codebooks from the talker's hidden state.
114
+
115
+ **Generation is slower than the parameter count suggests**, for two reasons that are this export's rather than the model's. The code predictor runs without a KV cache (it re-reads its own sixteen-position prefix each step, which is what lets it share a file with a cached talker), and the attention is materialised as MHA rather than GQA -- `k_proj`/`v_proj` are duplicated so the key/value head count matches the query's, +69 M parameters and a doubled cache, because the grouped form does not survive conversion.
116
+
117
+ **It is a big download**: 3.9 GB, F32, like the rest of this collection. Nearly a third of it is the 151936 x 2048 text embedding table. `loom-export --quantize Q8_0` on the upstream checkpoint packs the eligible weights to about 1.1 GB if you would rather have that.
118
+
119
+ ## Files
120
+
121
+ - `qwen3-tts-12hz-0.6b.gguf` -- the model, exported with loom-exporter.
qwen3-tts-12hz-0.6b.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:090979e436535e0ab7b0dbf0ffa67d9bb708945f269763f97747435e03747bf4
3
+ size 3942136000