Re-export qwen3-tts-12hz-0.6b (loom-exporter 1891005)
Browse files- README.md +9 -14
- qwen3-tts-12hz-0.6b.gguf +2 -2
README.md
CHANGED
|
@@ -50,24 +50,19 @@ pip install -U "loom-py-rt[hub]"
|
|
| 50 |
|
| 51 |
```python
|
| 52 |
import loom
|
| 53 |
-
import
|
| 54 |
|
| 55 |
model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")
|
| 56 |
|
| 57 |
-
# The voice is
|
| 58 |
-
#
|
| 59 |
-
|
| 60 |
-
x_vector = np.loadtxt("x_vector.txt", dtype=np.float32)
|
| 61 |
-
|
| 62 |
-
# The text goes in the prompt format this model was trained on, spelled out: the file declares no
|
| 63 |
-
# chat template, and the bare sentence makes it stop early or ramble (see "Known limitations").
|
| 64 |
-
role = "<|im_start|>assistant\n"
|
| 65 |
|
| 66 |
# What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
|
| 67 |
# `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
|
| 68 |
codes = model.text2codes.infer(
|
| 69 |
-
|
| 70 |
-
|
| 71 |
language_id=2050, # English; see "Known limitations" for the rest
|
| 72 |
max_new_tokens=200, seed=1234,
|
| 73 |
)
|
|
@@ -105,11 +100,11 @@ accepts for this model, and is the authority on it. See [loom-py](https://github
|
|
| 105 |
|
| 106 |
**This model does not produce audio.** It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- [`qwen3-tts-tokenizer-12hz-loom`](https://huggingface.co/loom-ai-org/qwen3-tts-tokenizer-12hz-loom), which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.
|
| 107 |
|
| 108 |
-
**The voice
|
| 109 |
|
| 110 |
-
**Pass the
|
| 111 |
|
| 112 |
-
**The reference TEXT is used if you pass it -- that is ICL mode.** Upstream's higher-fidelity clone mode conditions on a transcript of the reference clip plus the clip's own codec tokens, replayed as if the model had just spoken it. Pass `ref_audio=` (the same clip, 24 kHz) and `ref_tokens=` (its transcript, tokenized with this model's own tokenizer
|
| 113 |
|
| 114 |
**A reference clip is used in whole codec frames.** 1920 samples at 24 kHz, 80 ms: the driver trims to a multiple of that before encoding, so up to 79 ms of the clip's tail is not heard. Every convolution in the encode stack pads by a length-derived amount that would not otherwise trace, and on a frame boundary that amount is provably zero.
|
| 115 |
|
|
|
|
| 50 |
|
| 51 |
```python
|
| 52 |
import loom
|
| 53 |
+
import librosa
|
| 54 |
|
| 55 |
model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")
|
| 56 |
|
| 57 |
+
# The voice is a REFERENCE CLIP, not a speaker id -- this checkpoint carries no speaker table. A few
|
| 58 |
+
# clear seconds is enough. 24 kHz is what the speaker encoder expects, so resample on the way in.
|
| 59 |
+
reference, _ = librosa.load("reference.wav", sr=24000)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
|
| 61 |
# What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
|
| 62 |
# `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
|
| 63 |
codes = model.text2codes.infer(
|
| 64 |
+
"The quick brown fox jumps over the lazy dog.",
|
| 65 |
+
waveform=reference.tolist(),
|
| 66 |
language_id=2050, # English; see "Known limitations" for the rest
|
| 67 |
max_new_tokens=200, seed=1234,
|
| 68 |
)
|
|
|
|
| 100 |
|
| 101 |
**This model does not produce audio.** It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- [`qwen3-tts-tokenizer-12hz-loom`](https://huggingface.co/loom-ai-org/qwen3-tts-tokenizer-12hz-loom), which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.
|
| 102 |
|
| 103 |
+
**The voice comes from a reference clip, and there is no speaker table.** This checkpoint's `spk_id` is empty, so cloning is the only mode: pass `waveform=` and the speaker encoder extracts an x-vector from it. Pass `x_vector=` instead to reuse one you already have -- it is 1024 floats and it is the whole of what the voice contributes. Copies of this file downloaded before 2026-09-29 abort the process on `waveform=`; download it again.
|
| 104 |
|
| 105 |
+
**Pass the sentence as it is.** The talker was trained on `<|im_start|>assistant\n{text}<|im_end|>\n<|im_start|>assistant\n`, and this file wraps whatever it is given in that template. Ids you have already wrapped yourself -- `tokens=` that open with `<|im_start|>assistant\n` -- are passed through as they are, and so is a reference transcript in ICL mode below.
|
| 106 |
|
| 107 |
+
**The reference TEXT is used if you pass it -- that is ICL mode.** Upstream's higher-fidelity clone mode conditions on a transcript of the reference clip plus the clip's own codec tokens, replayed as if the model had just spoken it. Pass `ref_audio=` (the same clip, 24 kHz) and `ref_tokens=` (its transcript, tokenized with this model's own tokenizer) and this file draws those codes itself: the codec's ENCODE half rides inside this GGUF, since the prompt is its only caller. Pass `ref_code=` instead if you already hold them, frame-major and sixteen wide. Omit both and you get the x-vector mode above, which is what the rest of this card describes.
|
| 108 |
|
| 109 |
**A reference clip is used in whole codec frames.** 1920 samples at 24 kHz, 80 ms: the driver trims to a multiple of that before encoding, so up to 79 ms of the clip's tail is not heard. Every convolution in the encode stack pads by a length-derived amount that would not otherwise trace, and on a frame boundary that amount is provably zero.
|
| 110 |
|
qwen3-tts-12hz-0.6b.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ceb1bb8e543d2eacc82b9892fef2ab4746891a1e8d0c901c0ab8024b82c75540
|
| 3 |
+
size 4132304416
|