fdemelo commited on
Commit
b3c0dc2
·
verified ·
1 Parent(s): 0673639

Re-export qwen3-tts-12hz-0.6b (loom-exporter 1891005)

Browse files
Files changed (2) hide show
  1. README.md +9 -14
  2. qwen3-tts-12hz-0.6b.gguf +2 -2
README.md CHANGED
@@ -50,24 +50,19 @@ pip install -U "loom-py-rt[hub]"
50
 
51
  ```python
52
  import loom
53
- import numpy as np
54
 
55
  model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")
56
 
57
- # The voice is an X-VECTOR -- 1024 floats -- not a speaker id: this checkpoint carries no speaker
58
- # table. Making one from a reference clip inside loom is temporarily withdrawn (see "Known
59
- # limitations", which also gives the upstream recipe that wrote this file).
60
- x_vector = np.loadtxt("x_vector.txt", dtype=np.float32)
61
-
62
- # The text goes in the prompt format this model was trained on, spelled out: the file declares no
63
- # chat template, and the bare sentence makes it stop early or ramble (see "Known limitations").
64
- role = "<|im_start|>assistant\n"
65
 
66
  # What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
67
  # `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
68
  codes = model.text2codes.infer(
69
- f"{role}The quick brown fox jumps over the lazy dog.<|im_end|>\n{role}",
70
- x_vector=x_vector.tolist(),
71
  language_id=2050, # English; see "Known limitations" for the rest
72
  max_new_tokens=200, seed=1234,
73
  )
@@ -105,11 +100,11 @@ accepts for this model, and is the authority on it. See [loom-py](https://github
105
 
106
  **This model does not produce audio.** It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- [`qwen3-tts-tokenizer-12hz-loom`](https://huggingface.co/loom-ai-org/qwen3-tts-tokenizer-12hz-loom), which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.
107
 
108
- **The voice is an x-vector, and there is no speaker table.** This checkpoint's `spk_id` is empty, so cloning is the only mode: pass `x_vector=`, 1024 floats that are the whole of what the voice contributes. **Making one from a clip inside loom is temporarily withdrawn:** this file's speaker encoder (`waveform=`) currently aborts the process -- a native assertion, not a Python exception -- so do not pass `waveform=` until a release notes the fix. Until then make the x-vector once with the upstream `qwen-tts` package and save it: `spk = Qwen3TTSForConditionalGeneration.from_pretrained("Qwen/Qwen3-TTS-12Hz-0.6B-Base").extract_speaker_embedding(audio=wav, sr=24000)` on a few clear seconds of the voice resampled to 24 kHz, then `np.savetxt("x_vector.txt", spk.float().numpy().reshape(-1))`. The same file serves every later call.
109
 
110
- **Pass the text in the model's prompt format, as the example does.** The talker was trained on `<|im_start|>assistant\n{text}<|im_end|>\n<|im_start|>assistant\n`, and this file does not declare that template, so `text2codes` sends exactly what you give it. Spelled out, the prompt tokenizes to upstream's ids and a greedy run (`temperature=0, subtalker_temperature=0`) matches `transformers` on every code; the bare sentence makes the model stop after a word or two, or run to `max_new_tokens`.
111
 
112
- **The reference TEXT is used if you pass it -- that is ICL mode.** Upstream's higher-fidelity clone mode conditions on a transcript of the reference clip plus the clip's own codec tokens, replayed as if the model had just spoken it. Pass `ref_audio=` (the same clip, 24 kHz) and `ref_tokens=` (its transcript, tokenized with this model's own tokenizer and template) alongside `x_vector=`, and this file draws those codes itself: the codec's ENCODE half rides inside this GGUF, since the prompt is its only caller, and it is unaffected by the speaker encoder's abort. Pass `ref_code=` instead if you already hold them, frame-major and sixteen wide. Omit both and you get the x-vector mode above, which is what the rest of this card describes.
113
 
114
  **A reference clip is used in whole codec frames.** 1920 samples at 24 kHz, 80 ms: the driver trims to a multiple of that before encoding, so up to 79 ms of the clip's tail is not heard. Every convolution in the encode stack pads by a length-derived amount that would not otherwise trace, and on a frame boundary that amount is provably zero.
115
 
 
50
 
51
  ```python
52
  import loom
53
+ import librosa
54
 
55
  model = loom.Model.from_pretrained("loom-ai-org/qwen3-tts-12hz-0.6b-loom")
56
 
57
+ # The voice is a REFERENCE CLIP, not a speaker id -- this checkpoint carries no speaker table. A few
58
+ # clear seconds is enough. 24 kHz is what the speaker encoder expects, so resample on the way in.
59
+ reference, _ = librosa.load("reference.wav", sr=24000)
 
 
 
 
 
60
 
61
  # What comes back is codec TOKENS, not audio -- frame-major, one row per frame, 16 codebooks wide.
62
  # `max_new_tokens` counts AUDIO FRAMES, at 12.5 per second.
63
  codes = model.text2codes.infer(
64
+ "The quick brown fox jumps over the lazy dog.",
65
+ waveform=reference.tolist(),
66
  language_id=2050, # English; see "Known limitations" for the rest
67
  max_new_tokens=200, seed=1234,
68
  )
 
100
 
101
  **This model does not produce audio.** It emits sixteen streams of Qwen3-TTS codec tokens, and the codec turns those into a waveform -- [`qwen3-tts-tokenizer-12hz-loom`](https://huggingface.co/loom-ai-org/qwen3-tts-tokenizer-12hz-loom), which ships inside this same checkpoint upstream. They stay separate because one codec serves every size and variant of this talker, and because the codes are the useful intermediate.
102
 
103
+ **The voice comes from a reference clip, and there is no speaker table.** This checkpoint's `spk_id` is empty, so cloning is the only mode: pass `waveform=` and the speaker encoder extracts an x-vector from it. Pass `x_vector=` instead to reuse one you already have -- it is 1024 floats and it is the whole of what the voice contributes. Copies of this file downloaded before 2026-09-29 abort the process on `waveform=`; download it again.
104
 
105
+ **Pass the sentence as it is.** The talker was trained on `<|im_start|>assistant\n{text}<|im_end|>\n<|im_start|>assistant\n`, and this file wraps whatever it is given in that template. Ids you have already wrapped yourself -- `tokens=` that open with `<|im_start|>assistant\n` -- are passed through as they are, and so is a reference transcript in ICL mode below.
106
 
107
+ **The reference TEXT is used if you pass it -- that is ICL mode.** Upstream's higher-fidelity clone mode conditions on a transcript of the reference clip plus the clip's own codec tokens, replayed as if the model had just spoken it. Pass `ref_audio=` (the same clip, 24 kHz) and `ref_tokens=` (its transcript, tokenized with this model's own tokenizer) and this file draws those codes itself: the codec's ENCODE half rides inside this GGUF, since the prompt is its only caller. Pass `ref_code=` instead if you already hold them, frame-major and sixteen wide. Omit both and you get the x-vector mode above, which is what the rest of this card describes.
108
 
109
  **A reference clip is used in whole codec frames.** 1920 samples at 24 kHz, 80 ms: the driver trims to a multiple of that before encoding, so up to 79 ms of the clip's tail is not heard. Every convolution in the encode stack pads by a length-derived amount that would not otherwise trace, and on a frame boundary that amount is provably zero.
110
 
qwen3-tts-12hz-0.6b.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5871b9bc7803465b0a7e90ce8a1bbee2887214fd9ab3474ef75808037276b0ca
3
- size 4132302976
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ceb1bb8e543d2eacc82b9892fef2ab4746891a1e8d0c901c0ab8024b82c75540
3
+ size 4132304416