Pocket TTS (English)
Kyutai's Pocket TTS (English, 2026-09 release), exported for loom.cpp: a 100M-parameter flow language model over Mimi codec latents. Encodes text itself and has a built-in voice.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from kyutai/pocket-tts. Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
License
cc-by-4.0, inherited from the base model above.
Language(s)
en
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
model = loom.Model.from_pretrained("loom-ai-org/pocket-tts-loom")
# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer) # kind, vocabulary size, default language
# sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")
# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).
Choosing a voice
The file carries one voice, alba, and uses it when you name none. The other 25 of Kyutai's voices
ship in this repo as voice files under voices/, and a name is enough -- it is fetched from this repo
the first time:
print(model.voices) # the built-in one first, then the voice files
audio = model.text2speech.infer("Hello there.", voice="marius")
audio.save("marius.wav")
A voice file is the model's own state after hearing the speaker, stamped with a fingerprint of these
weights, so it only fits this model: loom refuses one made for a different release rather than
producing speech that never stops. voice= also takes a path to a voice file of your own.
| voice | licence of the recording | voice | licence of the recording |
|---|---|---|---|
alba (built in) |
CC-BY-4.0 | javert |
CC0-1.0 |
anna |
CC-BY-4.0 | jean |
CC-BY-NC-4.0 (non-commercial) |
azelma |
CC-BY-4.0 | juergen |
not stated by Kyutai |
bill_boerst |
CC0-1.0 | lola |
CC0-1.0 |
caro_davy |
CC0-1.0 | marius |
CC0-1.0 |
charles |
CC-BY-4.0 | mary |
CC-BY-4.0 |
cosette |
CC-BY-NC-4.0 (non-commercial) | michael |
CC-BY-4.0 |
eponine |
CC-BY-4.0 | paul |
CC-BY-4.0 |
estelle |
CC0-1.0 | peter_yearsley |
CC0-1.0 |
eve |
CC-BY-4.0 | rafael |
not stated by Kyutai |
fantine |
CC-BY-4.0 | stuart_bell |
CC0-1.0 |
george |
CC-BY-4.0 | vera |
CC-BY-4.0 |
giovanni |
CC0-1.0 | ||
jane |
CC-BY-4.0 |
Licences are per kyutai/tts-voices' README, by the
dataset each recording comes from; every file also records its own (loom.voice.license,
loom.voice.origin). CC-BY voices need attribution to their source dataset or speaker.
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
Kyutai's use restrictions apply. The upstream release prohibits, among other things, voice impersonation or cloning without explicit and lawful consent, and presenting generated audio as a genuine recording of a real person. They travel with the weights.
One voice is built in (alba); the other 25 are voice files under voices/, selected with voice= (see above). A voice is the model's own attention state after hearing the speaker, so a voice file only fits these weights, and loom refuses one made for another release. Each voice's licence is its RECORDING's, not the model's -- two are non-commercial -- and the table above lists them. Cloning a voice from a recording needs the Mimi encoder, which this export does not carry; a state you saved yourself with pocket-tts export-voice converts with python -m loom_exporter.pocket_tts_voices --from.
Sampled by default, at the checkpoint's own temperature (0.3), so two calls differ; pass seed to reproduce one. Verified against the reference with its random draws pinned: within 1.8e-06 rms of the reference waveform step for step.
Long text is split, as the reference splits it: into sentence chunks of at most 50 tokens, each generated separately from the same voice and joined. A single sentence longer than that is cut at its commas.
English only. Kyutai's other languages are separate checkpoints and are not this file.
Files
pocket-tts.gguf-- the model, exported with loom-exporter.voices/*.gguf-- Kyutai's other predefined voices as loom voice files, converted unchanged fromkyutai/pocket-tts'slanguages/english_2026-09/embeddings/byloom_exporter.pocket_tts_voices. Only needed to pick a voice other thanalba.
- Downloads last month
- 520
We're not able to determine the quantization variants.
Model tree for loom-ai-org/pocket-tts-loom
Base model
kyutai/pocket-tts