Pocket TTS (English)

Kyutai's Pocket TTS (English, 2026-09 release), exported for loom.cpp: a 100M-parameter flow language model over Mimi codec latents. Encodes text itself and has a built-in voice.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from kyutai/pocket-tts. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

cc-by-4.0, inherited from the base model above.

Language(s)

en

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/pocket-tts-loom")

# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer)                       # kind, vocabulary size, default language

# sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")

# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).

Choosing a voice

The file carries one voice, alba, and uses it when you name none. The other 25 of Kyutai's voices ship in this repo as voice files under voices/, and a name is enough -- it is fetched from this repo the first time:

print(model.voices)                          # the built-in one first, then the voice files

audio = model.text2speech.infer("Hello there.", voice="marius")
audio.save("marius.wav")

A voice file is the model's own state after hearing the speaker, stamped with a fingerprint of these weights, so it only fits this model: loom refuses one made for a different release rather than producing speech that never stops. voice= also takes a path to a voice file of your own.

voice licence of the recording voice licence of the recording
alba (built in) CC-BY-4.0 javert CC0-1.0
anna CC-BY-4.0 jean CC-BY-NC-4.0 (non-commercial)
azelma CC-BY-4.0 juergen not stated by Kyutai
bill_boerst CC0-1.0 lola CC0-1.0
caro_davy CC0-1.0 marius CC0-1.0
charles CC-BY-4.0 mary CC-BY-4.0
cosette CC-BY-NC-4.0 (non-commercial) michael CC-BY-4.0
eponine CC-BY-4.0 paul CC-BY-4.0
estelle CC0-1.0 peter_yearsley CC0-1.0
eve CC-BY-4.0 rafael not stated by Kyutai
fantine CC-BY-4.0 stuart_bell CC0-1.0
george CC-BY-4.0 vera CC-BY-4.0
giovanni CC0-1.0
jane CC-BY-4.0

Licences are per kyutai/tts-voices' README, by the dataset each recording comes from; every file also records its own (loom.voice.license, loom.voice.origin). CC-BY voices need attribution to their source dataset or speaker.

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

Kyutai's use restrictions apply. The upstream release prohibits, among other things, voice impersonation or cloning without explicit and lawful consent, and presenting generated audio as a genuine recording of a real person. They travel with the weights.

One voice is built in (alba); the other 25 are voice files under voices/, selected with voice= (see above). A voice is the model's own attention state after hearing the speaker, so a voice file only fits these weights, and loom refuses one made for another release. Each voice's licence is its RECORDING's, not the model's -- two are non-commercial -- and the table above lists them. Cloning a voice from a recording needs the Mimi encoder, which this export does not carry; a state you saved yourself with pocket-tts export-voice converts with python -m loom_exporter.pocket_tts_voices --from.

Sampled by default, at the checkpoint's own temperature (0.3), so two calls differ; pass seed to reproduce one. Verified against the reference with its random draws pinned: within 1.8e-06 rms of the reference waveform step for step.

Long text is split, as the reference splits it: into sentence chunks of at most 50 tokens, each generated separately from the same voice and joined. A single sentence longer than that is cut at its commas.

English only. Kyutai's other languages are separate checkpoints and are not this file.

Files

  • pocket-tts.gguf -- the model, exported with loom-exporter.
  • voices/*.gguf -- Kyutai's other predefined voices as loom voice files, converted unchanged from kyutai/pocket-tts's languages/english_2026-09/embeddings/ by loom_exporter.pocket_tts_voices. Only needed to pick a voice other than alba.
Downloads last month
520
GGUF
Model size
0.1B params
Architecture
loom-pocket-tts
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/pocket-tts-loom

Quantized
(49)
this model

Collection including loom-ai-org/pocket-tts-loom