--- license: cc-by-4.0 language: - en base_model: - kyutai/pocket-tts pipeline_tag: text-to-speech tags: - loom - text-to-speech library_name: loom-py-rt --- # Pocket TTS (English) Kyutai's Pocket TTS (English, 2026-09 release), exported for loom.cpp: a 100M-parameter flow language model over Mimi codec latents. Encodes text itself and has a built-in voice. This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by [loom-exporter](https://github.com/loom-ai-org/loom-exporter). ## Original model Exported from [`kyutai/pocket-tts`](https://huggingface.co/kyutai/pocket-tts). Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format. ## License `cc-by-4.0`, inherited from the base model above. ## Language(s) `en` ## Usage Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI: ```sh pip install -U "loom-py-rt[hub]" ``` ```python import loom model = loom.Model.from_pretrained("loom-ai-org/pocket-tts-loom") # This model encodes text itself -- no phonemiser needed at all. print(model.tokenizer) # kind, vocabulary size, default language # sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value # you have to know from the model's documentation and pass. It is used only if the GGUF declares none; # a wrong rate does not fail, it plays the voice at the wrong speed. audio = model.text2speech.infer("hello world", sample_rate=24000) audio.save("out.wav") # That uses the voice the file itself defaults to. Whether it carries others is under "Known # limitations" (and, where it does, a section below says how to pick one). ``` ### Choosing a voice The file carries one voice, `alba`, and uses it when you name none. The other 25 of Kyutai's voices ship in this repo as voice files under `voices/`, and a name is enough -- it is fetched from this repo the first time: ```python print(model.voices) # the built-in one first, then the voice files audio = model.text2speech.infer("Hello there.", voice="marius") audio.save("marius.wav") ``` A voice file is the model's own state after hearing the speaker, stamped with a fingerprint of these weights, so it only fits this model: loom refuses one made for a different release rather than producing speech that never stops. `voice=` also takes a path to a voice file of your own. | voice | licence of the recording | voice | licence of the recording | |---|---|---|---| | `alba` (built in) | CC-BY-4.0 | `javert` | CC0-1.0 | | `anna` | CC-BY-4.0 | `jean` | **CC-BY-NC-4.0 (non-commercial)** | | `azelma` | CC-BY-4.0 | `juergen` | not stated by Kyutai | | `bill_boerst` | CC0-1.0 | `lola` | CC0-1.0 | | `caro_davy` | CC0-1.0 | `marius` | CC0-1.0 | | `charles` | CC-BY-4.0 | `mary` | CC-BY-4.0 | | `cosette` | **CC-BY-NC-4.0 (non-commercial)** | `michael` | CC-BY-4.0 | | `eponine` | CC-BY-4.0 | `paul` | CC-BY-4.0 | | `estelle` | CC0-1.0 | `peter_yearsley` | CC0-1.0 | | `eve` | CC-BY-4.0 | `rafael` | not stated by Kyutai | | `fantine` | CC-BY-4.0 | `stuart_bell` | CC0-1.0 | | `george` | CC-BY-4.0 | `vera` | CC-BY-4.0 | | `giovanni` | CC0-1.0 | | | | `jane` | CC-BY-4.0 | | | Licences are per [`kyutai/tts-voices`](https://huggingface.co/kyutai/tts-voices)' README, by the dataset each recording comes from; every file also records its own (`loom.voice.license`, `loom.voice.origin`). CC-BY voices need attribution to their source dataset or speaker. ### The layer underneath The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)` passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name. `model.driver_source` prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two. ## Known limitations **Kyutai's use restrictions apply.** The upstream release prohibits, among other things, voice impersonation or cloning without explicit and lawful consent, and presenting generated audio as a genuine recording of a real person. They travel with the weights. **One voice is built in (`alba`); the other 25 are voice files** under `voices/`, selected with `voice=` (see above). A voice is the model's own attention state after hearing the speaker, so a voice file only fits these weights, and loom refuses one made for another release. Each voice's licence is its RECORDING's, not the model's -- two are non-commercial -- and the table above lists them. Cloning a voice from a recording needs the Mimi encoder, which this export does not carry; a state you saved yourself with `pocket-tts export-voice` converts with `python -m loom_exporter.pocket_tts_voices --from`. **Sampled by default**, at the checkpoint's own temperature (0.3), so two calls differ; pass `seed` to reproduce one. Verified against the reference with its random draws pinned: within 1.8e-06 rms of the reference waveform step for step. **Long text is split, as the reference splits it**: into sentence chunks of at most 50 tokens, each generated separately from the same voice and joined. A single sentence longer than that is cut at its commas. **English only.** Kyutai's other languages are separate checkpoints and are not this file. ## Files - `pocket-tts.gguf` -- the model, exported with loom-exporter. - `voices/*.gguf` -- Kyutai's other predefined voices as loom voice files, converted unchanged from `kyutai/pocket-tts`'s `languages/english_2026-09/embeddings/` by `loom_exporter.pocket_tts_voices`. Only needed to pick a voice other than `alba`.