--- license: cc-by-nc-4.0 language: - en - fr - es - pt - it - nl - de - ar - hi base_model: - mistralai/Voxtral-4B-TTS-2603 pipeline_tag: text-to-speech tags: - loom - text-to-speech library_name: loom-py-rt --- # Voxtral 4B TTS Mistral's Voxtral-4B-TTS-2603, exported for loom.cpp: a 3.4B Ministral backbone over frames of audio codes, a flow-matching acoustic head and a causal codec, 9 languages, 24 kHz. Encodes text itself and has a built-in voice. This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by [loom-exporter](https://github.com/loom-ai-org/loom-exporter). ## Original model Exported from [`mistralai/Voxtral-4B-TTS-2603`](https://huggingface.co/mistralai/Voxtral-4B-TTS-2603). Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format. ## License `cc-by-nc-4.0`, inherited from the base model above. ## Language(s) `en`, `fr`, `es`, `pt`, `it`, `nl`, `de`, `ar`, `hi` ## Usage Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI: ```sh pip install -U "loom-py-rt[hub]" ``` ```python import loom model = loom.Model.from_pretrained("loom-ai-org/voxtral-4b-tts-2603-loom") # This model encodes text itself -- no phonemiser needed at all. print(model.tokenizer) # kind, vocabulary size, default language # sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value # you have to know from the model's documentation and pass. It is used only if the GGUF declares none; # a wrong rate does not fail, it plays the voice at the wrong speed. audio = model.text2speech.infer("hello world", sample_rate=24000) audio.save("out.wav") # That uses the voice the file itself defaults to. Whether it carries others is under "Known # limitations" (and, where it does, a section below says how to pick one). ``` ### Choosing a voice The file carries one voice, `casual_male`, and uses it when you name none. Mistral's other 19 preset voices ship in this repo as voice files under `voices/`, and a name is enough -- it is fetched from this repo the first time: ```python print(model.voices) # the built-in one first, then the voice files audio = model.text2speech.infer("Hello world.", voice="neutral_female") audio.save("neutral_female.wav") ``` Pick a voice in the language you are speaking (`voice="fr_female"` for French, and so on): each preset was recorded in one. A voice is the rows the model reads in place of a reference recording, stamped with a fingerprint of these weights, so it only fits this model and loom refuses one made for another. | voices | language | |---|---| | `casual_male` (built in), `casual_female`, `cheerful_female`, `neutral_male`, `neutral_female` | English | | `fr_male`, `fr_female` | French | | `es_male`, `es_female` | Spanish | | `de_male`, `de_female` | German | | `it_male`, `it_female` | Italian | | `pt_male`, `pt_female` | Portuguese | | `nl_male`, `nl_female` | Dutch | | `ar_male` | Arabic | | `hi_male`, `hi_female` | Hindi | Every voice is **CC BY-NC 4.0**, like the model: Mistral's card says the references come from the EARS, CML-TTS, IndicVoices-R and Arabic Natural Audio datasets. ### The layer underneath The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)` passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name. `model.driver_source` prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two. ## Known limitations **Non-commercial (CC BY-NC 4.0).** The weights and every voice carry the licence of the voice recordings they were built from, per Mistral's own card. **Preset voices only; no voice cloning.** Mistral's open checkpoint ships no codec encoder, which is what turns a recording into a voice, so this file has none either (upstream says the same of its own release). The 20 presets are the built-in `casual_male` and 19 voice files under `voices/`. **Sampled, so two calls differ.** Each 80 ms frame's acoustic half starts from a random draw, integrated over 7 guided steps (classifier-free guidance 1.2, the reference's default); its semantic half is the most likely code. Pass `seed` to reproduce a call. Verified against the reference (vLLM-Omni's own flow head and codec) with its draws pinned: the same codes for all 142 frames of an 11 s sentence, and a waveform within 2e-08 rms. Like the reference, it occasionally keeps talking after the sentence; another `seed` fixes it. **Large, and slow on a CPU.** 4B parameters: 16 GB at F32, about 17 GB of RAM to run. On a 24-core x86 desktop a second of audio takes about 5.5 seconds; a 4-bit build is not published yet. One call speaks up to 2048 frames (164 s), and long text is not split. **Needs loom 1.0.0-rc11 or later.** The text front end is Mistral's Tekken tokenizer, which older engines refuse by name rather than tokenize differently. ## Files - `voxtral-4b-tts-2603.gguf` -- the model, exported with loom-exporter. - `voices/*.gguf` -- Mistral's other preset voices as loom voice files, converted unchanged from `mistralai/Voxtral-4B-TTS-2603`'s `voice_embedding/*.pt` by `loom_exporter.voxtral_tts_voices`. Only needed to pick a voice other than `casual_male`.