Voxtral 4B TTS
Mistral's Voxtral-4B-TTS-2603, exported for loom.cpp: a 3.4B Ministral backbone over frames of audio codes, a flow-matching acoustic head and a causal codec, 9 languages, 24 kHz. Encodes text itself and has a built-in voice.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from mistralai/Voxtral-4B-TTS-2603. Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
License
cc-by-nc-4.0, inherited from the base model above.
Language(s)
en, fr, es, pt, it, nl, de, ar, hi
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
model = loom.Model.from_pretrained("loom-ai-org/voxtral-4b-tts-2603-loom")
# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer) # kind, vocabulary size, default language
# sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")
# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).
Choosing a voice
The file carries one voice, casual_male, and uses it when you name none. Mistral's other 19 preset
voices ship in this repo as voice files under voices/, and a name is enough -- it is fetched from this
repo the first time:
print(model.voices) # the built-in one first, then the voice files
audio = model.text2speech.infer("Hello world.", voice="neutral_female")
audio.save("neutral_female.wav")
Pick a voice in the language you are speaking (voice="fr_female" for French, and so on): each preset
was recorded in one. A voice is the rows the
model reads in place of a reference recording, stamped with a fingerprint of these weights, so it only
fits this model and loom refuses one made for another.
| voices | language |
|---|---|
casual_male (built in), casual_female, cheerful_female, neutral_male, neutral_female |
English |
fr_male, fr_female |
French |
es_male, es_female |
Spanish |
de_male, de_female |
German |
it_male, it_female |
Italian |
pt_male, pt_female |
Portuguese |
nl_male, nl_female |
Dutch |
ar_male |
Arabic |
hi_male, hi_female |
Hindi |
Every voice is CC BY-NC 4.0, like the model: Mistral's card says the references come from the EARS, CML-TTS, IndicVoices-R and Arabic Natural Audio datasets.
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
Non-commercial (CC BY-NC 4.0). The weights and every voice carry the licence of the voice recordings they were built from, per Mistral's own card.
Preset voices only; no voice cloning. Mistral's open checkpoint ships no codec encoder, which is what turns a recording into a voice, so this file has none either (upstream says the same of its own release). The 20 presets are the built-in casual_male and 19 voice files under voices/.
Sampled, so two calls differ. Each 80 ms frame's acoustic half starts from a random draw, integrated over 7 guided steps (classifier-free guidance 1.2, the reference's default); its semantic half is the most likely code. Pass seed to reproduce a call. Verified against the reference (vLLM-Omni's own flow head and codec) with its draws pinned: the same codes for all 142 frames of an 11 s sentence, and a waveform within 2e-08 rms. Like the reference, it occasionally keeps talking after the sentence; another seed fixes it.
Large, and slow on a CPU. 4B parameters: 16 GB at F32, about 17 GB of RAM to run. On a 24-core x86 desktop a second of audio takes about 5.5 seconds; a 4-bit build is not published yet. One call speaks up to 2048 frames (164 s), and long text is not split.
Needs loom 1.0.0-rc11 or later. The text front end is Mistral's Tekken tokenizer, which older engines refuse by name rather than tokenize differently.
Files
voxtral-4b-tts-2603.gguf-- the model, exported with loom-exporter.voices/*.gguf-- Mistral's other preset voices as loom voice files, converted unchanged frommistralai/Voxtral-4B-TTS-2603'svoice_embedding/*.ptbyloom_exporter.voxtral_tts_voices. Only needed to pick a voice other thancasual_male.
- Downloads last month
- 307
We're not able to determine the quantization variants.
Model tree for loom-ai-org/voxtral-4b-tts-2603-loom
Base model
mistralai/Ministral-3-3B-Base-2512