Voxtral 4B TTS

Mistral's Voxtral-4B-TTS-2603, exported for loom.cpp: a 3.4B Ministral backbone over frames of audio codes, a flow-matching acoustic head and a causal codec, 9 languages, 24 kHz. Encodes text itself and has a built-in voice.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from mistralai/Voxtral-4B-TTS-2603. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

cc-by-nc-4.0, inherited from the base model above.

Language(s)

en, fr, es, pt, it, nl, de, ar, hi

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/voxtral-4b-tts-2603-loom")

# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer)                       # kind, vocabulary size, default language

# sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")

# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).

Choosing a voice

The file carries one voice, casual_male, and uses it when you name none. Mistral's other 19 preset voices ship in this repo as voice files under voices/, and a name is enough -- it is fetched from this repo the first time:

print(model.voices)                          # the built-in one first, then the voice files

audio = model.text2speech.infer("Hello world.", voice="neutral_female")
audio.save("neutral_female.wav")

Pick a voice in the language you are speaking (voice="fr_female" for French, and so on): each preset was recorded in one. A voice is the rows the model reads in place of a reference recording, stamped with a fingerprint of these weights, so it only fits this model and loom refuses one made for another.

voices language
casual_male (built in), casual_female, cheerful_female, neutral_male, neutral_female English
fr_male, fr_female French
es_male, es_female Spanish
de_male, de_female German
it_male, it_female Italian
pt_male, pt_female Portuguese
nl_male, nl_female Dutch
ar_male Arabic
hi_male, hi_female Hindi

Every voice is CC BY-NC 4.0, like the model: Mistral's card says the references come from the EARS, CML-TTS, IndicVoices-R and Arabic Natural Audio datasets.

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

Non-commercial (CC BY-NC 4.0). The weights and every voice carry the licence of the voice recordings they were built from, per Mistral's own card.

Preset voices only; no voice cloning. Mistral's open checkpoint ships no codec encoder, which is what turns a recording into a voice, so this file has none either (upstream says the same of its own release). The 20 presets are the built-in casual_male and 19 voice files under voices/.

Sampled, so two calls differ. Each 80 ms frame's acoustic half starts from a random draw, integrated over 7 guided steps (classifier-free guidance 1.2, the reference's default); its semantic half is the most likely code. Pass seed to reproduce a call. Verified against the reference (vLLM-Omni's own flow head and codec) with its draws pinned: the same codes for all 142 frames of an 11 s sentence, and a waveform within 2e-08 rms. Like the reference, it occasionally keeps talking after the sentence; another seed fixes it.

Large, and slow on a CPU. 4B parameters: 16 GB at F32, about 17 GB of RAM to run. On a 24-core x86 desktop a second of audio takes about 5.5 seconds; a 4-bit build is not published yet. One call speaks up to 2048 frames (164 s), and long text is not split.

Needs loom 1.0.0-rc11 or later. The text front end is Mistral's Tekken tokenizer, which older engines refuse by name rather than tokenize differently.

Files

  • voxtral-4b-tts-2603.gguf -- the model, exported with loom-exporter.
  • voices/*.gguf -- Mistral's other preset voices as loom voice files, converted unchanged from mistralai/Voxtral-4B-TTS-2603's voice_embedding/*.pt by loom_exporter.voxtral_tts_voices. Only needed to pick a voice other than casual_male.
Downloads last month
307
GGUF
Model size
4B params
Architecture
loom-voxtral_tts
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/voxtral-4b-tts-2603-loom

Quantized
(11)
this model

Collection including loom-ai-org/voxtral-4b-tts-2603-loom