F5-TTS v1 Base

F5-TTS's flow-matching voice-cloning TTS model, exported for loom.cpp. Takes a reference clip, its transcript and the text to speak; encodes characters itself.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from SWivid/F5-TTS. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

cc-by-nc-4.0, inherited from the base model above.

Language(s)

en, zh

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom
import librosa

model = loom.Model.from_pretrained("loom-ai-org/f5-tts-v1-base-loom")

# The voice is a REFERENCE CLIP plus WHAT IT SAYS, word for word -- this model has no voice of its own.
# A few clear seconds (at most ~12) is enough. It reads audio at 24000 Hz, so resample on the way in.
reference, _ = librosa.load("reference.wav", sr=24000)
reference_text = ("And so, my fellow Americans, ask not what your country can do for you; "
                  "ask what you can do for your country.")

# The model in-fills ONE spectrogram whose first frames are the reference; the door joins the transcript
# to the text and returns only the GENERATED speech, at 24000 Hz.
audio = model.text2speech.infer("hello world", reference=reference.tolist(),
                                reference_text=reference_text, seed=42)
audio.save("out.wav")

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

It clones a voice, so it needs one. Every call takes a reference clip at 24 kHz, the transcript of that clip, and the text to speak -- the model in-fills one spectrogram whose first frames are the reference, so there is no way to synthesise without a prompt.

The door is text2speech.infer(text, reference=, reference_text=), which needs loom-py-rt 1.0.0rc12 or later. It joins the transcript to the text and passes the file's own inputs: waveform (the clip, mono, 24 kHz), text_ids (the transcript's ids, a space, then the text's) and n_ref_text (how many of those ids are the transcript -- the duration estimate is a ratio of the two lengths, and nothing in the ids marks the join). On 1.0.0rc11 and earlier call model.infer(waveform=, text_ids=, n_ref_text=) with that join spelled out; the audio is identical. Text alone is refused: this model has no voice of its own. loom_cli --wav ref.wav --ref-text "..." --prompt "..." --out out.wav does the same from the shell. Optional knobs: n_steps, cfg_scale, sway_coef, speed, duration (total frames), seed.

Chinese needs pinyin conversion this file cannot do. F5-TTS's own front end runs rjieba word segmentation and pypinyin before a single id is looked up. What ships here is the character table, which reproduces that function exactly for ordinary space-separated English prose (2000/2000 generated sentences) and differs by one inserted space for text carrying multi-character punctuation runs (--, ...) or hyphen-joined digit groups (2026-09-18). Text containing CJK is refused by name; pass ids from the reference's own convert_char_to_pinyin instead.

The default duration is an estimate, not a prediction. There is no duration model: the frame count is the reference clip's own characters-per-frame rate applied to the text to speak, and duration overrides it when the result is clipped or padded.

Files

  • f5-tts-v1-base.gguf -- the model, exported with loom-exporter.
Downloads last month
66
GGUF
Model size
0.4B params
Architecture
loom-f5-tts
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/f5-tts-v1-base-loom

Base model

SWivid/F5-TTS
Quantized
(7)
this model

Collection including loom-ai-org/f5-tts-v1-base-loom