F5-TTS v1 Base
F5-TTS's flow-matching voice-cloning TTS model, exported for loom.cpp. Takes a reference clip, its transcript and the text to speak; encodes characters itself.
This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.
Original model
Exported from SWivid/F5-TTS. Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
License
cc-by-nc-4.0, inherited from the base model above.
Language(s)
en, zh
Usage
Run it with loom-py -- loom-py-rt on PyPI:
pip install -U "loom-py-rt[hub]"
import loom
import librosa
model = loom.Model.from_pretrained("loom-ai-org/f5-tts-v1-base-loom")
# The voice is a REFERENCE CLIP plus WHAT IT SAYS, word for word -- this model has no voice of its own.
# A few clear seconds (at most ~12) is enough. It reads audio at 24000 Hz, so resample on the way in.
reference, _ = librosa.load("reference.wav", sr=24000)
reference_text = ("And so, my fellow Americans, ask not what your country can do for you; "
"ask what you can do for your country.")
# The model in-fills ONE spectrogram whose first frames are the reference; the door joins the transcript
# to the text and returns only the GENERATED speech, at 24000 Hz.
audio = model.text2speech.infer("hello world", reference=reference.tolist(),
reference_text=reference_text, seed=42)
audio.save("out.wav")
The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...)
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
model.driver_source prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See loom-py for the API and
loom.cpp for what the engine does between the two.
Known limitations
It clones a voice, so it needs one. Every call takes a reference clip at 24 kHz, the transcript of that clip, and the text to speak -- the model in-fills one spectrogram whose first frames are the reference, so there is no way to synthesise without a prompt.
The door is text2speech.infer(text, reference=, reference_text=), which needs loom-py-rt 1.0.0rc12 or later. It joins the transcript to the text and passes the file's own inputs: waveform (the clip, mono, 24 kHz), text_ids (the transcript's ids, a space, then the text's) and n_ref_text (how many of those ids are the transcript -- the duration estimate is a ratio of the two lengths, and nothing in the ids marks the join). On 1.0.0rc11 and earlier call model.infer(waveform=, text_ids=, n_ref_text=) with that join spelled out; the audio is identical. Text alone is refused: this model has no voice of its own. loom_cli --wav ref.wav --ref-text "..." --prompt "..." --out out.wav does the same from the shell. Optional knobs: n_steps, cfg_scale, sway_coef, speed, duration (total frames), seed.
Chinese needs pinyin conversion this file cannot do. F5-TTS's own front end runs rjieba word segmentation and pypinyin before a single id is looked up. What ships here is the character table, which reproduces that function exactly for ordinary space-separated English prose (2000/2000 generated sentences) and differs by one inserted space for text carrying multi-character punctuation runs (--, ...) or hyphen-joined digit groups (2026-09-18). Text containing CJK is refused by name; pass ids from the reference's own convert_char_to_pinyin instead.
The default duration is an estimate, not a prediction. There is no duration model: the frame count is the reference clip's own characters-per-frame rate applied to the text to speak, and duration overrides it when the result is clipped or padded.
Files
f5-tts-v1-base.gguf-- the model, exported with loom-exporter.
- Downloads last month
- 66
We're not able to determine the quantization variants.
Model tree for loom-ai-org/f5-tts-v1-base-loom
Base model
SWivid/F5-TTS