VoxCPM2

OpenBMB's VoxCPM2, exported for loom.cpp: a 2B diffusion-autoregressive TTS over continuous AudioVAE latents, 30 languages, 48 kHz. Encodes text itself.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from openbmb/VoxCPM2. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

apache-2.0, inherited from the base model above.

Language(s)

en, zh, ar, my, da, nl, fi, fr, de, el, he, hi, id, it, ja, km, ko, lo, ms, no, pl, pt, ru, es, sw, sv, tl, th, tr, vi

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/voxcpm2-loom")

# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer)                       # kind, vocabulary size, default language

# sample_rate=48000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=48000)
audio.save("out.wav")

# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).

Designing a voice

VoxCPM2 has no built-in speaker: every call invents a voice to fit the text. To steer it, describe the voice in parentheses at the start of the text -- the description is not spoken:

designed = "(A calm older man, speaking slowly)Welcome back. The results are in."
model.text2speech.infer(designed, seed=7).save("designed.wav")

Pass seed to get the same voice again.

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

No voice cloning in this file. The reference clones a voice from a recording through its AudioVAE's encoder, which this export does not carry. Voices are zero-shot or designed in the text (see above).

Sampled, so two calls differ. Each 160 ms patch of audio starts from a random draw, integrated over 10 guided steps (CFG-Zero*, guidance 2.0). Pass seed to reproduce a call. Verified against the reference with its draws pinned: within 1.2e-06 rms of its latents step for step.

Large, and slow on a small CPU. 2.3B parameters: 9.3 GB at F32. On a 2-core x86 laptop a second of 48 kHz audio takes about 20 seconds.

A run that never stops is retried, as the reference retries it: a generation that uses its whole length budget (six patches per text token) is drawn again, up to three times.

Files

  • voxcpm2.gguf -- the model, exported with loom-exporter.
Downloads last month
78
GGUF
Model size
2B params
Architecture
loom-voxcpm2
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/voxcpm2-loom

Base model

openbmb/VoxCPM2
Quantized
(14)
this model

Collection including loom-ai-org/voxcpm2-loom