--- license: apache-2.0 language: - en - zh - ar - my - da - nl - fi - fr - de - el - he - hi - id - it - ja - km - ko - lo - ms - no - pl - pt - ru - es - sw - sv - tl - th - tr - vi base_model: - openbmb/VoxCPM2 pipeline_tag: text-to-speech tags: - loom - text-to-speech library_name: loom-py-rt --- # VoxCPM2 OpenBMB's VoxCPM2, exported for loom.cpp: a 2B diffusion-autoregressive TTS over continuous AudioVAE latents, 30 languages, 48 kHz. Encodes text itself. This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by [loom-exporter](https://github.com/loom-ai-org/loom-exporter). ## Original model Exported from [`openbmb/VoxCPM2`](https://huggingface.co/openbmb/VoxCPM2). Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format. ## License `apache-2.0`, inherited from the base model above. ## Language(s) `en`, `zh`, `ar`, `my`, `da`, `nl`, `fi`, `fr`, `de`, `el`, `he`, `hi`, `id`, `it`, `ja`, `km`, `ko`, `lo`, `ms`, `no`, `pl`, `pt`, `ru`, `es`, `sw`, `sv`, `tl`, `th`, `tr`, `vi` ## Usage Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI: ```sh pip install -U "loom-py-rt[hub]" ``` ```python import loom model = loom.Model.from_pretrained("loom-ai-org/voxcpm2-loom") # This model encodes text itself -- no phonemiser needed at all. print(model.tokenizer) # kind, vocabulary size, default language # sample_rate=48000: a rate is not something a checkpoint necessarily carries, so it is a value # you have to know from the model's documentation and pass. It is used only if the GGUF declares none; # a wrong rate does not fail, it plays the voice at the wrong speed. audio = model.text2speech.infer("hello world", sample_rate=48000) audio.save("out.wav") # That uses the voice the file itself defaults to. Whether it carries others is under "Known # limitations" (and, where it does, a section below says how to pick one). ``` ### Designing a voice VoxCPM2 has no built-in speaker: every call invents a voice to fit the text. To steer it, describe the voice in parentheses at the start of the text -- the description is not spoken: ```python designed = "(A calm older man, speaking slowly)Welcome back. The results are in." model.text2speech.infer(designed, seed=7).save("designed.wav") ``` Pass `seed` to get the same voice again. ### The layer underneath The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)` passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name. `model.driver_source` prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two. ## Known limitations **No voice cloning in this file.** The reference clones a voice from a recording through its AudioVAE's encoder, which this export does not carry. Voices are zero-shot or designed in the text (see above). **Sampled, so two calls differ.** Each 160 ms patch of audio starts from a random draw, integrated over 10 guided steps (CFG-Zero\*, guidance 2.0). Pass `seed` to reproduce a call. Verified against the reference with its draws pinned: within 1.2e-06 rms of its latents step for step. **Large, and slow on a small CPU.** 2.3B parameters: 9.3 GB at F32. On a 2-core x86 laptop a second of 48 kHz audio takes about 20 seconds. **A run that never stops is retried**, as the reference retries it: a generation that uses its whole length budget (six patches per text token) is drawn again, up to three times. ## Files - `voxcpm2.gguf` -- the model, exported with loom-exporter.