|
Download README.md from loom-ai-org/voxcpm2-loom: direct link, hf CLI and curl.
- Browser
- Download file 4.06 kB
-
https://huggingface.co/loom-ai-org/voxcpm2-loom/resolve/main/README.md
- Command line
-
hf download hf://loom-ai-org/voxcpm2-loom/README.md
-
curl -L -o README.md https://huggingface.co/loom-ai-org/voxcpm2-loom/resolve/main/README.md
4.06 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| - zh | |
| - ar | |
| - my | |
| - da | |
| - nl | |
| - fi | |
| - fr | |
| - de | |
| - el | |
| - he | |
| - hi | |
| - id | |
| - it | |
| - ja | |
| - km | |
| - ko | |
| - lo | |
| - ms | |
| - no | |
| - pl | |
| - pt | |
| - ru | |
| - es | |
| - sw | |
| - sv | |
| - tl | |
| - th | |
| - tr | |
| - vi | |
| base_model: | |
| - openbmb/VoxCPM2 | |
| pipeline_tag: text-to-speech | |
| tags: | |
| - loom | |
| - text-to-speech | |
| library_name: loom-py-rt | |
| # VoxCPM2 | |
| OpenBMB's VoxCPM2, exported for loom.cpp: a 2B diffusion-autoregressive TTS over continuous AudioVAE latents, 30 languages, 48 kHz. Encodes text itself. | |
| This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF | |
| that carries its own graph topologies, tokenizer (if any) and driver script, produced by | |
| [loom-exporter](https://github.com/loom-ai-org/loom-exporter). | |
| ## Original model | |
| Exported from [`openbmb/VoxCPM2`](https://huggingface.co/openbmb/VoxCPM2). Weights are unmodified; this repo packages the same parameters into | |
| loom.cpp's GGUF format. | |
| ## License | |
| `apache-2.0`, inherited from the base model above. | |
| ## Language(s) | |
| `en`, `zh`, `ar`, `my`, `da`, `nl`, `fi`, `fr`, `de`, `el`, `he`, `hi`, `id`, `it`, `ja`, `km`, `ko`, `lo`, `ms`, `no`, `pl`, `pt`, `ru`, `es`, `sw`, `sv`, `tl`, `th`, `tr`, `vi` | |
| ## Usage | |
| Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI: | |
| ```sh | |
| pip install -U "loom-py-rt[hub]" | |
| ``` | |
| ```python | |
| import loom | |
| model = loom.Model.from_pretrained("loom-ai-org/voxcpm2-loom") | |
| # This model encodes text itself -- no phonemiser needed at all. | |
| print(model.tokenizer) # kind, vocabulary size, default language | |
| # sample_rate=48000: a rate is not something a checkpoint necessarily carries, so it is a value | |
| # you have to know from the model's documentation and pass. It is used only if the GGUF declares none; | |
| # a wrong rate does not fail, it plays the voice at the wrong speed. | |
| audio = model.text2speech.infer("hello world", sample_rate=48000) | |
| audio.save("out.wav") | |
| # That uses the voice the file itself defaults to. Whether it carries others is under "Known | |
| # limitations" (and, where it does, a section below says how to pick one). | |
| ``` | |
| ### Designing a voice | |
| VoxCPM2 has no built-in speaker: every call invents a voice to fit the text. To steer it, describe the | |
| voice in parentheses at the start of the text -- the description is not spoken: | |
| ```python | |
| designed = "(A calm older man, speaking slowly)Welcome back. The results are in." | |
| model.text2speech.infer(designed, seed=7).save("designed.wav") | |
| ``` | |
| Pass `seed` to get the same voice again. | |
| ### The layer underneath | |
| The call above is the high-level door: one per task, named for the modality pair it maps between, with | |
| the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)` | |
| passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the | |
| door does not name. | |
| `model.driver_source` prints that driver, including a header comment documenting every argument it | |
| accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and | |
| [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two. | |
| ## Known limitations | |
| **No voice cloning in this file.** The reference clones a voice from a recording through its AudioVAE's encoder, which this export does not carry. Voices are zero-shot or designed in the text (see above). | |
| **Sampled, so two calls differ.** Each 160 ms patch of audio starts from a random draw, integrated over 10 guided steps (CFG-Zero\*, guidance 2.0). Pass `seed` to reproduce a call. Verified against the reference with its draws pinned: within 1.2e-06 rms of its latents step for step. | |
| **Large, and slow on a small CPU.** 2.3B parameters: 9.3 GB at F32. On a 2-core x86 laptop a second of 48 kHz audio takes about 20 seconds. | |
| **A run that never stops is retried**, as the reference retries it: a generation that uses its whole length budget (six patches per text token) is drawn again, up to three times. | |
| ## Files | |
| - `voxcpm2.gguf` -- the model, exported with loom-exporter. | |