File size: 4,060 Bytes
e164bae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
---
license: apache-2.0
language:
- en
- zh
- ar
- my
- da
- nl
- fi
- fr
- de
- el
- he
- hi
- id
- it
- ja
- km
- ko
- lo
- ms
- no
- pl
- pt
- ru
- es
- sw
- sv
- tl
- th
- tr
- vi
base_model:
- openbmb/VoxCPM2
pipeline_tag: text-to-speech
tags:
- loom
- text-to-speech
library_name: loom-py-rt
---
# VoxCPM2

OpenBMB's VoxCPM2, exported for loom.cpp: a 2B diffusion-autoregressive TTS over continuous AudioVAE latents, 30 languages, 48 kHz. Encodes text itself.

This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
that carries its own graph topologies, tokenizer (if any) and driver script, produced by
[loom-exporter](https://github.com/loom-ai-org/loom-exporter).

## Original model

Exported from [`openbmb/VoxCPM2`](https://huggingface.co/openbmb/VoxCPM2). Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.

## License

`apache-2.0`, inherited from the base model above.

## Language(s)

`en`, `zh`, `ar`, `my`, `da`, `nl`, `fi`, `fr`, `de`, `el`, `he`, `hi`, `id`, `it`, `ja`, `km`, `ko`, `lo`, `ms`, `no`, `pl`, `pt`, `ru`, `es`, `sw`, `sv`, `tl`, `th`, `tr`, `vi`

## Usage

Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:

```sh
pip install -U "loom-py-rt[hub]"
```

```python
import loom

model = loom.Model.from_pretrained("loom-ai-org/voxcpm2-loom")

# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer)                       # kind, vocabulary size, default language

# sample_rate=48000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=48000)
audio.save("out.wav")

# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).
```

### Designing a voice

VoxCPM2 has no built-in speaker: every call invents a voice to fit the text. To steer it, describe the
voice in parentheses at the start of the text -- the description is not spoken:

```python
designed = "(A calm older man, speaking slowly)Welcome back. The results are in."
model.text2speech.infer(designed, seed=7).save("designed.wav")
```

Pass `seed` to get the same voice again.

### The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.

`model.driver_source` prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
[loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.

## Known limitations

**No voice cloning in this file.** The reference clones a voice from a recording through its AudioVAE's encoder, which this export does not carry. Voices are zero-shot or designed in the text (see above).

**Sampled, so two calls differ.** Each 160 ms patch of audio starts from a random draw, integrated over 10 guided steps (CFG-Zero\*, guidance 2.0). Pass `seed` to reproduce a call. Verified against the reference with its draws pinned: within 1.2e-06 rms of its latents step for step.

**Large, and slow on a small CPU.** 2.3B parameters: 9.3 GB at F32. On a 2-core x86 laptop a second of 48 kHz audio takes about 20 seconds.

**A run that never stops is retried**, as the reference retries it: a generation that uses its whole length budget (six patches per text token) is drawn again, up to three times.

## Files

- `voxcpm2.gguf` -- the model, exported with loom-exporter.