File size: 4,433 Bytes
83dbfbf
 
 
 
 
 
 
3cc6f31
 
 
83dbfbf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32ef4fa
83dbfbf
 
b8130b9
 
 
 
 
 
0c31030
83dbfbf
 
 
 
0bc2c03
 
b8130b9
 
 
 
 
 
 
 
 
 
 
21d5d7d
b8130b9
 
3cc6f31
 
 
 
b8130b9
 
 
21d5d7d
 
 
 
0bc2c03
83dbfbf
 
0bc2c03
 
 
 
 
 
 
 
 
 
83dbfbf
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
---
license: mit
language:
- en
base_model:
- yl4579/StyleTTS2-LJSpeech
pipeline_tag: text-to-speech
tags:
- loom
- text-to-speech
library_name: loom-py-rt
---
# StyleTTS2 (LJSpeech)

yl4579's StyleTTS2 LJSpeech checkpoint, exported for loom.cpp. Takes phoneme ids, not text.

This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
that carries its own graph topologies, tokenizer (if any) and driver script, produced by
[loom-exporter](https://github.com/loom-ai-org/loom-exporter).

## Original model

Exported from [`yl4579/StyleTTS2-LJSpeech`](https://huggingface.co/yl4579/StyleTTS2-LJSpeech). Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.

## License

`mit`, inherited from the base model above.

## Language(s)

`en`

the HF repo carries no `license:`/`language:` tags; MIT per the upstream GitHub repo's LICENSE (github.com/yl4579/StyleTTS2)

## Usage

Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:

```sh
pip install -U "loom-py-rt[hub,phonemes]"
```

**NOTE:** This is a work in progress. For now, in order to avoid license conflicts and keep
dependencies at a minimum, we opted for orthography2ipa as our "swiss-knife" phonemizer. For
deep-orthography languages like English, to get stressing rules and context-based phonemization, the
phonemizer must register a reference lexicon (e.g., ipa-dict) to get highly accurate phonemization.
Full quality can be achieved by phonemizing the text yourself using your engine of choice and feeding
it to the model via the argument `phonemes` as in the example below.

```python
import loom

model = loom.Model.from_pretrained("loom-ai-org/styletts2-ljspeech-loom")

# styletts2-ljspeech is trained on phonemes. Its symbol table ships in the GGUF, so the only piece that is not in
# the file is grapheme-to-phoneme -- a property of the language rather than of this checkpoint, which
# is why it is the `phonemes` extra above rather than part of the model.

# THE FULL-QUALITY PATH: phonemes you produced yourself, with whatever G2P you trust. The symbol table
# in the GGUF is what encodes them, so anything that emits IPA works.
audio = model.text2speech.infer(phonemes="həˈloʊ wˈɜːld", sample_rate=24000)
audio.save("out.wav")

# THE BUILT-IN PATH: text straight in, phonemized by the bundled rule-based G2P. Good enough for
# shallow orthographies; for English see the note above, and give it a lexicon so it has stress and
# real vowels to work with -- "time" is /tɪm/ without one.
#
# open-dict-data/ipa-dict (MIT) publishes ~65k-entry wordlists WITH stress for en_UK and en_US, in
# almost the right shape: its IPA is wrapped in slashes and a rare entry carries two comma-separated
# variants, both of which the loader rejects. Take the RAW file: a github.com/.../blob/... URL serves
# an HTML PAGE, and the sed below will turn that into a .tsv that parses to zero entries -- which is
# indistinguishable from no lexicon at all except for the warning `set_lexicon` raises. Two lines:
#     curl -LO https://raw.githubusercontent.com/open-dict-data/ipa-dict/master/data/en_UK.txt
#     sed 's:/::g; s/\t\([^,]*\),.*/\t\1/' en_UK.txt > en_UK.tsv
loom.phonemizers.set_lexicon("en_UK.tsv")    # a path, an http(s):// URL, or hf://<repo>/<path>

# sample_rate=24000: this checkpoint does not carry its own rate, so it is a value you have to
# know from the model's documentation and pass. It is used only if the GGUF declares none; a wrong rate
# does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")
```

### The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.

`model.driver_source` prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
[loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.

## Files

- `styletts2-ljspeech.gguf` -- the model, exported with loom-exporter.