File size: 4,433 Bytes
83dbfbf 3cc6f31 83dbfbf 32ef4fa 83dbfbf b8130b9 0c31030 83dbfbf 0bc2c03 b8130b9 21d5d7d b8130b9 3cc6f31 b8130b9 21d5d7d 0bc2c03 83dbfbf 0bc2c03 83dbfbf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 | ---
license: mit
language:
- en
base_model:
- yl4579/StyleTTS2-LJSpeech
pipeline_tag: text-to-speech
tags:
- loom
- text-to-speech
library_name: loom-py-rt
---
# StyleTTS2 (LJSpeech)
yl4579's StyleTTS2 LJSpeech checkpoint, exported for loom.cpp. Takes phoneme ids, not text.
This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
that carries its own graph topologies, tokenizer (if any) and driver script, produced by
[loom-exporter](https://github.com/loom-ai-org/loom-exporter).
## Original model
Exported from [`yl4579/StyleTTS2-LJSpeech`](https://huggingface.co/yl4579/StyleTTS2-LJSpeech). Weights are unmodified; this repo packages the same parameters into
loom.cpp's GGUF format.
## License
`mit`, inherited from the base model above.
## Language(s)
`en`
the HF repo carries no `license:`/`language:` tags; MIT per the upstream GitHub repo's LICENSE (github.com/yl4579/StyleTTS2)
## Usage
Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:
```sh
pip install -U "loom-py-rt[hub,phonemes]"
```
**NOTE:** This is a work in progress. For now, in order to avoid license conflicts and keep
dependencies at a minimum, we opted for orthography2ipa as our "swiss-knife" phonemizer. For
deep-orthography languages like English, to get stressing rules and context-based phonemization, the
phonemizer must register a reference lexicon (e.g., ipa-dict) to get highly accurate phonemization.
Full quality can be achieved by phonemizing the text yourself using your engine of choice and feeding
it to the model via the argument `phonemes` as in the example below.
```python
import loom
model = loom.Model.from_pretrained("loom-ai-org/styletts2-ljspeech-loom")
# styletts2-ljspeech is trained on phonemes. Its symbol table ships in the GGUF, so the only piece that is not in
# the file is grapheme-to-phoneme -- a property of the language rather than of this checkpoint, which
# is why it is the `phonemes` extra above rather than part of the model.
# THE FULL-QUALITY PATH: phonemes you produced yourself, with whatever G2P you trust. The symbol table
# in the GGUF is what encodes them, so anything that emits IPA works.
audio = model.text2speech.infer(phonemes="həˈloʊ wˈɜːld", sample_rate=24000)
audio.save("out.wav")
# THE BUILT-IN PATH: text straight in, phonemized by the bundled rule-based G2P. Good enough for
# shallow orthographies; for English see the note above, and give it a lexicon so it has stress and
# real vowels to work with -- "time" is /tɪm/ without one.
#
# open-dict-data/ipa-dict (MIT) publishes ~65k-entry wordlists WITH stress for en_UK and en_US, in
# almost the right shape: its IPA is wrapped in slashes and a rare entry carries two comma-separated
# variants, both of which the loader rejects. Take the RAW file: a github.com/.../blob/... URL serves
# an HTML PAGE, and the sed below will turn that into a .tsv that parses to zero entries -- which is
# indistinguishable from no lexicon at all except for the warning `set_lexicon` raises. Two lines:
# curl -LO https://raw.githubusercontent.com/open-dict-data/ipa-dict/master/data/en_UK.txt
# sed 's:/::g; s/\t\([^,]*\),.*/\t\1/' en_UK.txt > en_UK.tsv
loom.phonemizers.set_lexicon("en_UK.tsv") # a path, an http(s):// URL, or hf://<repo>/<path>
# sample_rate=24000: this checkpoint does not carry its own rate, so it is a value you have to
# know from the model's documentation and pass. It is used only if the GGUF declares none; a wrong rate
# does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")
```
### The layer underneath
The call above is the high-level door: one per task, named for the modality pair it maps between, with
the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)`
passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the
door does not name.
`model.driver_source` prints that driver, including a header comment documenting every argument it
accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and
[loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two.
## Files
- `styletts2-ljspeech.gguf` -- the model, exported with loom-exporter.
|