Initial commit
Browse files- .gitattributes +1 -0
- README.md +54 -0
- styletts2-ljspeech.gguf +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
styletts2-ljspeech.gguf filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
base_model:
|
| 6 |
+
- yl4579/StyleTTS2-LJSpeech
|
| 7 |
+
pipeline_tag: text-to-speech
|
| 8 |
+
library_name: loom-py-rt
|
| 9 |
+
---
|
| 10 |
+
# StyleTTS2 (LJSpeech)
|
| 11 |
+
|
| 12 |
+
yl4579's StyleTTS2 LJSpeech checkpoint, exported for loom.cpp. Takes phoneme ids, not text.
|
| 13 |
+
|
| 14 |
+
This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF
|
| 15 |
+
that carries its own graph topologies, tokenizer (if any) and driver script, produced by
|
| 16 |
+
[loom-exporter](https://github.com/loom-ai-org/loom-exporter).
|
| 17 |
+
|
| 18 |
+
## Original model
|
| 19 |
+
|
| 20 |
+
Exported from [`yl4579/StyleTTS2-LJSpeech`](https://huggingface.co/yl4579/StyleTTS2-LJSpeech). Weights are unmodified; this repo packages the same parameters into
|
| 21 |
+
loom.cpp's GGUF format.
|
| 22 |
+
|
| 23 |
+
## License
|
| 24 |
+
|
| 25 |
+
`mit`, inherited from the base model above.
|
| 26 |
+
|
| 27 |
+
## Language(s)
|
| 28 |
+
|
| 29 |
+
`en`
|
| 30 |
+
|
| 31 |
+
the HF repo carries no `license:`/`language:` tags; MIT per the upstream GitHub repo's LICENSE (github.com/yl4579/StyleTTS2)
|
| 32 |
+
|
| 33 |
+
## Usage
|
| 34 |
+
|
| 35 |
+
Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI:
|
| 36 |
+
|
| 37 |
+
```sh
|
| 38 |
+
pip install loom-py-rt[hub]
|
| 39 |
+
```
|
| 40 |
+
|
| 41 |
+
```python
|
| 42 |
+
import loom
|
| 43 |
+
|
| 44 |
+
model = loom.Model.from_pretrained("loom-ai-org/styletts2-ljspeech-loom")
|
| 45 |
+
# styletts2-ljspeech takes phoneme ids, not text -- see model.driver_source for the exact driver inputs.
|
| 46 |
+
audio = model.infer(tokens=[16, 40, 22, 30, 12, 3], n_steps=4, seed=1234)
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
`model.driver_source` prints the exact driver script this GGUF embeds, including a header comment
|
| 50 |
+
documenting every argument `model.infer()`/`model.generate()` accepts for this model.
|
| 51 |
+
|
| 52 |
+
## Files
|
| 53 |
+
|
| 54 |
+
- `styletts2-ljspeech.gguf` -- the model, exported with loom-exporter.
|
styletts2-ljspeech.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:99ad89f69ec20dc9dab905cfad3bc8e8d3d5e81c69e78d936c6e9c606272c272
|
| 3 |
+
size 411051200
|