|
Download README.md from loom-ai-org/matcha-tts-ljspeech-loom: direct link, hf CLI and curl.
- Browser
- Download file 4.28 kB
-
https://huggingface.co/loom-ai-org/matcha-tts-ljspeech-loom/resolve/main/README.md
- Command line
-
hf download hf://loom-ai-org/matcha-tts-ljspeech-loom/README.md
-
curl -L -o README.md https://huggingface.co/loom-ai-org/matcha-tts-ljspeech-loom/resolve/main/README.md
4.28 kB
| license: mit | |
| language: | |
| - en | |
| pipeline_tag: text-to-speech | |
| tags: | |
| - loom | |
| - text-to-speech | |
| library_name: loom-py-rt | |
| # Matcha-TTS (LJSpeech) | |
| Matcha-TTS's LJSpeech flow-matching TTS checkpoint, exported for loom.cpp. Takes phoneme ids, not text. | |
| This is a [loom.cpp](https://github.com/loom-ai-org/loom.cpp) export: a single self-describing GGUF | |
| that carries its own graph topologies, tokenizer (if any) and driver script, produced by | |
| [loom-exporter](https://github.com/loom-ai-org/loom-exporter). | |
| ## Original model | |
| Exported from [Matcha-TTS (LJSpeech checkpoint)](https://github.com/shivammehta25/Matcha-TTS). Weights are unmodified; this repo packages the same parameters into | |
| loom.cpp's GGUF format. | |
| ## License | |
| `mit`, inherited from the base model above. | |
| ## Language(s) | |
| `en` | |
| ## Usage | |
| Run it with [loom-py](https://github.com/loom-ai-org/loom-py) -- `loom-py-rt` on PyPI: | |
| ```sh | |
| pip install -U "loom-py-rt[hub,phonemes]" | |
| ``` | |
| **NOTE:** This is a work in progress. For now, in order to avoid license conflicts and keep | |
| dependencies at a minimum, we opted for orthography2ipa as our "swiss-knife" phonemizer. For | |
| deep-orthography languages like English, to get stressing rules and context-based phonemization, the | |
| phonemizer must register a reference lexicon (e.g., ipa-dict) to get highly accurate phonemization. | |
| Full quality can be achieved by phonemizing the text yourself using your engine of choice and feeding | |
| it to the model via the argument `phonemes` as in the example below. | |
| ```python | |
| import loom | |
| model = loom.Model.from_pretrained("loom-ai-org/matcha-tts-ljspeech-loom") | |
| # matcha-tts-ljspeech is trained on phonemes. Its symbol table ships in the GGUF, so the only piece that is not in | |
| # the file is grapheme-to-phoneme -- a property of the language rather than of this checkpoint, which | |
| # is why it is the `phonemes` extra above rather than part of the model. | |
| # THE FULL-QUALITY PATH: phonemes you produced yourself, with whatever G2P you trust. The symbol table | |
| # in the GGUF is what encodes them, so anything that emits IPA works. | |
| audio = model.text2speech.infer(phonemes="həˈloʊ wˈɜːld", sample_rate=22050) | |
| audio.save("out.wav") | |
| # THE BUILT-IN PATH: text straight in, phonemized by the bundled rule-based G2P. Good enough for | |
| # shallow orthographies; for English see the note above, and give it a lexicon so it has stress and | |
| # real vowels to work with -- "time" is /tɪm/ without one. | |
| # | |
| # open-dict-data/ipa-dict (MIT) publishes ~65k-entry wordlists WITH stress for en_UK and en_US, in | |
| # almost the right shape: its IPA is wrapped in slashes and a rare entry carries two comma-separated | |
| # variants, both of which the loader rejects. Take the RAW file: a github.com/.../blob/... URL serves | |
| # an HTML PAGE, and the sed below will turn that into a .tsv that parses to zero entries -- which is | |
| # indistinguishable from no lexicon at all except for the warning `set_lexicon` raises. Two lines: | |
| # curl -LO https://raw.githubusercontent.com/open-dict-data/ipa-dict/master/data/en_UK.txt | |
| # sed 's:/::g; s/\t\([^,]*\),.*/\t\1/' en_UK.txt > en_UK.tsv | |
| loom.phonemizers.set_lexicon("en_UK.tsv") # a path, an http(s):// URL, or hf://<repo>/<path> | |
| # sample_rate=22050: this checkpoint does not carry its own rate, so it is a value you have to | |
| # know from the model's documentation and pass. It is used only if the GGUF declares none; a wrong rate | |
| # does not fail, it plays the voice at the wrong speed. | |
| audio = model.text2speech.infer("hello world", sample_rate=22050) | |
| audio.save("out.wav") | |
| ``` | |
| ### The layer underneath | |
| The call above is the high-level door: one per task, named for the modality pair it maps between, with | |
| the windowing, sampling and assembly this model needs already applied. Under it, `model.infer(...)` | |
| passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the | |
| door does not name. | |
| `model.driver_source` prints that driver, including a header comment documenting every argument it | |
| accepts for this model, and is the authority on it. See [loom-py](https://github.com/loom-ai-org/loom-py) for the API and | |
| [loom.cpp](https://github.com/loom-ai-org/loom.cpp) for what the engine does between the two. | |
| ## Files | |
| - `matcha-tts-ljspeech.gguf` -- the model, exported with loom-exporter. | |