|
Download README.md from drmhse/qwen3-tts-swahili: direct link, hf CLI and curl.
- Browser
- Download file 5.27 kB
-
https://huggingface.co/drmhse/qwen3-tts-swahili/resolve/main/README.md
- Command line
-
hf download hf://drmhse/qwen3-tts-swahili/README.md
-
curl -L -o README.md https://huggingface.co/drmhse/qwen3-tts-swahili/resolve/main/README.md
5.27 kB
| license: cc-by-sa-4.0 | |
| base_model: Qwen/Qwen3-TTS-12Hz-1.7B-Base | |
| language: | |
| - sw | |
| tags: | |
| - text-to-speech | |
| - lora | |
| - swahili | |
| - kenyan-swahili | |
| - voice-cloning | |
| # Qwen3-TTS Swahili (Kenyan) | |
| A LoRA adapter and a trained language row that teach | |
| [Qwen3-TTS-12Hz-1.7B-Base](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base) Kenyan | |
| Swahili, which is not one of its ten languages. It is built for the | |
| [dream-tts](https://github.com/drmhse/dream-tts) engine (Rust, Metal), which overlays the adapter | |
| on the base checkpoint at load time. | |
| | file | what it is | | |
| |---|---| | |
| | `swahili-adapter.safetensors` | LoRA pairs (rank 16, q/k/v/o/gate/up/down on the talker), codec row 2074 trained as the Swahili language tag, and the orthography scheme (`meta::scheme`). Exported at gain 0.6. 82 MB. | | |
| | `swahili.json` | The orthography scheme on its own: how Swahili text is cut into the pieces the talker was trained on. | | |
| | `voices/mc-swahili-qwen3tts/` | A dream-tts voice asset: 16 s of Kenyan Swahili news read by the author. | | |
| | `voices/sw-male-qwen3tts/`, `voices/sw-female-qwen3tts/` | Generic Kenyan Swahili voices, male and female. They are synthetic: designed with Qwen3-TTS VoiceDesign, not cloned from anyone. | | |
| | `samples/` | Renders with the voice above: three paragraphs, and a chapter through the full narration pipeline. | | |
| ## Use | |
| ```sh | |
| dream-tts speak --engine qwen3tts --voice voices/mc-swahili-qwen3tts \ | |
| --set adapter=swahili-adapter.safetensors --set language=swahili \ | |
| --text-file chapter.md --text-language swahili --out chapter.wav | |
| ``` | |
| For a voice whose reference clip is in another language, add `--set clone=xvector`. It | |
| conditions on the speaker embedding only, because continuing an English clip carries its English | |
| accent into the Swahili. With the English-reference `male` voice, CER falls from 6.0% to 2.4% over | |
| 36 renders. | |
| `--set adapter_gain=G` applies the LoRA at another strength. The engine reads the orthography from | |
| the adapter and tokenizes both the target text and the voice's transcript with it. It also adds a | |
| short lead-in before each segment (a sentence-initial *ng'* was otherwise dropped 15 times in 20), | |
| continues each segment from the previous one, and levels segments to one loudness. | |
| ## Training | |
| - **Data.** 11.5 h of publicly available TV broadcast speech (500+ speakers). Clips were cut at | |
| pauses and transcribed by Gemini, then kept only when a CTC forced alignment confirmed the text. | |
| Also 3.8 h of | |
| [WAXAL](https://huggingface.co/datasets/google/WaxalNLP) `swa_tts` (CC-BY-SA-4.0), limited to | |
| clips as fluent as the broadcast speech. | |
| - **Method.** In-context pairs: a reference clip and a target clip from the same speaker, as in | |
| voice cloning. Pairs were matched by recording channel and speaker-embedding similarity, and | |
| every clip was levelled to -20 dBFS. Gradients flow through the whole prompt, which trains the | |
| language row too. The code predictor is frozen. 2000 steps on an Apple M4 (MLX). | |
| - **Orthography.** Open syllables. Prenasal onsets (*mb, nd, ng, nj, nz, mv*) are split off, *j* | |
| is split from its vowel, and *ng'* becomes a single `ŋ` token. | |
| ## Evaluation | |
| The evaluation used 12 sentences from a held-out broadcast and three voices: the author's | |
| (Swahili reference) and two with English reference clips. CER is from Gemini transcripts. | |
| | | author's voice | `cosy-default` (English reference) | `male` (English reference) | | |
| |---|---|---|---| | |
| | CER | 1.6% | 5.1% | 6.0% (36 renders) | | |
| | sentence-to-sentence speaker similarity within a passage (base model without adapter) | 0.782 (0.747) | 0.699 (0.707) | 0.672 (0.684) | | |
| The earlier adapter paused mid-phrase 12.4 times per 100 words; this one pauses 5.0, against 6.8 | |
| for the original speaker of the same sentences. A chapter through the dream-tts narration | |
| pipeline aligned 142 of 142 words. | |
| ## The generic voices | |
| Each generic voice was made in three steps. First it was designed from a short description (a | |
| Kenyan man or woman, warm and clear, like a radio news presenter). It then spoke a Swahili passage | |
| through the adapter in x-vector mode, and the most fluent take became its own reference clip. | |
| Because they clone from Swahili, nothing English is inherited. Over three paragraphs, | |
| sentence-to-sentence speaker similarity is 0.782 (male) and 0.818 (female), with CER 2.6% and | |
| 3.1%. The English-reference voices score 0.672 and 0.699 when continuing their clips. | |
| ## Limitations | |
| - **Voices with English reference clips keep an English accent** when cloned from their clip, since | |
| cloning continues it accent included. Use `clone=xvector` for them, or a Swahili reference clip. | |
| - A sentence-initial *ng'* is right about 13 times in 20, even with the lead-in. | |
| - It is tested only with dream-tts. Swahili numbers, dates and currency are read as words by | |
| dream-tts's text normaliser (`--text-language swahili`), not by the model. | |
| ## The voice asset | |
| `voices/mc-swahili-qwen3tts` is the author's own voice, shared by the author for use with this | |
| adapter. Do not use it to impersonate the author, or to publish speech attributed to them without | |
| their consent. | |
| ## Licence | |
| CC-BY-SA-4.0, following WAXAL's share-alike terms; the base model is Apache-2.0. Attribution: | |
| WAXAL (Google), Qwen3-TTS (Alibaba Qwen team). | |