Qwen3-TTS Swahili (Kenyan)
A LoRA adapter and a trained language row that teach Qwen3-TTS-12Hz-1.7B-Base Kenyan Swahili, which is not one of its ten languages. It is built for the dream-tts engine (Rust, Metal), which overlays the adapter on the base checkpoint at load time.
| file | what it is |
|---|---|
swahili-adapter.safetensors |
LoRA pairs (rank 16, q/k/v/o/gate/up/down on the talker), codec row 2074 trained as the Swahili language tag, and the orthography scheme (meta::scheme). Exported at gain 0.6. 82 MB. |
swahili.json |
The orthography scheme on its own: how Swahili text is cut into the pieces the talker was trained on. |
voices/mc-swahili-qwen3tts/ |
A dream-tts voice asset: 16 s of Kenyan Swahili news read by the author. |
voices/sw-male-qwen3tts/, voices/sw-female-qwen3tts/ |
Generic Kenyan Swahili voices, male and female. They are synthetic: designed with Qwen3-TTS VoiceDesign, not cloned from anyone. |
samples/ |
Renders with the voice above: three paragraphs, and a chapter through the full narration pipeline. |
Use
dream-tts speak --engine qwen3tts --voice voices/mc-swahili-qwen3tts \
--set adapter=swahili-adapter.safetensors --set language=swahili \
--text-file chapter.md --text-language swahili --out chapter.wav
For a voice whose reference clip is in another language, add --set clone=xvector. It
conditions on the speaker embedding only, because continuing an English clip carries its English
accent into the Swahili. With the English-reference male voice, CER falls from 6.0% to 2.4% over
36 renders.
--set adapter_gain=G applies the LoRA at another strength. The engine reads the orthography from
the adapter and tokenizes both the target text and the voice's transcript with it. It also adds a
short lead-in before each segment (a sentence-initial ng' was otherwise dropped 15 times in 20),
continues each segment from the previous one, and levels segments to one loudness.
Training
- Data. 11.5 h of publicly available TV broadcast speech (500+ speakers). Clips were cut at
pauses and transcribed by Gemini, then kept only when a CTC forced alignment confirmed the text.
Also 3.8 h of
WAXAL
swa_tts(CC-BY-SA-4.0), limited to clips as fluent as the broadcast speech. - Method. In-context pairs: a reference clip and a target clip from the same speaker, as in voice cloning. Pairs were matched by recording channel and speaker-embedding similarity, and every clip was levelled to -20 dBFS. Gradients flow through the whole prompt, which trains the language row too. The code predictor is frozen. 2000 steps on an Apple M4 (MLX).
- Orthography. Open syllables. Prenasal onsets (mb, nd, ng, nj, nz, mv) are split off, j
is split from its vowel, and ng' becomes a single
ŋtoken.
Evaluation
The evaluation used 12 sentences from a held-out broadcast and three voices: the author's (Swahili reference) and two with English reference clips. CER is from Gemini transcripts.
| author's voice | cosy-default (English reference) |
male (English reference) |
|
|---|---|---|---|
| CER | 1.6% | 5.1% | 6.0% (36 renders) |
| sentence-to-sentence speaker similarity within a passage (base model without adapter) | 0.782 (0.747) | 0.699 (0.707) | 0.672 (0.684) |
The earlier adapter paused mid-phrase 12.4 times per 100 words; this one pauses 5.0, against 6.8 for the original speaker of the same sentences. A chapter through the dream-tts narration pipeline aligned 142 of 142 words.
The generic voices
Each generic voice was made in three steps. First it was designed from a short description (a Kenyan man or woman, warm and clear, like a radio news presenter). It then spoke a Swahili passage through the adapter in x-vector mode, and the most fluent take became its own reference clip. Because they clone from Swahili, nothing English is inherited. Over three paragraphs, sentence-to-sentence speaker similarity is 0.782 (male) and 0.818 (female), with CER 2.6% and 3.1%. The English-reference voices score 0.672 and 0.699 when continuing their clips.
Limitations
- Voices with English reference clips keep an English accent when cloned from their clip, since
cloning continues it accent included. Use
clone=xvectorfor them, or a Swahili reference clip. - A sentence-initial ng' is right about 13 times in 20, even with the lead-in.
- It is tested only with dream-tts. Swahili numbers, dates and currency are read as words by
dream-tts's text normaliser (
--text-language swahili), not by the model.
The voice asset
voices/mc-swahili-qwen3tts is the author's own voice, shared by the author for use with this
adapter. Do not use it to impersonate the author, or to publish speech attributed to them without
their consent.
Licence
CC-BY-SA-4.0, following WAXAL's share-alike terms; the base model is Apache-2.0. Attribution: WAXAL (Google), Qwen3-TTS (Alibaba Qwen team).
Model tree for drmhse/qwen3-tts-swahili
Base model
Qwen/Qwen3-TTS-12Hz-1.7B-Base