Qwen3-TTS Swahili (Kenyan)

A LoRA adapter and a trained language row that teach Qwen3-TTS-12Hz-1.7B-Base Kenyan Swahili, which is not one of its ten languages. It is built for the dream-tts engine (Rust, Metal), which overlays the adapter on the base checkpoint at load time.

file what it is
swahili-adapter.safetensors LoRA pairs (rank 16, q/k/v/o/gate/up/down on the talker), codec row 2074 trained as the Swahili language tag, and the orthography scheme (meta::scheme). Exported at gain 0.6. 82 MB.
swahili.json The orthography scheme on its own: how Swahili text is cut into the pieces the talker was trained on.
voices/mc-swahili-qwen3tts/ A dream-tts voice asset: 16 s of Kenyan Swahili news read by the author.
voices/sw-male-qwen3tts/, voices/sw-female-qwen3tts/ Generic Kenyan Swahili voices, male and female. They are synthetic: designed with Qwen3-TTS VoiceDesign, not cloned from anyone.
samples/ Renders with the voice above: three paragraphs, and a chapter through the full narration pipeline.

Use

dream-tts speak --engine qwen3tts --voice voices/mc-swahili-qwen3tts \
  --set adapter=swahili-adapter.safetensors --set language=swahili \
  --text-file chapter.md --text-language swahili --out chapter.wav

For a voice whose reference clip is in another language, add --set clone=xvector. It conditions on the speaker embedding only, because continuing an English clip carries its English accent into the Swahili. With the English-reference male voice, CER falls from 6.0% to 2.4% over 36 renders.

--set adapter_gain=G applies the LoRA at another strength. The engine reads the orthography from the adapter and tokenizes both the target text and the voice's transcript with it. It also adds a short lead-in before each segment (a sentence-initial ng' was otherwise dropped 15 times in 20), continues each segment from the previous one, and levels segments to one loudness.

Training

  • Data. 11.5 h of publicly available TV broadcast speech (500+ speakers). Clips were cut at pauses and transcribed by Gemini, then kept only when a CTC forced alignment confirmed the text. Also 3.8 h of WAXAL swa_tts (CC-BY-SA-4.0), limited to clips as fluent as the broadcast speech.
  • Method. In-context pairs: a reference clip and a target clip from the same speaker, as in voice cloning. Pairs were matched by recording channel and speaker-embedding similarity, and every clip was levelled to -20 dBFS. Gradients flow through the whole prompt, which trains the language row too. The code predictor is frozen. 2000 steps on an Apple M4 (MLX).
  • Orthography. Open syllables. Prenasal onsets (mb, nd, ng, nj, nz, mv) are split off, j is split from its vowel, and ng' becomes a single ŋ token.

Evaluation

The evaluation used 12 sentences from a held-out broadcast and three voices: the author's (Swahili reference) and two with English reference clips. CER is from Gemini transcripts.

author's voice cosy-default (English reference) male (English reference)
CER 1.6% 5.1% 6.0% (36 renders)
sentence-to-sentence speaker similarity within a passage (base model without adapter) 0.782 (0.747) 0.699 (0.707) 0.672 (0.684)

The earlier adapter paused mid-phrase 12.4 times per 100 words; this one pauses 5.0, against 6.8 for the original speaker of the same sentences. A chapter through the dream-tts narration pipeline aligned 142 of 142 words.

The generic voices

Each generic voice was made in three steps. First it was designed from a short description (a Kenyan man or woman, warm and clear, like a radio news presenter). It then spoke a Swahili passage through the adapter in x-vector mode, and the most fluent take became its own reference clip. Because they clone from Swahili, nothing English is inherited. Over three paragraphs, sentence-to-sentence speaker similarity is 0.782 (male) and 0.818 (female), with CER 2.6% and 3.1%. The English-reference voices score 0.672 and 0.699 when continuing their clips.

Limitations

  • Voices with English reference clips keep an English accent when cloned from their clip, since cloning continues it accent included. Use clone=xvector for them, or a Swahili reference clip.
  • A sentence-initial ng' is right about 13 times in 20, even with the lead-in.
  • It is tested only with dream-tts. Swahili numbers, dates and currency are read as words by dream-tts's text normaliser (--text-language swahili), not by the model.

The voice asset

voices/mc-swahili-qwen3tts is the author's own voice, shared by the author for use with this adapter. Do not use it to impersonate the author, or to publish speech attributed to them without their consent.

Licence

CC-BY-SA-4.0, following WAXAL's share-alike terms; the base model is Apache-2.0. Attribution: WAXAL (Google), Qwen3-TTS (Alibaba Qwen team).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drmhse/qwen3-tts-swahili

Adapter
(8)
this model