Text-to-Speech
NeMo
speech-synthesis
multilingual
swahili
continual-learning
codec-language-model
malcomzw commited on
Commit
00e1da3
·
verified ·
1 Parent(s): 3c2df87

Upload folder using huggingface_hub

Browse files
Files changed (5) hide show
  1. .gitattributes +1 -0
  2. LICENSE +27 -0
  3. NOTICE +18 -0
  4. README.md +124 -0
  5. magpie_tts_13lang_357m.nemo +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ magpie_tts_13lang_357m.nemo filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ NVIDIA OPEN MODEL LICENSE
2
+ =========================
3
+
4
+ This model, "Magpie-TTS 13-Language (357M)", is a derivative of NVIDIA
5
+ Magpie-TTS-Multilingual (nvidia/magpie_tts_multilingual_357m) and is distributed
6
+ under the NVIDIA Open Model License Agreement.
7
+
8
+ The full and authoritative text of the NVIDIA Open Model License is available at:
9
+
10
+ https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
11
+
12
+ Required attribution (see NOTICE):
13
+
14
+ "Licensed by NVIDIA Corporation under the NVIDIA Open Model License."
15
+
16
+ Summary of key permissions and conditions (the linked text governs):
17
+ - Redistribution and commercial use are permitted.
18
+ - Derivative models are permitted (this model is one such derivative).
19
+ - You must retain this attribution/notice and comply with the license terms,
20
+ including its use restrictions and trustworthy-AI provisions.
21
+
22
+ The Swahili language extension in this derivative was produced by New Emerging
23
+ Technologies (a subsidiary of Infinia Technologies) and is not provided,
24
+ supported, or endorsed by NVIDIA Corporation.
25
+
26
+ Training data are licensed CC-BY-4.0 (Google FLEURS; Bateesa Kiswahili TTS) and
27
+ are credited in the model card and NOTICE.
NOTICE ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Magpie-TTS 13-Language (357M)
2
+ Copyright (c) 2026 New Emerging Technologies (a subsidiary of Infinia Technologies).
3
+
4
+ This model is a derivative of NVIDIA Magpie-TTS-Multilingual
5
+ (nvidia/magpie_tts_multilingual_357m), extended with Swahili support.
6
+
7
+ Licensed by NVIDIA Corporation under the NVIDIA Open Model License.
8
+ See the LICENSE file for the full terms.
9
+
10
+ The Swahili language extension was added by the community and is not provided,
11
+ supported, or endorsed by NVIDIA Corporation.
12
+
13
+ Training data:
14
+ - FLEURS (Conneau et al., 2022), CC-BY-4.0
15
+ - Bateesa Kiswahili TTS dataset, CC-BY-4.0
16
+
17
+ Audio generated by this model carries a watermark and a synthetic-speech
18
+ disclosure inherited from the base model.
README.md ADDED
@@ -0,0 +1,124 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: nvidia-open-model-license
4
+ license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
5
+ library_name: nemo
6
+ pipeline_tag: text-to-speech
7
+ base_model: nvidia/magpie_tts_multilingual_357m
8
+ language:
9
+ - ar
10
+ - de
11
+ - en
12
+ - es
13
+ - fr
14
+ - hi
15
+ - it
16
+ - ja
17
+ - ko
18
+ - pt
19
+ - sw
20
+ - vi
21
+ - zh
22
+ datasets:
23
+ - google/fleurs
24
+ - Bateesa/kiswahili-tts-dataset
25
+ tags:
26
+ - text-to-speech
27
+ - speech-synthesis
28
+ - multilingual
29
+ - swahili
30
+ - continual-learning
31
+ - codec-language-model
32
+ - nemo
33
+ ---
34
+
35
+ # Magpie-TTS 13-Language (357M) — with Swahili
36
+
37
+ A single **13-language** text-to-speech checkpoint: the twelve languages of NVIDIA's
38
+ [`magpie_tts_multilingual_357m`](https://huggingface.co/nvidia/magpie_tts_multilingual_357m)
39
+ (Arabic, German, English, Spanish, French, Hindi, Italian, Japanese, Korean,
40
+ Portuguese, Vietnamese, Chinese) **plus Swahili (Kiswahili)** — added by a
41
+ community team **without regressing any of the original twelve**.
42
+
43
+ Swahili was grafted onto the frozen base model with a dedicated byte-level
44
+ tokenizer, a **warm-started** input-embedding surgery, and multilingual rehearsal.
45
+ On held-out Swahili the model reaches **8.8% character error rate** (median 0.0%,
46
+ MMS-Swahili ASR), and the twelve base languages show **no measurable regression**
47
+ versus the untouched base. Full method: see the accompanying paper *"Cross-Lingual
48
+ Grafting: Extending Frozen Codec Speech Language Models to Unseen Low-Resource
49
+ Languages"* (New Emerging Technologies, a subsidiary of Infinia Technologies).
50
+
51
+ > ⚠️ **Community model.** Swahili support was added by the community and is **not**
52
+ > provided or endorsed by NVIDIA. The twelve base languages follow the base model.
53
+
54
+ ## Highlights
55
+
56
+ | | |
57
+ |---|---|
58
+ | Base | Magpie-TTS-Multilingual 357M (Koel-TTS family) |
59
+ | Audio codec | Low Frame-rate Speech Codec (22.05 kHz, 1.89 kbps, 21.5 fps) |
60
+ | Languages | 13 (12 base + Swahili) |
61
+ | Swahili quality | 8.8% mean / 0.0% median CER (MMS-sw) |
62
+ | Base regression | ~0 (mean recognizer CER identical to base) |
63
+ | Voices | 5 baked speakers (2 female, 3 male), selected by index |
64
+ | Trained on | 1× 128 GB unified-memory GPU |
65
+
66
+ ## Usage (NVIDIA NeMo)
67
+
68
+ ```python
69
+ from nemo.collections.tts.models import MagpieTTSModel
70
+ import soundfile as sf, numpy as np, torch, random
71
+
72
+ m = MagpieTTSModel.from_pretrained("infiniatechnologies/magpie-tts-13lang-357m").eval().cuda()
73
+
74
+ # Deterministic inference: pin the seed so identical inputs give identical audio
75
+ def seed(s=1234):
76
+ random.seed(s); np.random.seed(s); torch.manual_seed(s); torch.cuda.manual_seed_all(s)
77
+
78
+ seed(1234)
79
+ audio, alen = m.do_tts(
80
+ "Habari ya asubuhi. Karibu kwenye jaribio la sauti ya Kiswahili.",
81
+ language="sw", # native Swahili code; also en, de, es, fr, it, vi,
82
+ # zh, hi, ja, pt-BR, ko, ar-MSA, ar-SA, ar-AE
83
+ speaker_index=0, # 0 Aria(F) 1 Jason(M) 2 John(M) 3 Leo(M) 4 Sofia(F)
84
+ apply_TN=False, use_cfg=True,
85
+ )
86
+ w = np.squeeze(audio.detach().float().cpu().numpy())
87
+ sf.write("out.wav", (w[0] if w.ndim > 1 else w).astype(np.float32), 22050)
88
+ ```
89
+
90
+ **Determinism.** The released tokenizers use `phoneme_probability=1.0` (deterministic
91
+ phonemization) and the sampler defaults (temperature 0.7, top-k 80, CFG 2.5) match
92
+ the base model. With a fixed RNG seed, identical (text, language, voice, CFG, seed)
93
+ inputs produce byte-identical audio; change the seed for a different rendering.
94
+ For the occasional degenerate silent generation, resample at a new seed
95
+ (retry-on-silence).
96
+
97
+ ## Language codes
98
+ `en, de, es, fr, it, vi, zh, hi, ja, pt-BR, ko, ar-MSA, ar-SA, ar-AE, sw`.
99
+
100
+ ## Training data
101
+ - **Swahili**: [Bateesa/kiswahili-tts-dataset](https://huggingface.co/datasets/Bateesa/kiswahili-tts-dataset) (CC-BY) + [FLEURS](https://huggingface.co/datasets/google/fleurs) `sw_ke` (CC-BY), ~27.8 h, resampled to 22.05 kHz.
102
+ - **Rehearsal (12 base languages)**: FLEURS, ~500 utterances/language, each routed to its native tokenizer.
103
+
104
+ Please cite FLEURS (Conneau et al., *arXiv:2205.12446*) and the Koel-TTS
105
+ (*arXiv:2502.05236*) and Low Frame-rate Speech Codec (*arXiv:2409.12117*) papers.
106
+
107
+ ## Intended use & limitations
108
+ Research and product speech synthesis for the 13 supported languages. **Out of
109
+ scope:** impersonation of real individuals, deceptive or harmful synthetic media.
110
+ The model synthesizes a **fixed set of 5 baked voices**; it does **not** support
111
+ zero-shot voice cloning (the base model's context encoder is not included).
112
+ Swahili data skews toward read/literary speech and one dominant speaker; FLEURS
113
+ adds multi-speaker breadth at upsampled 16 kHz. 357M parameters bound absolute
114
+ quality.
115
+
116
+ ## Responsible use
117
+ Generated audio carries the base model's watermark and synthetic-speech
118
+ disclosure. Do not use to deceive. Disclose that audio is AI-generated.
119
+
120
+ ## License & attribution
121
+ Released under the **NVIDIA Open Model License** (see `LICENSE` and `NOTICE`).
122
+ Derived from `nvidia/magpie_tts_multilingual_357m`. *"Licensed by NVIDIA
123
+ Corporation under the NVIDIA Open Model License."* Training data: FLEURS (CC-BY),
124
+ Bateesa Kiswahili TTS (CC-BY).
magpie_tts_13lang_357m.nemo ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:be4d96080ee69b6d270767576bd4ce883eb6282fcc4791b28064ece0985a7f95
3
+ size 1460951040