FraPiz commited on
Commit
209ef39
·
verified ·
1 Parent(s): 58e55b6

Publish Moldovan Romanian XTTS-v2 model

Browse files
Files changed (9) hide show
  1. LICENSE.txt +84 -0
  2. README.md +126 -0
  3. best_model.pth +3 -0
  4. config.json +160 -0
  5. dvae.pth +3 -0
  6. generate_tts.py +297 -0
  7. mel_stats.pth +3 -0
  8. requirements.txt +5 -0
  9. vocab.json +0 -0
LICENSE.txt ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Coqui Public Model License 1.0.0
2
+ https://coqui.ai/cpml.txt
3
+
4
+
5
+ This license allows only non-commercial use of a machine learning model and its outputs.
6
+
7
+
8
+ ## Acceptance
9
+
10
+
11
+ In order to get any license under these terms, you must agree to them as both strict obligations and conditions to all your licenses.
12
+
13
+
14
+ ## Licenses
15
+
16
+
17
+ The licensor grants you a copyright license to do everything you might do with the model that would otherwise infringe the licensor's copyright in it, for any non-commercial purpose. The licensor grants you a patent license that covers patent claims the licensor can license, or becomes able to license, that you would infringe by using the model in the form provided by
18
+ the licensor, for any non-commercial purpose.
19
+
20
+
21
+ ## Non-commercial Purpose
22
+
23
+
24
+ Non-commercial purposes include any of the following uses of the model or its output, but only so far as you do not receive any direct or indirect payment arising from the use of the model or its output.
25
+
26
+
27
+ ### Personal use for research, experiment, and testing for the benefit of public knowledge, personal study, private entertainment, hobby projects, amateur pursuits, or religious
28
+ observance.
29
+
30
+
31
+ ### Use by commercial or for-profit entities for testing, evaluation, or non-commercial research and development. Use of the model to train other models for commercial use is not a non-commercial purpose.
32
+
33
+
34
+ ### Use by any charitable organization for charitable purposes, or for testing or evaluation. Use for revenue-generating activity, including projects directly funded by government grants, is not a non-commercial purpose.
35
+
36
+
37
+ ## Notices
38
+
39
+
40
+ You must ensure that anyone who gets a copy of any part of the model, or any modification of the model, or their output, from you also gets a copy of these terms or the URL for them above.
41
+
42
+
43
+ ## No Other Rights
44
+
45
+
46
+ These terms do not allow you to sublicense or transfer any of your licenses to anyone else, or prevent the licensor from granting licenses to anyone else. These terms do not imply
47
+ any other licenses.
48
+
49
+
50
+ ## Patent Defense
51
+
52
+
53
+ If you make any written claim that the model infringes or contributes to infringement of any patent, your licenses for the model granted under these terms ends immediately. If your company makes such a claim, your patent license ends immediately for work on behalf of your company.
54
+
55
+
56
+ ## Violations
57
+
58
+
59
+ The first time you are notified in writing that you have violated any of these terms, or done anything with the model or its output that is not covered by your licenses, your licenses can nonetheless continue if you come into full compliance with these terms, and take practical steps to correct past violations, within 30 days of receiving notice. Otherwise, all your licenses
60
+ end immediately.
61
+
62
+
63
+ ## No Liability
64
+
65
+
66
+ ***As far as the law allows, the model and its output come as is, without any warranty or condition, and the licensor will not be liable to you for any damages arising out of these terms or the use or nature of the model or its output, under any kind of legal claim. If this provision is not enforceable in your jurisdiction, your licenses are void.***
67
+
68
+
69
+ ## Definitions
70
+
71
+
72
+ The **licensor** is the individual or entity offering these terms, and the **model** is the model the licensor makes available under these terms, including any documentation or similar information about the model.
73
+
74
+
75
+ **You** refers to the individual or entity agreeing to these terms.
76
+
77
+
78
+ **Your company** is any legal entity, sole proprietorship, or other kind of organization that you work for, plus all organizations that have control over, are under the control of, or are under common control with that organization. **Control** means ownership of substantially all the assets of an entity, or the power to direct its management and policies by vote, contract, or otherwise. Control can be direct or indirect.
79
+
80
+
81
+ **Your licenses** are all the licenses granted to you under these terms.
82
+
83
+
84
+ **Use** means anything you do with the model or its output requiring one of your licenses.
README.md ADDED
@@ -0,0 +1,126 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ro
4
+ license: other
5
+ license_name: coqui-public-model-license-1.0.0
6
+ license_link: https://coqui.ai/cpml.txt
7
+ library_name: coqui
8
+ pipeline_tag: text-to-speech
9
+ base_model: coqui/XTTS-v2
10
+ tags:
11
+ - xtts-v2
12
+ - text-to-speech
13
+ - romanian
14
+ - moldovan-dialect
15
+ - voice-cloning
16
+ - educational-speech
17
+ datasets:
18
+ - other
19
+ ---
20
+
21
+ # XTTS-v2 for Moldovan Dialectal Romanian
22
+
23
+ This checkpoint adapts
24
+ [`coqui/XTTS-v2`](https://huggingface.co/coqui/XTTS-v2) to Romanian educational
25
+ speech with Moldovan dialectal characteristics. XTTS-v2 is a multilingual
26
+ voice-cloning model and requires a short reference recording at inference time.
27
+
28
+ ## Training data and configuration
29
+
30
+ TTS preparation started from a 35,760-segment Romanian educational-speech
31
+ corpus. Segments exceeding the configured maximum waveform length were removed,
32
+ leaving 31,472 eligible utterances (55.05 hours). With seed 42, these were split
33
+ as follows:
34
+
35
+ | Partition | Utterances | Duration |
36
+ |---|---:|---:|
37
+ | Train | 25,178 | 44.06 h |
38
+ | Evaluation | 3,147 | 5.47 h |
39
+ | Test | 3,147 | 5.52 h |
40
+
41
+ Training used an NVIDIA A100-SXM4-80GB GPU for 30 epochs, batch size 4,
42
+ gradient accumulation 63 (effective batch size 252), and learning rate
43
+ `5e-6`. The recorded execution lasted 8 hours and 6 minutes.
44
+
45
+ Corpus release:
46
+ https://drive.google.com/drive/folders/1Va0tmpp-Q8A2gGbFwLignYVtf0KD6Uth
47
+
48
+ ## Files
49
+
50
+ - `best_model.pth`: fine-tuned XTTS-v2 checkpoint.
51
+ - `config.json`: model and inference configuration.
52
+ - `vocab.json`: tokenizer vocabulary.
53
+ - `dvae.pth` and `mel_stats.pth`: XTTS-v2 acoustic assets.
54
+ - `generate_tts.py`: local inference example, including Romanian text
55
+ normalization used during evaluation.
56
+
57
+ No reference-speaker recording is included. Supply a reference WAV for a voice
58
+ that you have permission to use.
59
+
60
+ ## Installation and usage
61
+
62
+ Use Python 3.10 or 3.11 and install the dependencies from `requirements.txt`.
63
+ Download the repository, place a reference recording at
64
+ `reference_voice/reference.wav`, edit `TEXT` in `generate_tts.py`, and run:
65
+
66
+ ```bash
67
+ python generate_tts.py
68
+ ```
69
+
70
+ The included script loads `best_model.pth` explicitly, applies Romanian
71
+ cedilla-to-comma-below normalization, and synthesizes at 24 kHz. A CUDA-capable
72
+ GPU is strongly recommended.
73
+
74
+ ## Evaluation
75
+
76
+ Automatic intelligibility proxies on 10 generated prompts produced 1.68% WER
77
+ (95% CI 0.00-4.13), 0.36% CER (0.00-0.86), and generation RTF 0.388. These
78
+ scores were obtained by transcribing generated speech with the accompanying
79
+ fine-tuned ASR model and are not independent measures of speech quality.
80
+
81
+ A randomized, model-blinded pilot evaluation involved 10 listeners (8 native
82
+ Moldovan listeners), 10 matched prompts, five female and five male reference
83
+ voices, and 20 samples per listener. Naturalness and timbre similarity favored
84
+ the fine-tuned model after listener-level multiplicity correction. Moldovan
85
+ dialectal adequacy showed a positive but statistically uncertain trend;
86
+ pronunciation favored the Romanian baseline, and sentence-ending stability was
87
+ nearly tied.
88
+
89
+ ## Limitations and responsible use
90
+
91
+ - The listening study is preliminary and contains only 10 listeners.
92
+ - The training data emphasize planned educational speech, not spontaneous
93
+ conversation.
94
+ - The automatic WER/CER comparison contains only 10 prompts.
95
+ - Output quality depends strongly on reference-audio quality, text
96
+ normalization, and inference parameters.
97
+ - Voice cloning must only be performed with informed consent or another valid
98
+ legal basis. Do not use this model for impersonation, deception, fraud, or
99
+ unauthorized biometric processing.
100
+
101
+ ## Authors and project
102
+
103
+ Marius Cerescu, Alexandr Parahonco, Olesea Caftanatov, Tudor Bumbu, Nichita
104
+ Degteariov, and Ion Bostan. Vladimir Andrunachievici Institute of Mathematics
105
+ and Computer Science, Moldova State University, Chisinau, Republic of Moldova.
106
+
107
+ This work was elaborated within the project *Platforma Educationala bazata pe
108
+ Inteligenta Artificiala "Guguta"*, registered in the State Register of projects
109
+ in science and innovation, code 25.80012.0807.21TC, project leader Tudor Bumbu,
110
+ Dr.
111
+
112
+ ## Citation
113
+
114
+ ```bibtex
115
+ @inproceedings{cerescu2026moldovan,
116
+ title = {Fine-Tuning Neural Speech Models for Romanian Speech Recognition and Synthesis with Moldovan Dialectal Data},
117
+ author = {Cerescu, Marius and Parahonco, Alexandr and Caftanatov, Olesea and Bumbu, Tudor and Degteariov, Nichita and Bostan, Ion},
118
+ booktitle = {International Conference on System Analysis and Intelligent Information Technologies (SAIIT)},
119
+ year = {2026}
120
+ }
121
+ ```
122
+
123
+ ## License
124
+
125
+ This derivative model is distributed under the Coqui Public Model License 1.0.
126
+ Review `LICENSE.txt` before downloading or using the checkpoint.
best_model.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4db1d793c27cdae9956b2f0ebb5519a982268e51f3becf204278f27f4dd3086f
3
+ size 5608050938
config.json ADDED
@@ -0,0 +1,160 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "output_path": "output",
3
+ "logger_uri": null,
4
+ "run_name": "run",
5
+ "project_name": null,
6
+ "run_description": "\ud83d\udc38Coqui trainer run.",
7
+ "print_step": 25,
8
+ "plot_step": 100,
9
+ "model_param_stats": false,
10
+ "wandb_entity": null,
11
+ "dashboard_logger": "tensorboard",
12
+ "save_on_interrupt": true,
13
+ "log_model_step": null,
14
+ "save_step": 10000,
15
+ "save_n_checkpoints": 5,
16
+ "save_checkpoints": true,
17
+ "save_all_best": false,
18
+ "save_best_after": 10000,
19
+ "target_loss": null,
20
+ "print_eval": false,
21
+ "test_delay_epochs": 0,
22
+ "run_eval": true,
23
+ "run_eval_steps": null,
24
+ "distributed_backend": "nccl",
25
+ "distributed_url": "tcp://localhost:54321",
26
+ "mixed_precision": false,
27
+ "precision": "fp16",
28
+ "epochs": 1000,
29
+ "batch_size": 32,
30
+ "eval_batch_size": 16,
31
+ "grad_clip": 0.0,
32
+ "scheduler_after_epoch": true,
33
+ "lr": 0.001,
34
+ "optimizer": "radam",
35
+ "optimizer_params": null,
36
+ "lr_scheduler": null,
37
+ "lr_scheduler_params": {},
38
+ "use_grad_scaler": false,
39
+ "allow_tf32": false,
40
+ "cudnn_enable": true,
41
+ "cudnn_deterministic": false,
42
+ "cudnn_benchmark": false,
43
+ "training_seed": 54321,
44
+ "model": "xtts",
45
+ "num_loader_workers": 0,
46
+ "num_eval_loader_workers": 0,
47
+ "use_noise_augment": false,
48
+ "audio": {
49
+ "sample_rate": 22050,
50
+ "output_sample_rate": 24000
51
+ },
52
+ "use_phonemes": false,
53
+ "phonemizer": null,
54
+ "phoneme_language": null,
55
+ "compute_input_seq_cache": false,
56
+ "text_cleaner": null,
57
+ "enable_eos_bos_chars": false,
58
+ "test_sentences_file": "",
59
+ "phoneme_cache_path": null,
60
+ "characters": null,
61
+ "add_blank": false,
62
+ "batch_group_size": 0,
63
+ "loss_masking": null,
64
+ "min_audio_len": 1,
65
+ "max_audio_len": 1000000000,
66
+ "min_text_len": 1,
67
+ "max_text_len": 1000000000,
68
+ "compute_f0": false,
69
+ "compute_energy": false,
70
+ "compute_linear_spec": false,
71
+ "precompute_num_workers": 0,
72
+ "start_by_longest": false,
73
+ "shuffle": false,
74
+ "drop_last": false,
75
+ "datasets": [
76
+ {
77
+ "formatter": "",
78
+ "dataset_name": "",
79
+ "path": "",
80
+ "meta_file_train": "",
81
+ "ignored_speakers": null,
82
+ "language": "",
83
+ "phonemizer": "",
84
+ "meta_file_val": "",
85
+ "meta_file_attn_mask": ""
86
+ }
87
+ ],
88
+ "test_sentences": [],
89
+ "eval_split_max_size": null,
90
+ "eval_split_size": 0.01,
91
+ "use_speaker_weighted_sampler": false,
92
+ "speaker_weighted_sampler_alpha": 1.0,
93
+ "use_language_weighted_sampler": false,
94
+ "language_weighted_sampler_alpha": 1.0,
95
+ "use_length_weighted_sampler": false,
96
+ "length_weighted_sampler_alpha": 1.0,
97
+ "model_args": {
98
+ "gpt_batch_size": 1,
99
+ "enable_redaction": false,
100
+ "kv_cache": true,
101
+ "gpt_checkpoint": null,
102
+ "clvp_checkpoint": null,
103
+ "decoder_checkpoint": null,
104
+ "num_chars": 255,
105
+ "tokenizer_file": "",
106
+ "gpt_max_audio_tokens": 605,
107
+ "gpt_max_text_tokens": 402,
108
+ "gpt_max_prompt_tokens": 70,
109
+ "gpt_layers": 30,
110
+ "gpt_n_model_channels": 1024,
111
+ "gpt_n_heads": 16,
112
+ "gpt_number_text_tokens": 6681,
113
+ "gpt_start_text_token": null,
114
+ "gpt_stop_text_token": null,
115
+ "gpt_num_audio_tokens": 1026,
116
+ "gpt_start_audio_token": 1024,
117
+ "gpt_stop_audio_token": 1025,
118
+ "gpt_code_stride_len": 1024,
119
+ "gpt_use_masking_gt_prompt_approach": true,
120
+ "gpt_use_perceiver_resampler": true,
121
+ "input_sample_rate": 22050,
122
+ "output_sample_rate": 24000,
123
+ "output_hop_length": 256,
124
+ "decoder_input_dim": 1024,
125
+ "d_vector_dim": 512,
126
+ "cond_d_vector_in_each_upsampling_layer": true,
127
+ "duration_const": 102400
128
+ },
129
+ "model_dir": null,
130
+ "languages": [
131
+ "en",
132
+ "es",
133
+ "fr",
134
+ "de",
135
+ "it",
136
+ "pt",
137
+ "pl",
138
+ "tr",
139
+ "ru",
140
+ "nl",
141
+ "cs",
142
+ "ar",
143
+ "zh-cn",
144
+ "hu",
145
+ "ko",
146
+ "ja",
147
+ "hi",
148
+ "ro"
149
+ ],
150
+ "temperature": 0.75,
151
+ "length_penalty": 1.0,
152
+ "repetition_penalty": 5.0,
153
+ "top_k": 50,
154
+ "top_p": 0.85,
155
+ "num_gpt_outputs": 1,
156
+ "gpt_cond_len": 30,
157
+ "gpt_cond_chunk_len": 4,
158
+ "max_ref_len": 30,
159
+ "sound_norm_refs": false
160
+ }
dvae.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b29bc227d410d4991e0a8c09b858f77415013eeb9fba9650258e96095557d97a
3
+ size 210514388
generate_tts.py ADDED
@@ -0,0 +1,297 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ XTTS-v2 Moldovan TTS — Generare audio din text
3
+ Folosește modelul fine-tunat pe grai moldovenesc.
4
+
5
+ Prima rulare: extrage și salvează latentele vocii de referință (~10s).
6
+ Rulările următoare: încarcă latentele salvate (~instant).
7
+ """
8
+
9
+ import os
10
+ import sys
11
+ import time
12
+ import torch
13
+ import soundfile as sf
14
+ import numpy as np
15
+
16
+ # ══════════════════════════════════════════════════════════════
17
+ # TEXT DE GENERAT (modifică aici)
18
+ # ══════════════════════════════════════════════════════════════
19
+
20
+ TEXT = """
21
+ Bună ziua! Eu sunt un model de sinteză vocală antrenat pe grai moldovenesc.
22
+ """
23
+
24
+ # ══════════════════════════════════════════════════════════════
25
+ # CONFIGURARE
26
+ # ══════════════════════════════════════════════════════════════
27
+
28
+ SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
29
+
30
+ # Căi model
31
+ MODEL_DIR = SCRIPT_DIR
32
+ CHECKPOINT_PATH = os.path.join(MODEL_DIR, "best_model.pth")
33
+ CONFIG_PATH = os.path.join(MODEL_DIR, "config.json")
34
+ VOCAB_PATH = os.path.join(MODEL_DIR, "vocab.json")
35
+
36
+ # Voce de referință
37
+ REFERENCE_AUDIO = os.path.join(SCRIPT_DIR, "reference_voice", "reference.wav")
38
+ CACHED_LATENTS = os.path.join(SCRIPT_DIR, "reference_voice", "cached_latents.pth")
39
+
40
+ # Output
41
+ OUTPUT_DIR = os.path.join(SCRIPT_DIR, "output")
42
+ os.makedirs(OUTPUT_DIR, exist_ok=True)
43
+
44
+ # Parametri generare
45
+ TEMPERATURE = 0.7
46
+ TOP_P = 0.7
47
+ TOP_K = 30
48
+ LENGTH_PENALTY = 0.8
49
+ REPETITION_PENALTY = 10.0
50
+ GPT_COND_LEN = 6 # Secunde de audio referință pentru condiționare
51
+
52
+ # Cedilla → comma-below
53
+ CEDILLA_TO_COMMA = str.maketrans({
54
+ "\u015f": "\u0219", # ş -> ș
55
+ "\u0163": "\u021b", # ţ -> ț
56
+ "\u015e": "\u0218", # Ş -> Ș
57
+ "\u0162": "\u021a", # Ţ -> Ț
58
+ })
59
+
60
+
61
+ # ══════════════════════════════════════════════════════════════
62
+ # PATCH TOKENIZER ROMÂNESC
63
+ # ══════════════════════════════════════════════════════════════
64
+
65
+ def patch_tokenizer_for_romanian():
66
+ """Aplică patch pentru suport limba română în tokenizer-ul XTTS."""
67
+ import re
68
+ import TTS
69
+ tts_dir = os.path.dirname(TTS.__file__)
70
+ tokenizer_file = os.path.join(tts_dir, "tts", "layers", "xtts", "tokenizer.py")
71
+
72
+ if not os.path.exists(tokenizer_file):
73
+ print(f"WARN: tokenizer.py negăsit la {tokenizer_file}")
74
+ return
75
+
76
+ with open(tokenizer_file, "r", encoding="utf-8") as f:
77
+ content = f.read()
78
+ original = content
79
+
80
+ # Adaugă 'ro' la listele de limbi
81
+ if '"ro"' not in content and "'ro'" not in content:
82
+ content = re.sub(
83
+ r"""(['"])zh-cn\1(\s*)\]""",
84
+ lambda m: f'{m.group(1)}zh-cn{m.group(1)}, {m.group(1)}ro{m.group(1)}{m.group(2)}]',
85
+ content
86
+ )
87
+
88
+ # Safe .get() replacements
89
+ for old, new in {
90
+ "self.abbreviations[lang]": "self.abbreviations.get(lang, [])",
91
+ "self.symbols[lang]": "self.symbols.get(lang, [])",
92
+ "_abbreviations[lang]": "_abbreviations.get(lang, [])",
93
+ "_symbols_multilingual[lang]": "_symbols_multilingual.get(lang, [])",
94
+ 'and_equivalents[lang]': 'and_equivalents.get(lang, "and")',
95
+ '_ordinal_re[lang]': '_ordinal_re.get(lang, re.compile(r"(?!x)x"))',
96
+ }.items():
97
+ if old in content:
98
+ content = content.replace(old, new)
99
+
100
+ # Cedilla normalization block
101
+ if "_CEDILLA_TO_COMMA" not in content:
102
+ import_match = re.search(r'^import re\s*$', content, re.MULTILINE)
103
+ if import_match:
104
+ cedilla_block = (
105
+ '\n# Cedilla -> comma-below normalization for Romanian\n'
106
+ '_CEDILLA_TO_COMMA = str.maketrans({\n'
107
+ ' "\u015f": "\u0219",\n'
108
+ ' "\u0163": "\u021b",\n'
109
+ ' "\u015e": "\u0218",\n'
110
+ ' "\u0162": "\u021a",\n'
111
+ '})\n'
112
+ )
113
+ pos = import_match.end()
114
+ content = content[:pos] + cedilla_block + content[pos:]
115
+
116
+ # Inject cedilla normalization in multilingual_cleaners
117
+ if "_CEDILLA_TO_COMMA" in content:
118
+ mc_match = re.search(r'(def multilingual_cleaners\([^)]*\):\s*\n)', content)
119
+ if mc_match:
120
+ func_start = mc_match.end()
121
+ if "_CEDILLA_TO_COMMA" not in content[func_start:func_start + 500]:
122
+ injection = ' if lang == "ro":\n text = text.translate(_CEDILLA_TO_COMMA)\n'
123
+ content = content[:func_start] + injection + content[func_start:]
124
+
125
+ # Add 'ro' to _ordinal_re dict
126
+ if "_ordinal_re" in content:
127
+ ord_match = re.search(r'_ordinal_re\s*=\s*\{', content)
128
+ if ord_match:
129
+ brace_start = content.index('{', ord_match.start())
130
+ depth = 0
131
+ brace_end = brace_start
132
+ for ci in range(brace_start, len(content)):
133
+ if content[ci] == '{':
134
+ depth += 1
135
+ elif content[ci] == '}':
136
+ depth -= 1
137
+ if depth == 0:
138
+ brace_end = ci
139
+ break
140
+ ordinal_dict = content[brace_start:brace_end + 1]
141
+ if '"ro"' not in ordinal_dict and "'ro'" not in ordinal_dict:
142
+ new_dict = ordinal_dict[:-1] + ' "ro": re.compile(r"([0-9]+)\\.(?=\\s|$)"),\n}'
143
+ content = content[:brace_start] + new_dict + content[brace_end + 1:]
144
+
145
+ # Add 'ro' to preprocess_text
146
+ for old_pat, new_pat in [
147
+ ('"pt", "ru"', '"pt", "ro", "ru"'),
148
+ ("'pt', 'ru'", "'pt', 'ro', 'ru'"),
149
+ ]:
150
+ if old_pat in content:
151
+ content = content.replace(old_pat, new_pat, 1)
152
+ break
153
+
154
+ if content != original:
155
+ with open(tokenizer_file, "w", encoding="utf-8") as f:
156
+ f.write(content)
157
+ # Reload affected modules
158
+ mods_to_reload = [k for k in sys.modules if 'xtts' in k.lower() or 'tokenizer' in k.lower()]
159
+ for mod in mods_to_reload:
160
+ del sys.modules[mod]
161
+ print(" Tokenizer patch românesc aplicat!")
162
+ else:
163
+ print(" Tokenizer deja patch-uit.")
164
+
165
+
166
+ # ══════════════════════════════════════════════════════════════
167
+ # MAIN
168
+ # ══════════════════════════════════════════════════════════════
169
+
170
+ def main():
171
+ t_start = time.time()
172
+
173
+ # Verificări
174
+ if not os.path.exists(CHECKPOINT_PATH):
175
+ print(f"ERROR: best_model.pth nu există la {CHECKPOINT_PATH}")
176
+ sys.exit(1)
177
+ if not os.path.exists(REFERENCE_AUDIO):
178
+ print(f"ERROR: Audio de referință nu există la {REFERENCE_AUDIO}")
179
+ print(f"Copiază fișierul WAV în: {os.path.dirname(REFERENCE_AUDIO)}")
180
+ sys.exit(1)
181
+
182
+ # GPU check
183
+ if torch.cuda.is_available():
184
+ gpu_name = torch.cuda.get_device_name(0)
185
+ vram = torch.cuda.get_device_properties(0).total_memory / 1024**3
186
+ print(f"GPU: {gpu_name} ({vram:.1f} GB VRAM)")
187
+ else:
188
+ print("WARN: GPU nu a fost detectat! Generarea va fi foarte lentă.")
189
+
190
+ # Patch tokenizer
191
+ print("\nAplicare patch tokenizer românesc...")
192
+ patch_tokenizer_for_romanian()
193
+
194
+ # Patch torchaudio.load pentru a folosi soundfile (evită FFmpeg/torchcodec pe Windows)
195
+ import torchaudio
196
+ def _soundfile_load(filepath, **kwargs):
197
+ data, sr = sf.read(filepath, dtype='float32')
198
+ if data.ndim == 1:
199
+ data = data[np.newaxis, :]
200
+ else:
201
+ data = data.T
202
+ return torch.from_numpy(data), sr
203
+ torchaudio.load = _soundfile_load
204
+
205
+ # Încărcare model
206
+ print("\nÎncărcare model XTTS-v2...")
207
+ t_load = time.time()
208
+
209
+ from TTS.tts.configs.xtts_config import XttsConfig
210
+ from TTS.tts.models.xtts import Xtts
211
+
212
+ config = XttsConfig()
213
+ config.load_json(CONFIG_PATH)
214
+
215
+ model = Xtts.init_from_config(config)
216
+ model.load_checkpoint(
217
+ config,
218
+ checkpoint_path=CHECKPOINT_PATH,
219
+ vocab_path=VOCAB_PATH,
220
+ use_deepspeed=False,
221
+ )
222
+ model.cuda()
223
+ model.eval()
224
+ print(f" Model încărcat în {time.time() - t_load:.1f}s")
225
+
226
+ # Latente vocale (cache)
227
+ if os.path.exists(CACHED_LATENTS):
228
+ print(f"\nÎncărcare latente din cache...")
229
+ cached = torch.load(CACHED_LATENTS, weights_only=True)
230
+ gpt_cond_latent = cached["gpt_cond_latent"].cuda()
231
+ speaker_embedding = cached["speaker_embedding"].cuda()
232
+ print(" Latente încărcate din cache!")
233
+ else:
234
+ print(f"\nExtragere latente din: {os.path.basename(REFERENCE_AUDIO)}")
235
+ gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
236
+ audio_path=[REFERENCE_AUDIO],
237
+ gpt_cond_len=GPT_COND_LEN,
238
+ )
239
+ # Salvare cache
240
+ torch.save({
241
+ "gpt_cond_latent": gpt_cond_latent.cpu(),
242
+ "speaker_embedding": speaker_embedding.cpu(),
243
+ }, CACHED_LATENTS)
244
+ print(f" Latente salvate în cache: {CACHED_LATENTS}")
245
+
246
+ # Pregătire text
247
+ text = TEXT.strip()
248
+ if not text:
249
+ print("ERROR: TEXT este gol!")
250
+ sys.exit(1)
251
+
252
+ text_clean = text.translate(CEDILLA_TO_COMMA)
253
+
254
+ # Generare
255
+ print(f"\nGenerare audio...")
256
+ print(f" Text: {text_clean[:100]}{'...' if len(text_clean) > 100 else ''}")
257
+
258
+ t_gen = time.time()
259
+ with torch.no_grad():
260
+ output = model.inference(
261
+ text_clean,
262
+ "ro",
263
+ gpt_cond_latent,
264
+ speaker_embedding,
265
+ temperature=TEMPERATURE,
266
+ top_p=TOP_P,
267
+ top_k=TOP_K,
268
+ length_penalty=LENGTH_PENALTY,
269
+ repetition_penalty=REPETITION_PENALTY,
270
+ )
271
+
272
+ wav = output["wav"]
273
+ if isinstance(wav, torch.Tensor):
274
+ wav = wav.cpu().numpy()
275
+ wav = wav.squeeze()
276
+
277
+ gen_time = time.time() - t_gen
278
+ duration = len(wav) / 24000
279
+
280
+ # Salvare
281
+ from datetime import datetime
282
+ timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
283
+ out_path = os.path.join(OUTPUT_DIR, f"tts_{timestamp}.wav")
284
+ sf.write(out_path, wav, 24000)
285
+
286
+ # Raport
287
+ print(f"\n{'=' * 50}")
288
+ print(f" Audio generat: {out_path}")
289
+ print(f" Durată audio: {duration:.2f}s")
290
+ print(f" Timp generare: {gen_time:.2f}s")
291
+ print(f" Timp total: {time.time() - t_start:.1f}s")
292
+ print(f" RTF: {gen_time / duration:.2f}x")
293
+ print(f"{'=' * 50}")
294
+
295
+
296
+ if __name__ == "__main__":
297
+ main()
mel_stats.pth ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1f69422a8a8f344c4fca2f0c6b8d41d2151d6615b7321e48e6bb15ae949b119c
3
+ size 1067
requirements.txt ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ TTS==0.22.0
2
+ torch
3
+ torchaudio
4
+ soundfile
5
+ numpy
vocab.json ADDED
The diff for this file is too large to render. See raw diff