monatolmats commited on
Commit
f079b3f
·
verified ·
1 Parent(s): f2e3565

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +140 -0
README.md ADDED
@@ -0,0 +1,140 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - et
4
+ license: mit
5
+ base_model: sarulab-speech/UTMOSv2
6
+ tags:
7
+ - audio
8
+ - speech
9
+ - mos-prediction
10
+ - speech-quality-assessment
11
+ - text-to-speech
12
+ - estonian
13
+ - low-resource
14
+ ---
15
+
16
+ # UTMOSv2-Estonian
17
+
18
+ [UTMOSv2](https://github.com/sarulab-speech/UTMOSv2) (Baba et al., 2024) fine-tuned
19
+ on Estonian listening-test data to predict the naturalness mean opinion score
20
+ (MOS) of synthesized speech.
21
+
22
+ The mean opinion score quantifies the naturalness of synthesized speech by
23
+ averaging subjective ratings from multiple listeners, and is the standard measure
24
+ in speech synthesis evaluation. Obtaining it requires a listening test, which is
25
+ labour-intensive and time-consuming — a cost that falls hardest on languages with
26
+ few speakers, where recruiting sufficient raters is itself difficult.
27
+
28
+ **[Interactive demo](https://huggingface.co/spaces/monatolmats/utmosv2-estonian-demo)**
29
+
30
+ ## Scope of this work
31
+
32
+ The architecture and the English pretraining are the work of the UTMOSv2 authors.
33
+ The contribution here is the fine-tuning on Estonian data and its evaluation on
34
+ Estonian and Võro, carried out for the bachelor's thesis *Automatic Speech
35
+ Synthesis Quality Assessment for Finno-Ugric Languages* (Mona Tolmats, University
36
+ of Tartu, 2025; supervisor Liisa Rätsep).
37
+
38
+ ## Usage
39
+
40
+ UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a
41
+ corpus it has not seen. During fine-tuning the Estonian data occupied three of
42
+ UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2
43
+ authors, each clip should be scored under all three and the results averaged.
44
+ **Scoring under a single arbitrary domain will produce systematically different
45
+ values.**
46
+
47
+ ```bash
48
+ pip install git+https://github.com/sarulab-speech/UTMOSv2.git
49
+ huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights
50
+ ```
51
+
52
+ ```python
53
+ import utmosv2
54
+
55
+ model = utmosv2.create_model(
56
+ config="fusion_stage3",
57
+ checkpoint_path="weights/utmosv2_estonian.pth",
58
+ )
59
+
60
+ DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"]
61
+
62
+ runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS]
63
+ for rows in zip(*runs):
64
+ mos = sum(r["predicted_mos"] for r in rows) / len(rows)
65
+ print(f"{mos:.2f} {rows[0]['file_path']}")
66
+ ```
67
+
68
+ Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale.
69
+
70
+ ## Model details
71
+
72
+ | | |
73
+ | --- | --- |
74
+ | Architecture | UTMOSv2 `fusion_stage3`: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding |
75
+ | Initialised from | `fold0_s42_best_model.pth`, a single fold of the released 5-fold ensemble |
76
+ | Optimiser | AdamW, learning rate 1e-4, cosine annealing |
77
+ | Batch size | 4 |
78
+ | Input length | first 10 s of each clip |
79
+ | Stopping | epoch 33, on validation MSE |
80
+ | Hardware | NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster |
81
+
82
+ ## Training data
83
+
84
+ Listening-test results from evaluation campaigns run at the University of Tartu
85
+ (Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign
86
+ was held out entirely to assess cross-lingual transfer.
87
+
88
+ | Year | Language | Systems | Clips | Ratings | Role |
89
+ | --- | --- | --- | --- | --- | --- |
90
+ | 2020 | Estonian | 5 | 850 | 17 000 | fine-tuning |
91
+ | 2022 | Estonian | 7 | 1 400 | 5 600 | fine-tuning |
92
+ | 2024 | Estonian | 9 | 2 560 | 12 800 | fine-tuning |
93
+ | 2023 | Võro | 8 | 800 | 2 600 | held-out test |
94
+
95
+ The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation,
96
+ stratified by MOS and by source campaign. Ratings were averaged per clip, so
97
+ targets are continuous. The distribution is skewed toward mid and high scores and
98
+ was left unbalanced, to reflect the composition of real listening tests.
99
+
100
+ The data is not released; it belongs to the projects cited in the thesis.
101
+
102
+ ## Evaluation
103
+
104
+ Estonian validation set:
105
+
106
+ | Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
107
+ | --- | --- | --- | --- | --- |
108
+ | wav2vec 2.0 | 0.352 | 0.669 | 0.630 | 0.464 |
109
+ | SCOREQ | **0.224** | **0.802** | 0.768 | 0.589 |
110
+ | UTMOSv2-Estonian | 0.230 | 0.797 | **0.784** | **0.600** |
111
+
112
+ Võro test set, unseen during fine-tuning:
113
+
114
+ | Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
115
+ | --- | --- | --- | --- | --- |
116
+ | wav2vec 2.0 | 0.639 | 0.177 | 0.168 | 0.118 |
117
+ | SCOREQ | 0.632 | 0.202 | 0.191 | 0.132 |
118
+ | UTMOSv2-Estonian | **0.606** | **0.349** | **0.311** | **0.220** |
119
+
120
+ Utterance-level correlation on Võro is considerably lower than on Estonian. At
121
+ the system level the picture is better: of the three models compared, this was
122
+ the only one to rank all eight Võro synthesis systems in the same order as human
123
+ raters, and the only one whose 95 % confidence intervals overlapped the human
124
+ intervals for every system.
125
+
126
+ ## Limitations
127
+
128
+ - **Intended for ranking, not certification.** Like MOS predictors generally,
129
+ this model is suited to ordering systems rather than grading individual clips.
130
+ Differences below 0.2 MOS on a single utterance fall within noise; averaging
131
+ over at least 20 utterances is advisable when comparing systems.
132
+ - **Predictions are not deterministic.** The spectrogram branch samples random
133
+ windows, so repeated scoring of the same file varies slightly.
134
+ - **Cross-lingual transfer is demonstrated only for Võro**, which is closely
135
+ related to Estonian. No claim is made about more distant Finno-Ugric languages.
136
+ - **Narrow rating range in the training data.** Very low scores are almost absent,
137
+ so discrimination between good and excellent systems is weaker than between
138
+ poor and good ones.
139
+ - **Single fold**, not the 5-fold ensemble of the released UTMOSv2, so variance is
140
+ higher than the published UTMOSv2 figures.