File size: 5,647 Bytes
f079b3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
88ecebb
f079b3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
language:
  - et
license: mit
base_model: sarulab-speech/UTMOSv2
tags:
  - audio
  - speech
  - mos-prediction
  - speech-quality-assessment
  - text-to-speech
  - estonian
  - low-resource
---

# UTMOSv2-Estonian

[UTMOSv2](https://github.com/sarulab-speech/UTMOSv2) (Baba et al., 2024) fine-tuned
on Estonian listening-test data to predict the mean opinion score
(MOS) of synthesized speech.

The mean opinion score quantifies the naturalness of synthesized speech by
averaging subjective ratings from multiple listeners, and is the standard measure
in speech synthesis evaluation. Obtaining it requires a listening test, which is
labour-intensive and time-consuming — a cost that falls hardest on languages with
few speakers, where recruiting sufficient raters is itself difficult.

**[Interactive demo](https://huggingface.co/spaces/monatolmats/utmosv2-estonian-demo)**

## Scope of this work

The architecture and the English pretraining are the work of the UTMOSv2 authors.
The contribution here is the fine-tuning on Estonian data and its evaluation on
Estonian and Võro, carried out for the bachelor's thesis *Automatic Speech
Synthesis Quality Assessment for Finno-Ugric Languages* (Mona Tolmats, University
of Tartu, 2025; supervisor Liisa Rätsep).

## Usage

UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a
corpus it has not seen. During fine-tuning the Estonian data occupied three of
UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2
authors, each clip should be scored under all three and the results averaged.
**Scoring under a single arbitrary domain will produce systematically different
values.**

```bash
pip install git+https://github.com/sarulab-speech/UTMOSv2.git
huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights
```

```python
import utmosv2

model = utmosv2.create_model(
    config="fusion_stage3",
    checkpoint_path="weights/utmosv2_estonian.pth",
)

DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"]

runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS]
for rows in zip(*runs):
    mos = sum(r["predicted_mos"] for r in rows) / len(rows)
    print(f"{mos:.2f}  {rows[0]['file_path']}")
```

Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale.

## Model details

| | |
| --- | --- |
| Architecture | UTMOSv2 `fusion_stage3`: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding |
| Initialised from | `fold0_s42_best_model.pth`, a single fold of the released 5-fold ensemble |
| Optimiser | AdamW, learning rate 1e-4, cosine annealing |
| Batch size | 4 |
| Input length | first 10 s of each clip |
| Stopping | epoch 33, on validation MSE |
| Hardware | NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster |

## Training data

Listening-test results from evaluation campaigns run at the University of Tartu
(Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign
was held out entirely to assess cross-lingual transfer.

| Year | Language | Systems | Clips | Ratings | Role |
| --- | --- | --- | --- | --- | --- |
| 2020 | Estonian | 5 | 850 | 17 000 | fine-tuning |
| 2022 | Estonian | 7 | 1 400 | 5 600 | fine-tuning |
| 2024 | Estonian | 9 | 2 560 | 12 800 | fine-tuning |
| 2023 | Võro | 8 | 800 | 2 600 | held-out test |

The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation,
stratified by MOS and by source campaign. Ratings were averaged per clip, so
targets are continuous. The distribution is skewed toward mid and high scores and
was left unbalanced, to reflect the composition of real listening tests.

The data is not released; it belongs to the projects cited in the thesis.

## Evaluation

Estonian validation set:

| Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
| --- | --- | --- | --- | --- |
| wav2vec 2.0 | 0.352 | 0.669 | 0.630 | 0.464 |
| SCOREQ | **0.224** | **0.802** | 0.768 | 0.589 |
| UTMOSv2-Estonian | 0.230 | 0.797 | **0.784** | **0.600** |

Võro test set, unseen during fine-tuning:

| Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
| --- | --- | --- | --- | --- |
| wav2vec 2.0 | 0.639 | 0.177 | 0.168 | 0.118 |
| SCOREQ | 0.632 | 0.202 | 0.191 | 0.132 |
| UTMOSv2-Estonian | **0.606** | **0.349** | **0.311** | **0.220** |

Utterance-level correlation on Võro is considerably lower than on Estonian. At
the system level the picture is better: of the three models compared, this was
the only one to rank all eight Võro synthesis systems in the same order as human
raters, and the only one whose 95 % confidence intervals overlapped the human
intervals for every system.

## Limitations

- **Intended for ranking, not certification.** Like MOS predictors generally,
  this model is suited to ordering systems rather than grading individual clips.
  Differences below 0.2 MOS on a single utterance fall within noise; averaging
  over at least 20 utterances is advisable when comparing systems.
- **Predictions are not deterministic.** The spectrogram branch samples random
  windows, so repeated scoring of the same file varies slightly.
- **Cross-lingual transfer is demonstrated only for Võro**, which is closely
  related to Estonian. No claim is made about more distant Finno-Ugric languages.
- **Narrow rating range in the training data.** Very low scores are almost absent,
  so discrimination between good and excellent systems is weaker than between
  poor and good ones.
- **Single fold**, not the 5-fold ensemble of the released UTMOSv2, so variance is
  higher than the published UTMOSv2 figures.