File size: 5,647 Bytes
f079b3f 88ecebb f079b3f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | ---
language:
- et
license: mit
base_model: sarulab-speech/UTMOSv2
tags:
- audio
- speech
- mos-prediction
- speech-quality-assessment
- text-to-speech
- estonian
- low-resource
---
# UTMOSv2-Estonian
[UTMOSv2](https://github.com/sarulab-speech/UTMOSv2) (Baba et al., 2024) fine-tuned
on Estonian listening-test data to predict the mean opinion score
(MOS) of synthesized speech.
The mean opinion score quantifies the naturalness of synthesized speech by
averaging subjective ratings from multiple listeners, and is the standard measure
in speech synthesis evaluation. Obtaining it requires a listening test, which is
labour-intensive and time-consuming — a cost that falls hardest on languages with
few speakers, where recruiting sufficient raters is itself difficult.
**[Interactive demo](https://huggingface.co/spaces/monatolmats/utmosv2-estonian-demo)**
## Scope of this work
The architecture and the English pretraining are the work of the UTMOSv2 authors.
The contribution here is the fine-tuning on Estonian data and its evaluation on
Estonian and Võro, carried out for the bachelor's thesis *Automatic Speech
Synthesis Quality Assessment for Finno-Ugric Languages* (Mona Tolmats, University
of Tartu, 2025; supervisor Liisa Rätsep).
## Usage
UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a
corpus it has not seen. During fine-tuning the Estonian data occupied three of
UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2
authors, each clip should be scored under all three and the results averaged.
**Scoring under a single arbitrary domain will produce systematically different
values.**
```bash
pip install git+https://github.com/sarulab-speech/UTMOSv2.git
huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights
```
```python
import utmosv2
model = utmosv2.create_model(
config="fusion_stage3",
checkpoint_path="weights/utmosv2_estonian.pth",
)
DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"]
runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS]
for rows in zip(*runs):
mos = sum(r["predicted_mos"] for r in rows) / len(rows)
print(f"{mos:.2f} {rows[0]['file_path']}")
```
Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale.
## Model details
| | |
| --- | --- |
| Architecture | UTMOSv2 `fusion_stage3`: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding |
| Initialised from | `fold0_s42_best_model.pth`, a single fold of the released 5-fold ensemble |
| Optimiser | AdamW, learning rate 1e-4, cosine annealing |
| Batch size | 4 |
| Input length | first 10 s of each clip |
| Stopping | epoch 33, on validation MSE |
| Hardware | NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster |
## Training data
Listening-test results from evaluation campaigns run at the University of Tartu
(Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign
was held out entirely to assess cross-lingual transfer.
| Year | Language | Systems | Clips | Ratings | Role |
| --- | --- | --- | --- | --- | --- |
| 2020 | Estonian | 5 | 850 | 17 000 | fine-tuning |
| 2022 | Estonian | 7 | 1 400 | 5 600 | fine-tuning |
| 2024 | Estonian | 9 | 2 560 | 12 800 | fine-tuning |
| 2023 | Võro | 8 | 800 | 2 600 | held-out test |
The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation,
stratified by MOS and by source campaign. Ratings were averaged per clip, so
targets are continuous. The distribution is skewed toward mid and high scores and
was left unbalanced, to reflect the composition of real listening tests.
The data is not released; it belongs to the projects cited in the thesis.
## Evaluation
Estonian validation set:
| Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
| --- | --- | --- | --- | --- |
| wav2vec 2.0 | 0.352 | 0.669 | 0.630 | 0.464 |
| SCOREQ | **0.224** | **0.802** | 0.768 | 0.589 |
| UTMOSv2-Estonian | 0.230 | 0.797 | **0.784** | **0.600** |
Võro test set, unseen during fine-tuning:
| Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
| --- | --- | --- | --- | --- |
| wav2vec 2.0 | 0.639 | 0.177 | 0.168 | 0.118 |
| SCOREQ | 0.632 | 0.202 | 0.191 | 0.132 |
| UTMOSv2-Estonian | **0.606** | **0.349** | **0.311** | **0.220** |
Utterance-level correlation on Võro is considerably lower than on Estonian. At
the system level the picture is better: of the three models compared, this was
the only one to rank all eight Võro synthesis systems in the same order as human
raters, and the only one whose 95 % confidence intervals overlapped the human
intervals for every system.
## Limitations
- **Intended for ranking, not certification.** Like MOS predictors generally,
this model is suited to ordering systems rather than grading individual clips.
Differences below 0.2 MOS on a single utterance fall within noise; averaging
over at least 20 utterances is advisable when comparing systems.
- **Predictions are not deterministic.** The spectrogram branch samples random
windows, so repeated scoring of the same file varies slightly.
- **Cross-lingual transfer is demonstrated only for Võro**, which is closely
related to Estonian. No claim is made about more distant Finno-Ugric languages.
- **Narrow rating range in the training data.** Very low scores are almost absent,
so discrimination between good and excellent systems is weaker than between
poor and good ones.
- **Single fold**, not the 5-fold ensemble of the released UTMOSv2, so variance is
higher than the published UTMOSv2 figures. |