UTMOSv2-Estonian
UTMOSv2 (Baba et al., 2024) fine-tuned on Estonian listening-test data to predict the mean opinion score (MOS) of synthesized speech.
The mean opinion score quantifies the naturalness of synthesized speech by averaging subjective ratings from multiple listeners, and is the standard measure in speech synthesis evaluation. Obtaining it requires a listening test, which is labour-intensive and time-consuming — a cost that falls hardest on languages with few speakers, where recruiting sufficient raters is itself difficult.
Scope of this work
The architecture and the English pretraining are the work of the UTMOSv2 authors. The contribution here is the fine-tuning on Estonian data and its evaluation on Estonian and Võro, carried out for the bachelor's thesis Automatic Speech Synthesis Quality Assessment for Finno-Ugric Languages (Mona Tolmats, University of Tartu, 2025; supervisor Liisa Rätsep).
Usage
UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a corpus it has not seen. During fine-tuning the Estonian data occupied three of UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2 authors, each clip should be scored under all three and the results averaged. Scoring under a single arbitrary domain will produce systematically different values.
pip install git+https://github.com/sarulab-speech/UTMOSv2.git
huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights
import utmosv2
model = utmosv2.create_model(
config="fusion_stage3",
checkpoint_path="weights/utmosv2_estonian.pth",
)
DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"]
runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS]
for rows in zip(*runs):
mos = sum(r["predicted_mos"] for r in rows) / len(rows)
print(f"{mos:.2f} {rows[0]['file_path']}")
Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale.
Model details
| Architecture | UTMOSv2 fusion_stage3: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding |
| Initialised from | fold0_s42_best_model.pth, a single fold of the released 5-fold ensemble |
| Optimiser | AdamW, learning rate 1e-4, cosine annealing |
| Batch size | 4 |
| Input length | first 10 s of each clip |
| Stopping | epoch 33, on validation MSE |
| Hardware | NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster |
Training data
Listening-test results from evaluation campaigns run at the University of Tartu (Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign was held out entirely to assess cross-lingual transfer.
| Year | Language | Systems | Clips | Ratings | Role |
|---|---|---|---|---|---|
| 2020 | Estonian | 5 | 850 | 17 000 | fine-tuning |
| 2022 | Estonian | 7 | 1 400 | 5 600 | fine-tuning |
| 2024 | Estonian | 9 | 2 560 | 12 800 | fine-tuning |
| 2023 | Võro | 8 | 800 | 2 600 | held-out test |
The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation, stratified by MOS and by source campaign. Ratings were averaged per clip, so targets are continuous. The distribution is skewed toward mid and high scores and was left unbalanced, to reflect the composition of real listening tests.
The data is not released; it belongs to the projects cited in the thesis.
Evaluation
Estonian validation set:
| Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
|---|---|---|---|---|
| wav2vec 2.0 | 0.352 | 0.669 | 0.630 | 0.464 |
| SCOREQ | 0.224 | 0.802 | 0.768 | 0.589 |
| UTMOSv2-Estonian | 0.230 | 0.797 | 0.784 | 0.600 |
Võro test set, unseen during fine-tuning:
| Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ |
|---|---|---|---|---|
| wav2vec 2.0 | 0.639 | 0.177 | 0.168 | 0.118 |
| SCOREQ | 0.632 | 0.202 | 0.191 | 0.132 |
| UTMOSv2-Estonian | 0.606 | 0.349 | 0.311 | 0.220 |
Utterance-level correlation on Võro is considerably lower than on Estonian. At the system level the picture is better: of the three models compared, this was the only one to rank all eight Võro synthesis systems in the same order as human raters, and the only one whose 95 % confidence intervals overlapped the human intervals for every system.
Limitations
- Intended for ranking, not certification. Like MOS predictors generally, this model is suited to ordering systems rather than grading individual clips. Differences below 0.2 MOS on a single utterance fall within noise; averaging over at least 20 utterances is advisable when comparing systems.
- Predictions are not deterministic. The spectrogram branch samples random windows, so repeated scoring of the same file varies slightly.
- Cross-lingual transfer is demonstrated only for Võro, which is closely related to Estonian. No claim is made about more distant Finno-Ugric languages.
- Narrow rating range in the training data. Very low scores are almost absent, so discrimination between good and excellent systems is weaker than between poor and good ones.
- Single fold, not the 5-fold ensemble of the released UTMOSv2, so variance is higher than the published UTMOSv2 figures.
Model tree for monatolmats/utmosv2-estonian
Base model
sarulab-speech/UTMOSv2