utmosv2-estonian / README.md
monatolmats's picture
Update README.md
88ecebb verified
|
Raw History Blame Contribute Delete
5.65 kB
metadata
language:
  - et
license: mit
base_model: sarulab-speech/UTMOSv2
tags:
  - audio
  - speech
  - mos-prediction
  - speech-quality-assessment
  - text-to-speech
  - estonian
  - low-resource

UTMOSv2-Estonian

UTMOSv2 (Baba et al., 2024) fine-tuned on Estonian listening-test data to predict the mean opinion score (MOS) of synthesized speech.

The mean opinion score quantifies the naturalness of synthesized speech by averaging subjective ratings from multiple listeners, and is the standard measure in speech synthesis evaluation. Obtaining it requires a listening test, which is labour-intensive and time-consuming — a cost that falls hardest on languages with few speakers, where recruiting sufficient raters is itself difficult.

Interactive demo

Scope of this work

The architecture and the English pretraining are the work of the UTMOSv2 authors. The contribution here is the fine-tuning on Estonian data and its evaluation on Estonian and Võro, carried out for the bachelor's thesis Automatic Speech Synthesis Quality Assessment for Finno-Ugric Languages (Mona Tolmats, University of Tartu, 2025; supervisor Liisa Rätsep).

Usage

UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a corpus it has not seen. During fine-tuning the Estonian data occupied three of UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2 authors, each clip should be scored under all three and the results averaged. Scoring under a single arbitrary domain will produce systematically different values.

pip install git+https://github.com/sarulab-speech/UTMOSv2.git
huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights
import utmosv2

model = utmosv2.create_model(
    config="fusion_stage3",
    checkpoint_path="weights/utmosv2_estonian.pth",
)

DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"]

runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS]
for rows in zip(*runs):
    mos = sum(r["predicted_mos"] for r in rows) / len(rows)
    print(f"{mos:.2f}  {rows[0]['file_path']}")

Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale.

Model details

Architecture UTMOSv2 fusion_stage3: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding
Initialised from fold0_s42_best_model.pth, a single fold of the released 5-fold ensemble
Optimiser AdamW, learning rate 1e-4, cosine annealing
Batch size 4
Input length first 10 s of each clip
Stopping epoch 33, on validation MSE
Hardware NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster

Training data

Listening-test results from evaluation campaigns run at the University of Tartu (Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign was held out entirely to assess cross-lingual transfer.

Year Language Systems Clips Ratings Role
2020 Estonian 5 850 17 000 fine-tuning
2022 Estonian 7 1 400 5 600 fine-tuning
2024 Estonian 9 2 560 12 800 fine-tuning
2023 Võro 8 800 2 600 held-out test

The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation, stratified by MOS and by source campaign. Ratings were averaged per clip, so targets are continuous. The distribution is skewed toward mid and high scores and was left unbalanced, to reflect the composition of real listening tests.

The data is not released; it belongs to the projects cited in the thesis.

Evaluation

Estonian validation set:

Model MSE ↓ LCC ↑ SRCC ↑ KTAU ↑
wav2vec 2.0 0.352 0.669 0.630 0.464
SCOREQ 0.224 0.802 0.768 0.589
UTMOSv2-Estonian 0.230 0.797 0.784 0.600

Võro test set, unseen during fine-tuning:

Model MSE ↓ LCC ↑ SRCC ↑ KTAU ↑
wav2vec 2.0 0.639 0.177 0.168 0.118
SCOREQ 0.632 0.202 0.191 0.132
UTMOSv2-Estonian 0.606 0.349 0.311 0.220

Utterance-level correlation on Võro is considerably lower than on Estonian. At the system level the picture is better: of the three models compared, this was the only one to rank all eight Võro synthesis systems in the same order as human raters, and the only one whose 95 % confidence intervals overlapped the human intervals for every system.

Limitations

  • Intended for ranking, not certification. Like MOS predictors generally, this model is suited to ordering systems rather than grading individual clips. Differences below 0.2 MOS on a single utterance fall within noise; averaging over at least 20 utterances is advisable when comparing systems.
  • Predictions are not deterministic. The spectrogram branch samples random windows, so repeated scoring of the same file varies slightly.
  • Cross-lingual transfer is demonstrated only for Võro, which is closely related to Estonian. No claim is made about more distant Finno-Ugric languages.
  • Narrow rating range in the training data. Very low scores are almost absent, so discrimination between good and excellent systems is weaker than between poor and good ones.
  • Single fold, not the 5-fold ensemble of the released UTMOSv2, so variance is higher than the published UTMOSv2 figures.