--- language: - et license: mit base_model: sarulab-speech/UTMOSv2 tags: - audio - speech - mos-prediction - speech-quality-assessment - text-to-speech - estonian - low-resource --- # UTMOSv2-Estonian [UTMOSv2](https://github.com/sarulab-speech/UTMOSv2) (Baba et al., 2024) fine-tuned on Estonian listening-test data to predict the mean opinion score (MOS) of synthesized speech. The mean opinion score quantifies the naturalness of synthesized speech by averaging subjective ratings from multiple listeners, and is the standard measure in speech synthesis evaluation. Obtaining it requires a listening test, which is labour-intensive and time-consuming — a cost that falls hardest on languages with few speakers, where recruiting sufficient raters is itself difficult. **[Interactive demo](https://huggingface.co/spaces/monatolmats/utmosv2-estonian-demo)** ## Scope of this work The architecture and the English pretraining are the work of the UTMOSv2 authors. The contribution here is the fine-tuning on Estonian data and its evaluation on Estonian and Võro, carried out for the bachelor's thesis *Automatic Speech Synthesis Quality Assessment for Finno-Ugric Languages* (Mona Tolmats, University of Tartu, 2025; supervisor Liisa Rätsep). ## Usage UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a corpus it has not seen. During fine-tuning the Estonian data occupied three of UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2 authors, each clip should be scored under all three and the results averaged. **Scoring under a single arbitrary domain will produce systematically different values.** ```bash pip install git+https://github.com/sarulab-speech/UTMOSv2.git huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights ``` ```python import utmosv2 model = utmosv2.create_model( config="fusion_stage3", checkpoint_path="weights/utmosv2_estonian.pth", ) DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"] runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS] for rows in zip(*runs): mos = sum(r["predicted_mos"] for r in rows) / len(rows) print(f"{mos:.2f} {rows[0]['file_path']}") ``` Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale. ## Model details | | | | --- | --- | | Architecture | UTMOSv2 `fusion_stage3`: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding | | Initialised from | `fold0_s42_best_model.pth`, a single fold of the released 5-fold ensemble | | Optimiser | AdamW, learning rate 1e-4, cosine annealing | | Batch size | 4 | | Input length | first 10 s of each clip | | Stopping | epoch 33, on validation MSE | | Hardware | NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster | ## Training data Listening-test results from evaluation campaigns run at the University of Tartu (Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign was held out entirely to assess cross-lingual transfer. | Year | Language | Systems | Clips | Ratings | Role | | --- | --- | --- | --- | --- | --- | | 2020 | Estonian | 5 | 850 | 17 000 | fine-tuning | | 2022 | Estonian | 7 | 1 400 | 5 600 | fine-tuning | | 2024 | Estonian | 9 | 2 560 | 12 800 | fine-tuning | | 2023 | Võro | 8 | 800 | 2 600 | held-out test | The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation, stratified by MOS and by source campaign. Ratings were averaged per clip, so targets are continuous. The distribution is skewed toward mid and high scores and was left unbalanced, to reflect the composition of real listening tests. The data is not released; it belongs to the projects cited in the thesis. ## Evaluation Estonian validation set: | Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ | | --- | --- | --- | --- | --- | | wav2vec 2.0 | 0.352 | 0.669 | 0.630 | 0.464 | | SCOREQ | **0.224** | **0.802** | 0.768 | 0.589 | | UTMOSv2-Estonian | 0.230 | 0.797 | **0.784** | **0.600** | Võro test set, unseen during fine-tuning: | Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ | | --- | --- | --- | --- | --- | | wav2vec 2.0 | 0.639 | 0.177 | 0.168 | 0.118 | | SCOREQ | 0.632 | 0.202 | 0.191 | 0.132 | | UTMOSv2-Estonian | **0.606** | **0.349** | **0.311** | **0.220** | Utterance-level correlation on Võro is considerably lower than on Estonian. At the system level the picture is better: of the three models compared, this was the only one to rank all eight Võro synthesis systems in the same order as human raters, and the only one whose 95 % confidence intervals overlapped the human intervals for every system. ## Limitations - **Intended for ranking, not certification.** Like MOS predictors generally, this model is suited to ordering systems rather than grading individual clips. Differences below 0.2 MOS on a single utterance fall within noise; averaging over at least 20 utterances is advisable when comparing systems. - **Predictions are not deterministic.** The spectrogram branch samples random windows, so repeated scoring of the same file varies slightly. - **Cross-lingual transfer is demonstrated only for Võro**, which is closely related to Estonian. No claim is made about more distant Finno-Ugric languages. - **Narrow rating range in the training data.** Very low scores are almost absent, so discrimination between good and excellent systems is weaker than between poor and good ones. - **Single fold**, not the 5-fold ensemble of the released UTMOSv2, so variance is higher than the published UTMOSv2 figures.