|
Download README.md from monatolmats/utmosv2-estonian: direct link, hf CLI and curl.
- Browser
- Download file 5.65 kB
-
https://huggingface.co/monatolmats/utmosv2-estonian/resolve/main/README.md
- Command line
-
hf download hf://monatolmats/utmosv2-estonian/README.md
-
curl -L -o README.md https://huggingface.co/monatolmats/utmosv2-estonian/resolve/main/README.md
5.65 kB
| language: | |
| - et | |
| license: mit | |
| base_model: sarulab-speech/UTMOSv2 | |
| tags: | |
| - audio | |
| - speech | |
| - mos-prediction | |
| - speech-quality-assessment | |
| - text-to-speech | |
| - estonian | |
| - low-resource | |
| # UTMOSv2-Estonian | |
| [UTMOSv2](https://github.com/sarulab-speech/UTMOSv2) (Baba et al., 2024) fine-tuned | |
| on Estonian listening-test data to predict the mean opinion score | |
| (MOS) of synthesized speech. | |
| The mean opinion score quantifies the naturalness of synthesized speech by | |
| averaging subjective ratings from multiple listeners, and is the standard measure | |
| in speech synthesis evaluation. Obtaining it requires a listening test, which is | |
| labour-intensive and time-consuming — a cost that falls hardest on languages with | |
| few speakers, where recruiting sufficient raters is itself difficult. | |
| **[Interactive demo](https://huggingface.co/spaces/monatolmats/utmosv2-estonian-demo)** | |
| ## Scope of this work | |
| The architecture and the English pretraining are the work of the UTMOSv2 authors. | |
| The contribution here is the fine-tuning on Estonian data and its evaluation on | |
| Estonian and Võro, carried out for the bachelor's thesis *Automatic Speech | |
| Synthesis Quality Assessment for Finno-Ugric Languages* (Mona Tolmats, University | |
| of Tartu, 2025; supervisor Liisa Rätsep). | |
| ## Usage | |
| UTMOSv2 conditions on a data-domain one-hot vector and cannot derive one for a | |
| corpus it has not seen. During fine-tuning the Estonian data occupied three of | |
| UTMOSv2's existing domain slots. Following the recommendation of the UTMOSv2 | |
| authors, each clip should be scored under all three and the results averaged. | |
| **Scoring under a single arbitrary domain will produce systematically different | |
| values.** | |
| ```bash | |
| pip install git+https://github.com/sarulab-speech/UTMOSv2.git | |
| huggingface-cli download monatolmats/utmosv2-estonian --local-dir weights | |
| ``` | |
| ```python | |
| import utmosv2 | |
| model = utmosv2.create_model( | |
| config="fusion_stage3", | |
| checkpoint_path="weights/utmosv2_estonian.pth", | |
| ) | |
| DOMAINS = ["somos", "blizzard2010-ES3", "blizzard2010-ES1"] | |
| runs = [model.predict(input_dir="wavs/", predict_dataset=d) for d in DOMAINS] | |
| for rows in zip(*runs): | |
| mos = sum(r["predicted_mos"] for r in rows) / len(rows) | |
| print(f"{mos:.2f} {rows[0]['file_path']}") | |
| ``` | |
| Input should be 16 kHz mono audio; output is a predicted MOS on the 1–5 scale. | |
| ## Model details | |
| | | | | |
| | --- | --- | | |
| | Architecture | UTMOSv2 `fusion_stage3`: a wav2vec 2.0 branch and four EfficientNetV2-S mel-spectrogram branches, fused with a data-domain embedding | | |
| | Initialised from | `fold0_s42_best_model.pth`, a single fold of the released 5-fold ensemble | | |
| | Optimiser | AdamW, learning rate 1e-4, cosine annealing | | |
| | Batch size | 4 | | |
| | Input length | first 10 s of each clip | | |
| | Stopping | epoch 33, on validation MSE | | |
| | Hardware | NVIDIA Tesla A100 40 GB, University of Tartu HPC cluster | | |
| ## Training data | |
| Listening-test results from evaluation campaigns run at the University of Tartu | |
| (Rätsep et al.). Fine-tuning used the three Estonian campaigns; the Võro campaign | |
| was held out entirely to assess cross-lingual transfer. | |
| | Year | Language | Systems | Clips | Ratings | Role | | |
| | --- | --- | --- | --- | --- | --- | | |
| | 2020 | Estonian | 5 | 850 | 17 000 | fine-tuning | | |
| | 2022 | Estonian | 7 | 1 400 | 5 600 | fine-tuning | | |
| | 2024 | Estonian | 9 | 2 560 | 12 800 | fine-tuning | | |
| | 2023 | Võro | 8 | 800 | 2 600 | held-out test | | |
| The 4 810 Estonian clips were split 90/10 into 4 329 training and 481 validation, | |
| stratified by MOS and by source campaign. Ratings were averaged per clip, so | |
| targets are continuous. The distribution is skewed toward mid and high scores and | |
| was left unbalanced, to reflect the composition of real listening tests. | |
| The data is not released; it belongs to the projects cited in the thesis. | |
| ## Evaluation | |
| Estonian validation set: | |
| | Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ | | |
| | --- | --- | --- | --- | --- | | |
| | wav2vec 2.0 | 0.352 | 0.669 | 0.630 | 0.464 | | |
| | SCOREQ | **0.224** | **0.802** | 0.768 | 0.589 | | |
| | UTMOSv2-Estonian | 0.230 | 0.797 | **0.784** | **0.600** | | |
| Võro test set, unseen during fine-tuning: | |
| | Model | MSE ↓ | LCC ↑ | SRCC ↑ | KTAU ↑ | | |
| | --- | --- | --- | --- | --- | | |
| | wav2vec 2.0 | 0.639 | 0.177 | 0.168 | 0.118 | | |
| | SCOREQ | 0.632 | 0.202 | 0.191 | 0.132 | | |
| | UTMOSv2-Estonian | **0.606** | **0.349** | **0.311** | **0.220** | | |
| Utterance-level correlation on Võro is considerably lower than on Estonian. At | |
| the system level the picture is better: of the three models compared, this was | |
| the only one to rank all eight Võro synthesis systems in the same order as human | |
| raters, and the only one whose 95 % confidence intervals overlapped the human | |
| intervals for every system. | |
| ## Limitations | |
| - **Intended for ranking, not certification.** Like MOS predictors generally, | |
| this model is suited to ordering systems rather than grading individual clips. | |
| Differences below 0.2 MOS on a single utterance fall within noise; averaging | |
| over at least 20 utterances is advisable when comparing systems. | |
| - **Predictions are not deterministic.** The spectrogram branch samples random | |
| windows, so repeated scoring of the same file varies slightly. | |
| - **Cross-lingual transfer is demonstrated only for Võro**, which is closely | |
| related to Estonian. No claim is made about more distant Finno-Ugric languages. | |
| - **Narrow rating range in the training data.** Very low scores are almost absent, | |
| so discrimination between good and excellent systems is weaker than between | |
| poor and good ones. | |
| - **Single fold**, not the 5-fold ensemble of the released UTMOSv2, so variance is | |
| higher than the published UTMOSv2 figures. |