PolyAI/minds14
Viewer • Updated • 16.3k • 14.4k • 107
This model is a fine-tuned version of
microsoft/speecht5_tts for the Hugging Face Audio Course
Unit 6 hands-on assignment.
Dataset: PolyAI/minds14, English (en-US).
A small subset was used for fine-tuning on a free Google Colab Tesla T4 GPU.
Speaker embeddings were extracted with
speechbrain/spkrec-xvect-voxceleb.
microsoft/speecht5_ttstext-to-speechFinal validation loss: 1.000769
The Hugging Face Audio Course Unit 6 assignment does not specify a minimum evaluation metric. Its objective is to practice SpeechT5 fine-tuning and publish the resulting TTS model to the Hub.
Hugging Face Audio Course — Unit 6: From Text to Speech.
Base model
microsoft/speecht5_tts