SpeechT5 fine-tuned on MINDS-14 en-US

This model is a fine-tuned version of microsoft/speecht5_tts for the Hugging Face Audio Course Unit 6 hands-on assignment.

Training data

Dataset: PolyAI/minds14, English (en-US).

A small subset was used for fine-tuning on a free Google Colab Tesla T4 GPU.

Speaker embeddings were extracted with speechbrain/spkrec-xvect-voxceleb.

Training

  • Base model: microsoft/speecht5_tts
  • Task: text-to-speech
  • Training steps: 100
  • Learning rate: 1e-5
  • FP16 training
  • Gradient checkpointing enabled during training

Evaluation

Final validation loss: 1.000769

The Hugging Face Audio Course Unit 6 assignment does not specify a minimum evaluation metric. Its objective is to practice SpeechT5 fine-tuning and publish the resulting TTS model to the Hub.

Audio Course

Hugging Face Audio Course — Unit 6: From Text to Speech.

Downloads last month
11
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for herurg/speecht5-minds14-en-us-audio-course

Finetuned
(1387)
this model

Dataset used to train herurg/speecht5-minds14-en-us-audio-course