Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -17,21 +17,90 @@ Fine-tuned **WavLM-Large** backbone + MLP scoring head for phoneme-level English
|
|
| 17 |
|
| 18 |
Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
|
| 19 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
## Architecture
|
| 21 |
|
| 22 |
```
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
|
|
|
| 35 |
```
|
| 36 |
|
| 37 |
- **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
|
|
@@ -40,41 +109,22 @@ Audio βββ CTC model (alignment) βββ Frame-level phoneme segments
|
|
| 40 |
|
| 41 |
## Performance
|
| 42 |
|
| 43 |
-
Evaluated on test set (8062 phonemes from 1727 audio files):
|
| 44 |
|
| 45 |
| Metric | GOP Baseline (v1.0) | This Model |
|
| 46 |
|--------|-------------------|------------|
|
| 47 |
| Phoneme Error AUC-ROC | 0.738 | **0.870** |
|
| 48 |
| Phoneme Error F1 | 0.476 | **0.595** |
|
| 49 |
| Phoneme Error Precision | 0.379 | **0.592** |
|
|
|
|
| 50 |
| Phone Score Pearson | 0.372 | **0.645** |
|
| 51 |
| Phone Score MAE | 27.44 | **16.47** |
|
| 52 |
|
| 53 |
-
## Usage
|
| 54 |
-
|
| 55 |
-
```python
|
| 56 |
-
from pipeline_v2 import PronunciationAssessorV2
|
| 57 |
-
|
| 58 |
-
assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
|
| 59 |
-
result = assessor.assess("audio.mp3", "Hello, Peter.")
|
| 60 |
-
|
| 61 |
-
# result["overall_score"] β 82.3
|
| 62 |
-
# result["words"][0]["phonemes"][0]["error"] β False
|
| 63 |
-
# result["words"][0]["phonemes"][0]["pherr_prob"] β 0.05
|
| 64 |
-
```
|
| 65 |
-
|
| 66 |
-
CLI:
|
| 67 |
-
```bash
|
| 68 |
-
python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
|
| 69 |
-
```
|
| 70 |
-
|
| 71 |
-
See full pipeline code at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction)
|
| 72 |
-
|
| 73 |
## Training Data
|
| 74 |
|
| 75 |
- 11,601 audio recordings of English learners (children)
|
| 76 |
- 53,926 phonemes with professional human evaluation labels
|
| 77 |
-
- Labels include per-phoneme scores (0-100) and error flags
|
| 78 |
|
| 79 |
## Training Details
|
| 80 |
|
|
@@ -84,9 +134,16 @@ See full pipeline code at: [github.com/Jianshu-She/Voice-correction](https://git
|
|
| 84 |
- Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
|
| 85 |
- Batch size 64, trained for 24 epochs (early stopped)
|
| 86 |
- Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
|
|
|
|
| 87 |
|
| 88 |
## Files
|
| 89 |
|
| 90 |
-
|
| 91 |
-
-
|
| 92 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
|
| 19 |
|
| 20 |
+
## Quick Start
|
| 21 |
+
|
| 22 |
+
### Install
|
| 23 |
+
|
| 24 |
+
```bash
|
| 25 |
+
pip install torch torchaudio transformers g2p-en huggingface_hub
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
### Python API
|
| 29 |
+
|
| 30 |
+
```python
|
| 31 |
+
from pipeline_v2 import PronunciationAssessorV2
|
| 32 |
+
|
| 33 |
+
# Auto-download model from HuggingFace
|
| 34 |
+
assessor = PronunciationAssessorV2.from_pretrained()
|
| 35 |
+
|
| 36 |
+
result = assessor.assess("audio.mp3", "Hello, Peter.")
|
| 37 |
+
|
| 38 |
+
print(result["overall_score"]) # 85.2
|
| 39 |
+
print(result["n_errors"]) # 0
|
| 40 |
+
for word in result["words"]:
|
| 41 |
+
print(f"{word['word']}: score={word['score']}")
|
| 42 |
+
for ph in word["phonemes"]:
|
| 43 |
+
err = " <- ERROR" if ph["error"] else ""
|
| 44 |
+
print(f" /{ph['phone']}/ score={ph['score']} pherr={ph['pherr_prob']:.2f}{err}")
|
| 45 |
+
```
|
| 46 |
+
|
| 47 |
+
### CLI
|
| 48 |
+
|
| 49 |
+
```bash
|
| 50 |
+
# Model downloads automatically on first run
|
| 51 |
+
python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
|
| 52 |
+
```
|
| 53 |
+
|
| 54 |
+
Output:
|
| 55 |
+
```
|
| 56 |
+
============================================================
|
| 57 |
+
Text: "Hello, Peter."
|
| 58 |
+
Overall Score: 85.2/100 (errors: 0/8)
|
| 59 |
+
============================================================
|
| 60 |
+
|
| 61 |
+
β Hello score= 87.7 errors=0/4
|
| 62 |
+
/hh / score= 98.6 GOP= -0.97 pherr=0.05
|
| 63 |
+
/ah / score= 73.1 GOP= -7.40 pherr=0.43
|
| 64 |
+
/l / score= 88.6 GOP= +4.00 pherr=0.29
|
| 65 |
+
/ow / score= 90.7 GOP= -6.05 pherr=0.13
|
| 66 |
+
|
| 67 |
+
β Peter score= 82.6 errors=0/4
|
| 68 |
+
/p / score= 95.7 GOP= +5.40 pherr=0.08
|
| 69 |
+
/iy / score= 90.2 GOP= +3.70 pherr=0.12
|
| 70 |
+
/t / score= 72.5 GOP= +0.50 pherr=0.55
|
| 71 |
+
/er / score= 71.8 GOP= -1.40 pherr=0.61
|
| 72 |
+
```
|
| 73 |
+
|
| 74 |
+
### Download Model Manually
|
| 75 |
+
|
| 76 |
+
```bash
|
| 77 |
+
# Via huggingface-cli
|
| 78 |
+
huggingface-cli download Jianshu001/wavlm-phoneme-scorer wavlm_finetuned.pt --local-dir .
|
| 79 |
+
|
| 80 |
+
# Via Python
|
| 81 |
+
from huggingface_hub import hf_hub_download
|
| 82 |
+
hf_hub_download(repo_id="Jianshu001/wavlm-phoneme-scorer", filename="wavlm_finetuned.pt", local_dir=".")
|
| 83 |
+
|
| 84 |
+
# Then use with local path
|
| 85 |
+
assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
|
| 86 |
+
```
|
| 87 |
+
|
| 88 |
## Architecture
|
| 89 |
|
| 90 |
```
|
| 91 |
+
Reference Text βββ G2P βββ Expected phoneme sequence
|
| 92 |
+
β
|
| 93 |
+
Audio βββ CTC model βββ Viterbi Forced Alignment βββ Frame segments
|
| 94 |
+
β β
|
| 95 |
+
ββββ WavLM-Large (fine-tuned) βββ Hidden states βββ Pool per segment
|
| 96 |
+
β
|
| 97 |
+
+ phone embedding (32d)
|
| 98 |
+
+ GOP score (1d)
|
| 99 |
+
+ n_frames (1d)
|
| 100 |
+
β
|
| 101 |
+
MLP (1058 β 512 β 512 β 256)
|
| 102 |
+
βββ score_head β phoneme score (0-100)
|
| 103 |
+
βββ pherr_head β error probability (0-1)
|
| 104 |
```
|
| 105 |
|
| 106 |
- **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
|
|
|
|
| 109 |
|
| 110 |
## Performance
|
| 111 |
|
| 112 |
+
Evaluated on test set (8062 phonemes from 1727 audio files, children's speech):
|
| 113 |
|
| 114 |
| Metric | GOP Baseline (v1.0) | This Model |
|
| 115 |
|--------|-------------------|------------|
|
| 116 |
| Phoneme Error AUC-ROC | 0.738 | **0.870** |
|
| 117 |
| Phoneme Error F1 | 0.476 | **0.595** |
|
| 118 |
| Phoneme Error Precision | 0.379 | **0.592** |
|
| 119 |
+
| Phoneme Error Recall | 0.638 | **0.598** |
|
| 120 |
| Phone Score Pearson | 0.372 | **0.645** |
|
| 121 |
| Phone Score MAE | 27.44 | **16.47** |
|
| 122 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 123 |
## Training Data
|
| 124 |
|
| 125 |
- 11,601 audio recordings of English learners (children)
|
| 126 |
- 53,926 phonemes with professional human evaluation labels
|
| 127 |
+
- Labels include per-phoneme scores (0-100) and error flags (pherr 0/1)
|
| 128 |
|
| 129 |
## Training Details
|
| 130 |
|
|
|
|
| 134 |
- Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
|
| 135 |
- Batch size 64, trained for 24 epochs (early stopped)
|
| 136 |
- Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
|
| 137 |
+
- Train/Val/Test split by audio file: 40K/5K/8K phonemes
|
| 138 |
|
| 139 |
## Files
|
| 140 |
|
| 141 |
+
| File | Description | Size |
|
| 142 |
+
|------|-------------|------|
|
| 143 |
+
| `wavlm_finetuned.pt` | Full checkpoint (backbone + head state dict) | 1.2GB |
|
| 144 |
+
| `pipeline_v2.py` | Inference pipeline with `from_pretrained()` support | 18KB |
|
| 145 |
+
| `finetune_wavlm.py` | Training script (reproducing the fine-tuning) | 25KB |
|
| 146 |
+
|
| 147 |
+
## Full Repository
|
| 148 |
+
|
| 149 |
+
See the complete project (data, evaluation, all experiments) at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction) (branch: `feature/wavlm-pipeline`)
|