Jianshu001 commited on
Commit
643bd25
Β·
verified Β·
1 Parent(s): bdae6d5

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +94 -37
README.md CHANGED
@@ -17,21 +17,90 @@ Fine-tuned **WavLM-Large** backbone + MLP scoring head for phoneme-level English
17
 
18
  Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
19
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
  ## Architecture
21
 
22
  ```
23
- Audio ──→ CTC model (alignment) ──→ Frame-level phoneme segments
24
- β”‚
25
- └──→ WavLM-Large (fine-tuned top 6 layers) ──→ Hidden states per segment
26
- β”‚
27
- + phone embedding (32d)
28
- + GOP score (1d)
29
- + n_frames (1d)
30
- β”‚
31
- MLP (1058 β†’ 512 β†’ 512 β†’ 256)
32
- β”‚
33
- β”œβ”€β”€ score_head β†’ phoneme score (0-100)
34
- └── pherr_head β†’ error probability (0-1)
 
35
  ```
36
 
37
  - **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
@@ -40,41 +109,22 @@ Audio ──→ CTC model (alignment) ──→ Frame-level phoneme segments
40
 
41
  ## Performance
42
 
43
- Evaluated on test set (8062 phonemes from 1727 audio files):
44
 
45
  | Metric | GOP Baseline (v1.0) | This Model |
46
  |--------|-------------------|------------|
47
  | Phoneme Error AUC-ROC | 0.738 | **0.870** |
48
  | Phoneme Error F1 | 0.476 | **0.595** |
49
  | Phoneme Error Precision | 0.379 | **0.592** |
 
50
  | Phone Score Pearson | 0.372 | **0.645** |
51
  | Phone Score MAE | 27.44 | **16.47** |
52
 
53
- ## Usage
54
-
55
- ```python
56
- from pipeline_v2 import PronunciationAssessorV2
57
-
58
- assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
59
- result = assessor.assess("audio.mp3", "Hello, Peter.")
60
-
61
- # result["overall_score"] β†’ 82.3
62
- # result["words"][0]["phonemes"][0]["error"] β†’ False
63
- # result["words"][0]["phonemes"][0]["pherr_prob"] β†’ 0.05
64
- ```
65
-
66
- CLI:
67
- ```bash
68
- python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
69
- ```
70
-
71
- See full pipeline code at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction)
72
-
73
  ## Training Data
74
 
75
  - 11,601 audio recordings of English learners (children)
76
  - 53,926 phonemes with professional human evaluation labels
77
- - Labels include per-phoneme scores (0-100) and error flags
78
 
79
  ## Training Details
80
 
@@ -84,9 +134,16 @@ See full pipeline code at: [github.com/Jianshu-She/Voice-correction](https://git
84
  - Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
85
  - Batch size 64, trained for 24 epochs (early stopped)
86
  - Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
 
87
 
88
  ## Files
89
 
90
- - `wavlm_finetuned.pt` β€” Full checkpoint (backbone + head state dict, 1.2GB)
91
- - `pipeline_v2.py` β€” Inference pipeline
92
- - `finetune_wavlm.py` β€” Training script
 
 
 
 
 
 
 
17
 
18
  Given a reference text and an audio recording, this model identifies which phonemes are pronounced correctly and which have errors.
19
 
20
+ ## Quick Start
21
+
22
+ ### Install
23
+
24
+ ```bash
25
+ pip install torch torchaudio transformers g2p-en huggingface_hub
26
+ ```
27
+
28
+ ### Python API
29
+
30
+ ```python
31
+ from pipeline_v2 import PronunciationAssessorV2
32
+
33
+ # Auto-download model from HuggingFace
34
+ assessor = PronunciationAssessorV2.from_pretrained()
35
+
36
+ result = assessor.assess("audio.mp3", "Hello, Peter.")
37
+
38
+ print(result["overall_score"]) # 85.2
39
+ print(result["n_errors"]) # 0
40
+ for word in result["words"]:
41
+ print(f"{word['word']}: score={word['score']}")
42
+ for ph in word["phonemes"]:
43
+ err = " <- ERROR" if ph["error"] else ""
44
+ print(f" /{ph['phone']}/ score={ph['score']} pherr={ph['pherr_prob']:.2f}{err}")
45
+ ```
46
+
47
+ ### CLI
48
+
49
+ ```bash
50
+ # Model downloads automatically on first run
51
+ python pipeline_v2.py --audio audio.mp3 --text "Hello, Peter."
52
+ ```
53
+
54
+ Output:
55
+ ```
56
+ ============================================================
57
+ Text: "Hello, Peter."
58
+ Overall Score: 85.2/100 (errors: 0/8)
59
+ ============================================================
60
+
61
+ βœ“ Hello score= 87.7 errors=0/4
62
+ /hh / score= 98.6 GOP= -0.97 pherr=0.05
63
+ /ah / score= 73.1 GOP= -7.40 pherr=0.43
64
+ /l / score= 88.6 GOP= +4.00 pherr=0.29
65
+ /ow / score= 90.7 GOP= -6.05 pherr=0.13
66
+
67
+ βœ“ Peter score= 82.6 errors=0/4
68
+ /p / score= 95.7 GOP= +5.40 pherr=0.08
69
+ /iy / score= 90.2 GOP= +3.70 pherr=0.12
70
+ /t / score= 72.5 GOP= +0.50 pherr=0.55
71
+ /er / score= 71.8 GOP= -1.40 pherr=0.61
72
+ ```
73
+
74
+ ### Download Model Manually
75
+
76
+ ```bash
77
+ # Via huggingface-cli
78
+ huggingface-cli download Jianshu001/wavlm-phoneme-scorer wavlm_finetuned.pt --local-dir .
79
+
80
+ # Via Python
81
+ from huggingface_hub import hf_hub_download
82
+ hf_hub_download(repo_id="Jianshu001/wavlm-phoneme-scorer", filename="wavlm_finetuned.pt", local_dir=".")
83
+
84
+ # Then use with local path
85
+ assessor = PronunciationAssessorV2(checkpoint_path="wavlm_finetuned.pt")
86
+ ```
87
+
88
  ## Architecture
89
 
90
  ```
91
+ Reference Text ──→ G2P ──→ Expected phoneme sequence
92
+ β”‚
93
+ Audio ──→ CTC model ──→ Viterbi Forced Alignment ──→ Frame segments
94
+ β”‚ β”‚
95
+ └──→ WavLM-Large (fine-tuned) ──→ Hidden states ──→ Pool per segment
96
+ β”‚
97
+ + phone embedding (32d)
98
+ + GOP score (1d)
99
+ + n_frames (1d)
100
+ β”‚
101
+ MLP (1058 β†’ 512 β†’ 512 β†’ 256)
102
+ β”œβ”€β”€ score_head β†’ phoneme score (0-100)
103
+ └── pherr_head β†’ error probability (0-1)
104
  ```
105
 
106
  - **Backbone**: WavLM-Large with top 6 transformer layers fine-tuned (76.6M trainable / 316.4M total)
 
109
 
110
  ## Performance
111
 
112
+ Evaluated on test set (8062 phonemes from 1727 audio files, children's speech):
113
 
114
  | Metric | GOP Baseline (v1.0) | This Model |
115
  |--------|-------------------|------------|
116
  | Phoneme Error AUC-ROC | 0.738 | **0.870** |
117
  | Phoneme Error F1 | 0.476 | **0.595** |
118
  | Phoneme Error Precision | 0.379 | **0.592** |
119
+ | Phoneme Error Recall | 0.638 | **0.598** |
120
  | Phone Score Pearson | 0.372 | **0.645** |
121
  | Phone Score MAE | 27.44 | **16.47** |
122
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
123
  ## Training Data
124
 
125
  - 11,601 audio recordings of English learners (children)
126
  - 53,926 phonemes with professional human evaluation labels
127
+ - Labels include per-phoneme scores (0-100) and error flags (pherr 0/1)
128
 
129
  ## Training Details
130
 
 
134
  - Gradient accumulation (4 steps), gradient clipping (max_norm=1.0)
135
  - Batch size 64, trained for 24 epochs (early stopped)
136
  - Multi-task loss: MSE/100 (score) + BCE with pos_weight=5.2 (error)
137
+ - Train/Val/Test split by audio file: 40K/5K/8K phonemes
138
 
139
  ## Files
140
 
141
+ | File | Description | Size |
142
+ |------|-------------|------|
143
+ | `wavlm_finetuned.pt` | Full checkpoint (backbone + head state dict) | 1.2GB |
144
+ | `pipeline_v2.py` | Inference pipeline with `from_pretrained()` support | 18KB |
145
+ | `finetune_wavlm.py` | Training script (reproducing the fine-tuning) | 25KB |
146
+
147
+ ## Full Repository
148
+
149
+ See the complete project (data, evaluation, all experiments) at: [github.com/Jianshu-She/Voice-correction](https://github.com/Jianshu-She/Voice-correction) (branch: `feature/wavlm-pipeline`)