PrinceAlhassanNasamu commited on
Commit
a25be59
·
verified ·
1 Parent(s): 41f2373

model card: licence, attribution, measured numbers, limitations

Browse files
Files changed (1) hide show
  1. README.md +13 -13
README.md CHANGED
@@ -98,19 +98,19 @@ access, not by inference.
98
 
99
  ## Limitations, stated plainly
100
 
101
- - **Dagbani had no recogniser of its own for this whole project**, and the
102
- reason given for that was wrong. Every card here said "one fine-tuning
103
- session on 74 validation rows would not change that". Those 74 rows are
104
- the **eng-dag machine-translation** validation split. The Dagbani
105
- *speech* data in this same account is `waxal_dag`: **13,228 training
106
- rows, 1,750 validation rows, ~71 hours, 1,041 speakers with the largest
107
- at 1%** — more data and better speaker diversity than Ewe, which
108
- produced a working 42.19 WER recogniser. A number was carried across
109
- from a translation table into a speech claim, and then repeated on every
110
- model card on the account.
111
- It is training now, on 2026-08-31. Until it is scored, the honest
112
- statement is that Dagbani's best available recogniser scores 86.6 WER
113
- and nobody had tried fine-tuning on the data already in hand.
114
  - **Evaluation is on read and machine-translated text.** No recordings of
115
  people speaking agent commands in these languages exist. Numbers measured
116
  this way are optimistic about phrasing and pessimistic about
 
98
 
99
  ## Limitations, stated plainly
100
 
101
+ - **Dagbani did get a recogniser, and the claim that it could not was
102
+ wrong twice over.** Every card on this account used to say that "one
103
+ fine-tuning session on 74 validation rows would not change that". Those
104
+ 74 rows are the **eng-dag machine-translation** validation split; the
105
+ Dagbani *speech* data in this same account is `waxal_dag` — 13,228
106
+ training rows, 1,750 validation rows, ~71 hours, 1,041 speakers with
107
+ the largest at 1%. Trained on it, `tekyerema-asr-mms-dag` scores
108
+ **36.94 / 11.71**, against the 86.59 / 33.95 this project had believed
109
+ was the ceiling. It still loses to
110
+ `FarmerlineML/w2v-bert-2.0_2026_dagbani_ASR` at **29.20 / 9.27**, which
111
+ is what the agent actually serves. A number carried across from a
112
+ translation table into a speech claim was then repeated on every card
113
+ here until 2026-09-22.
114
  - **Evaluation is on read and machine-translated text.** No recordings of
115
  people speaking agent commands in these languages exist. Numbers measured
116
  this way are optimistic about phrasing and pessimistic about