RoeiG commited on
Commit
13de6cf
·
verified ·
1 Parent(s): 92a5cbe

Model card: bold the best value per row; numeric chance for the tone scales

Browse files
Files changed (1) hide show
  1. README.md +16 -11
README.md CHANGED
@@ -105,21 +105,23 @@ Question types:
105
 
106
  Every evaluation set below was held out of training. he_bench v1 has 8 Hebrew tasks, each asked in 3 phrasings; the
107
  set is frozen by sha256. Brier and ECE are better when lower. "laya-multilingual" is Laya's published multilingual
108
- checkpoint. "Previous" is this project's previous checkpoint, trained without the reading and teacher data.
 
 
109
 
110
- | Metric | laya-multilingual | Previous | **Laya-Hebrew** |
111
  |---|---|---|---|
112
- | MASSIVE he, 20 intents (500) | 0.352 | 0.816 | **0.806** |
113
  | MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
114
- | MASSIVE en, 20 intents | 0.652 | 0.816 | **0.816** |
115
- | he_bench accuracy / Brier / ECE (4,290) | 0.439 / 0.701 / 0.266 | 0.527 / 0.538 / 0.119 | **0.595 / 0.499 / 0.125** |
116
  | Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
117
- | SIB-200-he topic | | 0.808 | **0.801** |
118
  | Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
119
- | Yes/no as a claim: same gap (390) | | 0.692 | **0.674** |
120
  | Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
121
  | Rule-direction probe (208) | | 0.798 | **0.832** |
122
- | Held-out soft-label set, Brier (900) | | 0.230 | **0.231** |
123
 
124
  he_bench accuracy by task:
125
 
@@ -131,8 +133,11 @@ he_bench accuracy by task:
131
  | sentiment | 0.633 | 0.33 |
132
  | winograd | 0.607 | 0.50 |
133
  | hellaswag | 0.447 | 0.25 |
134
- | tone arousal | 0.389 | 5 levels |
135
- | tone valence | 0.328 | 5 levels |
 
 
 
136
 
137
  **Run-to-run noise.** A second training seed, with the same data and settings, differed by these amounts:
138
  - he_bench overall: 0.2 points
@@ -216,4 +221,4 @@ and runtime are by Laya's authors (Apache-2.0).
216
 
217
  - Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
218
  - Laya's authors, for the architecture, the RLCD training method and the runtime
219
- - The creators of every dataset listed above
 
105
 
106
  Every evaluation set below was held out of training. he_bench v1 has 8 Hebrew tasks, each asked in 3 phrasings; the
107
  set is frozen by sha256. Brier and ECE are better when lower. "laya-multilingual" is Laya's published multilingual
108
+ checkpoint. "Previous" is this project's previous checkpoint, trained without the reading and teacher data. **Bold**
109
+ marks the best value in each row: the highest, or the lowest for Brier and ECE. Differences smaller than the run-to-run
110
+ noise below are not meaningful.
111
 
112
+ | Metric | laya-multilingual | Previous | Laya-Hebrew |
113
  |---|---|---|---|
114
+ | MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
115
  | MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
116
+ | MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
117
+ | he_bench accuracy / Brier / ECE (4,290) | 0.439 / 0.701 / 0.266 | 0.527 / 0.538 / **0.119** | **0.595** / **0.499** / 0.125 |
118
  | Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
119
+ | SIB-200-he topic | | **0.808** | 0.801 |
120
  | Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
121
+ | Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
122
  | Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
123
  | Rule-direction probe (208) | | 0.798 | **0.832** |
124
+ | Held-out soft-label set, Brier (900) | | **0.230** | 0.231 |
125
 
126
  he_bench accuracy by task:
127
 
 
133
  | sentiment | 0.633 | 0.33 |
134
  | winograd | 0.607 | 0.50 |
135
  | hellaswag | 0.447 | 0.25 |
136
+ | tone arousal | 0.389 | 0.21 |
137
+ | tone valence | 0.328 | 0.23 |
138
+
139
+ Chance is the accuracy of a uniformly random pick. The tone tasks are 5-level scales, and their chance is slightly above
140
+ 0.20 because some items tie between two levels.
141
 
142
  **Run-to-run noise.** A second training seed, with the same data and settings, differed by these amounts:
143
  - he_bench overall: 0.2 points
 
221
 
222
  - Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
223
  - Laya's authors, for the architecture, the RLCD training method and the runtime
224
+ - The creators of every dataset listed above