Model card: bold the best value per row; numeric chance for the tone scales
Browse files
README.md
CHANGED
|
@@ -105,21 +105,23 @@ Question types:
|
|
| 105 |
|
| 106 |
Every evaluation set below was held out of training. he_bench v1 has 8 Hebrew tasks, each asked in 3 phrasings; the
|
| 107 |
set is frozen by sha256. Brier and ECE are better when lower. "laya-multilingual" is Laya's published multilingual
|
| 108 |
-
checkpoint. "Previous" is this project's previous checkpoint, trained without the reading and teacher data.
|
|
|
|
|
|
|
| 109 |
|
| 110 |
-
| Metric | laya-multilingual | Previous |
|
| 111 |
|---|---|---|---|
|
| 112 |
-
| MASSIVE he, 20 intents (500) | 0.352 | 0.816 |
|
| 113 |
| MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
|
| 114 |
-
| MASSIVE en, 20 intents | 0.652 | 0.816 | **0.816** |
|
| 115 |
-
| he_bench accuracy / Brier / ECE (4,290) | 0.439 / 0.701 / 0.266 | 0.527 / 0.538 / 0.119 | **0.595 / 0.499 / 0.125
|
| 116 |
| Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
|
| 117 |
-
| SIB-200-he topic | | 0.808 |
|
| 118 |
| Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
|
| 119 |
-
| Yes/no as a claim: same gap (390) | | 0.692 |
|
| 120 |
| Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
|
| 121 |
| Rule-direction probe (208) | | 0.798 | **0.832** |
|
| 122 |
-
| Held-out soft-label set, Brier (900) | | 0.230 |
|
| 123 |
|
| 124 |
he_bench accuracy by task:
|
| 125 |
|
|
@@ -131,8 +133,11 @@ he_bench accuracy by task:
|
|
| 131 |
| sentiment | 0.633 | 0.33 |
|
| 132 |
| winograd | 0.607 | 0.50 |
|
| 133 |
| hellaswag | 0.447 | 0.25 |
|
| 134 |
-
| tone arousal | 0.389 |
|
| 135 |
-
| tone valence | 0.328 |
|
|
|
|
|
|
|
|
|
|
| 136 |
|
| 137 |
**Run-to-run noise.** A second training seed, with the same data and settings, differed by these amounts:
|
| 138 |
- he_bench overall: 0.2 points
|
|
@@ -216,4 +221,4 @@ and runtime are by Laya's authors (Apache-2.0).
|
|
| 216 |
|
| 217 |
- Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
|
| 218 |
- Laya's authors, for the architecture, the RLCD training method and the runtime
|
| 219 |
-
- The creators of every dataset listed above
|
|
|
|
| 105 |
|
| 106 |
Every evaluation set below was held out of training. he_bench v1 has 8 Hebrew tasks, each asked in 3 phrasings; the
|
| 107 |
set is frozen by sha256. Brier and ECE are better when lower. "laya-multilingual" is Laya's published multilingual
|
| 108 |
+
checkpoint. "Previous" is this project's previous checkpoint, trained without the reading and teacher data. **Bold**
|
| 109 |
+
marks the best value in each row: the highest, or the lowest for Brier and ECE. Differences smaller than the run-to-run
|
| 110 |
+
noise below are not meaningful.
|
| 111 |
|
| 112 |
+
| Metric | laya-multilingual | Previous | Laya-Hebrew |
|
| 113 |
|---|---|---|---|
|
| 114 |
+
| MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
|
| 115 |
| MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
|
| 116 |
+
| MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
|
| 117 |
+
| he_bench accuracy / Brier / ECE (4,290) | 0.439 / 0.701 / 0.266 | 0.527 / 0.538 / **0.119** | **0.595** / **0.499** / 0.125 |
|
| 118 |
| Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
|
| 119 |
+
| SIB-200-he topic | | **0.808** | 0.801 |
|
| 120 |
| Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
|
| 121 |
+
| Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
|
| 122 |
| Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
|
| 123 |
| Rule-direction probe (208) | | 0.798 | **0.832** |
|
| 124 |
+
| Held-out soft-label set, Brier (900) | | **0.230** | 0.231 |
|
| 125 |
|
| 126 |
he_bench accuracy by task:
|
| 127 |
|
|
|
|
| 133 |
| sentiment | 0.633 | 0.33 |
|
| 134 |
| winograd | 0.607 | 0.50 |
|
| 135 |
| hellaswag | 0.447 | 0.25 |
|
| 136 |
+
| tone arousal | 0.389 | 0.21 |
|
| 137 |
+
| tone valence | 0.328 | 0.23 |
|
| 138 |
+
|
| 139 |
+
Chance is the accuracy of a uniformly random pick. The tone tasks are 5-level scales, and their chance is slightly above
|
| 140 |
+
0.20 because some items tie between two levels.
|
| 141 |
|
| 142 |
**Run-to-run noise.** A second training seed, with the same data and settings, differed by these amounts:
|
| 143 |
- he_bench overall: 0.2 points
|
|
|
|
| 221 |
|
| 222 |
- Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
|
| 223 |
- Laya's authors, for the architecture, the RLCD training method and the runtime
|
| 224 |
+
- The creators of every dataset listed above
|