Model card: mark lower-is-better metrics; split he_bench into accuracy, Brier and ECE rows
Browse files
README.md
CHANGED
|
@@ -114,14 +114,16 @@ noise below are not meaningful.
|
|
| 114 |
| MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
|
| 115 |
| MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
|
| 116 |
| MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
|
| 117 |
-
| he_bench accuracy
|
|
|
|
|
|
|
| 118 |
| Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
|
| 119 |
| SIB-200-he topic | | **0.808** | 0.801 |
|
| 120 |
| Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
|
| 121 |
| Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
|
| 122 |
| Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
|
| 123 |
| Rule-direction probe (208) | | 0.798 | **0.832** |
|
| 124 |
-
| Held-out soft-label set, Brier (
|
| 125 |
|
| 126 |
he_bench accuracy by task:
|
| 127 |
|
|
|
|
| 114 |
| MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
|
| 115 |
| MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
|
| 116 |
| MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
|
| 117 |
+
| he_bench accuracy (4,290) | 0.439 | 0.527 | **0.595** |
|
| 118 |
+
| he_bench Brier (lower is better) | 0.701 | 0.538 | **0.499** |
|
| 119 |
+
| he_bench ECE (lower is better) | 0.266 | **0.119** | 0.125 |
|
| 120 |
| Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
|
| 121 |
| SIB-200-he topic | | **0.808** | 0.801 |
|
| 122 |
| Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
|
| 123 |
| Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
|
| 124 |
| Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
|
| 125 |
| Rule-direction probe (208) | | 0.798 | **0.832** |
|
| 126 |
+
| Held-out soft-label set (900), Brier (lower is better) | | **0.230** | 0.231 |
|
| 127 |
|
| 128 |
he_bench accuracy by task:
|
| 129 |
|