RoeiG commited on
Commit
599fb9a
·
verified ·
1 Parent(s): 13de6cf

Model card: mark lower-is-better metrics; split he_bench into accuracy, Brier and ECE rows

Browse files
Files changed (1) hide show
  1. README.md +4 -2
README.md CHANGED
@@ -114,14 +114,16 @@ noise below are not meaningful.
114
  | MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
115
  | MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
116
  | MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
117
- | he_bench accuracy / Brier / ECE (4,290) | 0.439 / 0.701 / 0.266 | 0.527 / 0.538 / **0.119** | **0.595** / **0.499** / 0.125 |
 
 
118
  | Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
119
  | SIB-200-he topic | | **0.808** | 0.801 |
120
  | Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
121
  | Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
122
  | Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
123
  | Rule-direction probe (208) | | 0.798 | **0.832** |
124
- | Held-out soft-label set, Brier (900) | | **0.230** | 0.231 |
125
 
126
  he_bench accuracy by task:
127
 
 
114
  | MASSIVE he, 20 intents (500) | 0.352 | **0.816** | 0.806 |
115
  | MASSIVE he, 4 intents | 0.710 | 0.928 | **0.938** |
116
  | MASSIVE en, 20 intents | 0.652 | **0.816** | **0.816** |
117
+ | he_bench accuracy (4,290) | 0.439 | 0.527 | **0.595** |
118
+ | he_bench Brier (lower is better) | 0.701 | 0.538 | **0.499** |
119
+ | he_bench ECE (lower is better) | 0.266 | **0.119** | 0.125 |
120
  | Belebele-he reading comprehension (900; chance 0.25) | | 0.468 | **0.767** |
121
  | SIB-200-he topic | | **0.808** | 0.801 |
122
  | Yes/no as a question: P(yes \| true) − P(yes \| false) | | 0.047 | **0.379** |
123
  | Yes/no as a claim: same gap (390) | | **0.692** | 0.674 |
124
  | Hebrew BoolQ, held out (875) | | 0.611 | **0.838** |
125
  | Rule-direction probe (208) | | 0.798 | **0.832** |
126
+ | Held-out soft-label set (900), Brier (lower is better) | | **0.230** | 0.231 |
127
 
128
  he_bench accuracy by task:
129