chaoliangUNSW commited on
Commit
d53c8f8
·
verified ·
1 Parent(s): e3a56b4

Card: lead with results beyond the training data; typed decisions vs Laya typed only, with teacher-noise reference

Browse files

Typed-decisions gold comes from one teacher, so the in-domain number is now compared only with Laya's typed checkpoint (same train split) and shown next to a teacher-noise reference and a teacher-margin split. The +6.4-over-Jev and Brier-vs-Jev claims are removed; the first screen now shows Banking77, held-out MASSIVE languages, tweet_topic and JevBench.

README.md CHANGED
@@ -44,20 +44,22 @@ tags:
44
 
45
  **Jev-style decisions on your laptop.** Give it any text and a question; it returns a calibrated probability for every option in one forward pass. 0.8B parameters, open weights, Apache-2.0.
46
 
47
- ![Jev-Style 0.8B Decision v3: the whole model is 0.53 GB in 4-bit, and it is ahead of Laya typed and Jev on 2,000 typed decisions](https://huggingface.co/chaoliangUNSW/Jev-Style-0.8B-Decision-v3/resolve/main/figures/banner.png)
48
 
49
- | | **Jev-Style v3 · 0.8B** | Laya typed | Jev (API) |
50
- |---|:---:|:---:|:---:|
51
- | Typed decisions, accuracy ↑ | **79.2%** | 76.6% | 72.7%¹ |
52
- | Probability error, Brier ↓ | **0.046** | 0.061 | 0.148¹ |
53
- | Runs on your own machine | **Yes, 0.53 GB (4-bit GGUF)** | Yes | No, API only |
54
- | Longest input per call | **25,600 tokens** | 1,024 by default² | not published |
 
 
55
 
56
- <sub>¹ Same 2,000 typed decisions (LocalLLaMA/typed-decisions). v3 and Laya typed were trained on its train split; Jev is zero-shot, with numbers from the dataset card. Protocol and paired confidence intervals: see [Typed decisions](#typed-decisions-08b-beats-the-2b-models-and-jev). ² Default input budget in the Laya README: 1,024 tokens for the multilingual and typed checkpoints, 512 for English.</sub>
57
 
58
  **Reads long documents in one call.** Up to 25,600 tokens of input, 25× Laya's 1,024-token default and 25× our 2B v2's prompt. On 1,280 real 24K-token items v3 answers **98.3%** correctly, and accuracy stays flat from 1K to 24K tokens (preregistered claim, passed).
59
 
60
- **Also:** +30.3 points over the best official Laya checkpoint on model routing · ahead of Laya multilingual in 51 of 51 languages.
61
 
62
  **[Try it in your browser →](https://huggingface.co/spaces/chaoliangUNSW/jev-style-v3)**
63
 
@@ -84,7 +86,7 @@ Other builds: [GGUF for llama.cpp](https://huggingface.co/chaoliangUNSW/Jev-Styl
84
 
85
  ## What's new in v3
86
 
87
- - **Smaller and stronger.** 0.8B instead of 2B, and 79.2% vs 73.5% for our 2B v2 on the same 2,000 typed decisions.
88
  - **No letter cap.** v1 and v2 read one option-letter token, so a question could have at most 26 options. v3
89
  scores a verdict slot per option, so the options are whatever you pass. Banking77 was run with all 77 intents
90
  in one pass.
@@ -128,14 +130,70 @@ Other builds: [GGUF for llama.cpp](https://huggingface.co/chaoliangUNSW/Jev-Styl
128
 
129
  ## Results
130
 
131
- ### Typed decisions: 0.8B beats the 2B models and Jev
132
 
133
- ![Typed-decisions accuracy and Brier score: Jev-Style 0.8B v3 vs Jev, Laya typed and the 2B v1/v2](figures/headline_typed.png)
134
 
135
- At 0.8B parameters, v3 scores **79.2%** on the 2,000 typed decisions. That is **+6.4 points over Jev**, +5.7 over
136
- our 2B v2 and +2.6 over Laya's typed checkpoint, and the Brier score is **3.2× lower than Jev's** (0.046 vs 0.148).
 
 
137
 
138
- <sub>Typed-decisions test set (LocalLLaMA/typed-decisions), 2,000 decisions from 400 states. In-domain for v3 and Laya typed (both trained on its train split); zero-shot for Jev (numbers from the dataset card, measured through the Jev API on all 2,000 decisions). Laya: official typed-decisions checkpoint re-run by us on identical rows with its shipped temperature. 2B v1/v2: teacher agreement as reported on the v2 card (same 2,000 decisions, scored by that card's harness; v1 was not trained on typed decisions, v2's training pool included typed workflow decisions). Jev's accuracy is published as 0.727, so the gap is 6.40–6.50 points. v3: 1,583 / 2,000 correct, 95% CI 77.3–80.9% (Wilson); v3 minus Laya typed, paired bootstrap 95% CI +1.0 to +4.2 points.</sub>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
139
 
140
  **Head-to-head against Laya's typed checkpoint.** Both models trained on this dataset's train split, and v3 wins on all four
141
  metrics, each with a paired 95% CI that excludes zero:
@@ -149,30 +207,34 @@ metrics, each with a paired 95% CI that excludes zero:
149
 
150
  <sub>Both in-domain; v3 also trained on 27,300 synthetic typed items from other workflows. Laya: official checkpoint re-run by us on identical rows with its shipped temperature. Paired case-cluster bootstrap within suites, 2,000 resamples.</sub>
151
 
152
- <details>
153
- <summary><strong>More results:</strong> +30 points over Laya · 51 languages · calibration · 24K-token documents · JevBench · zero-shot topics · speed · 4-bit parity</summary>
154
-
155
- ### Beyond Laya: up to +30 points
156
 
157
- ![v3 vs the best official Laya checkpoint on five decision tasks](figures/beyond_laya.png)
158
 
159
- **On five decision tasks scored on identical rows, the 0.8B v3 beats the best official Laya checkpoint on every
160
- one:** +19.0 points on 77-way Banking77, +7.2 balanced accuracy on jailbreak detection, +29.5 macro-F1 on
161
- toxicity, +30.3 on model routing and +29.4 across the 37 locales held out of MASSIVE training.
 
 
162
 
163
- <sub>Laya numbers: official checkpoints (English, typed-decisions, multilingual) re-run by us on identical rows with their shipped temperatures and default token budgets; the best of the three is shown per task. v3 trained on tasks of the same kind from other datasets, never on these evaluation rows: intent (CLINC150/HWU64; Banking77 never trained), jailbreak (other permissive sets plus teacher data), toxicity (civil_comments plus teacher data; toxic-chat is evaluation-only), routing (teacher-written; the gsm8k/mbpp/AG rows are evaluation-only), MASSIVE in 14 other locales (no MASSIVE rows in these 37). n = 400 / 400 / 400 / 399 / 3,700 (37 × 100). Every gap's paired 95% bootstrap CI excludes zero.</sub>
 
 
164
 
165
- ### 51 languages, 51 wins
 
 
 
 
166
 
167
- ![Per-language MASSIVE intent accuracy, v3 vs Laya multilingual, 51 languages](figures/multilingual.png)
 
 
 
168
 
169
- **One 0.8B model, 51 languages, 51 wins over Laya.** On MASSIVE intent (20 options per question) v3 averages
170
- **71.7%** across 51 languages, against 40.1% for the official Laya multilingual checkpoint (+31.7 points). It
171
- beats Laya multilingual in every one of the 51 languages, by at least 11 points, and stays above 3× chance in all
172
- of them. That includes the 37 locales held out of MASSIVE training (65.5% vs 36.1%), 32 of them outside the 19
173
- fine-tuning languages.
174
 
175
- <sub>MASSIVE intent (mteb/amazon_massive_intent) test rows, 100 per language, 20 candidate intents per row (chance 5%, 3× chance 15%); accuracy = top-scored option. v3: in-domain for the 14 trained locales, held out for the other 37 (vi/th/el/ur had about 1.3K translated-NLI training rows each; zh-TW shares Chinese with zh-CN; a 69-row multilingual jailbreak set in training may include a few prompts in other held-out languages). Laya: official multilingual checkpoint re-run by us on identical rows with its shipped temperature and default token budget (held-out for Laya). v3 is also ahead of the best of the three official Laya checkpoints in all 51 languages (per-language point estimates on 100 rows each, smallest gap 10 points). Paired 95% CI of the 51-language macro difference: +30.1 to +33.1 points.</sub>
 
176
 
177
  ### Probabilities you can act on
178
 
@@ -202,29 +264,6 @@ at 24K), so the answers cannot be recovered from the question alone.
202
 
203
  <sub>v3 only. Laya's default input budget is 512 tokens (English) / 1,024 (multilingual, typed) per the Laya README, so Laya is not plotted. Suite long_grid_plus, English and Chinese documents: preregistered 2026-09-24 and amended before any model was scored (+96 items per 24K depth decile, thresholds unchanged); 320 items per bin, 1,280 at 24K. Controlled accuracy = the real item is correct AND its question-only and state-swap controls pass; both controls are at chance in every length bin. 25K claim rule: |24K − 2K–4K reference| ≤ 5 points and every 24K evidence-depth decile within 10 points of it.</sub>
204
 
205
- ### JevBench: ahead of Laya and every Qwen3.5-0.8B-based system
206
-
207
- ![JevBench v1.4.1 public accuracy: v3 vs Laya and the Qwen3.5-0.8B-based systems](figures/jevbench.png)
208
-
209
- **On the 231 public JevBench v1.4.1 items, v3 scores 64.1% zero-shot**: 5.6 points above Laya, and ahead of every
210
- Qwen3.5-0.8B-based system on the board, including a dedicated 0.8B decision fine-tune (+4.8 points) and
211
- SimpleJev on the same base (+9.5 points). Every answer is a valid option (231 of 231), because v3 can only score
212
- the options it is given.
213
-
214
- <sub>JevBench v1.4.1, public items only (231). v3: self-run zero-shot with the vendored official harness (commit 24b9b5c), 148 / 231 correct, 95% CI 57.7–70.0% (Wilson); training-pool contamination scan: 0 hits; not an official leaderboard entry. Other rows: public accuracy as published in the board's [v1.4.1 results file](https://github.com/fstandhartinger/jevbench). Shown: Laya plus every Qwen3.5-0.8B-based system on the board; other board systems are not shown. Laya's and M. Ghafiri's scores lie inside v3's 95% CI, so those two leads are point estimates, not significant at n = 231.</sub>
215
-
216
- ### Zero-shot topics: +12 points over English Laya
217
-
218
- ![Zero-shot tweet_topic and fin_topic accuracy: v3 vs English Laya, with Jev on tweet_topic](figures/zeroshot.png)
219
-
220
- **On two topic sets it never trained on, v3 leads English Laya by +12.3 points on tweet_topic** (75.5% vs 63.2%)
221
- **and +12.5 points on the 20-way fin_topic** (46.7% vs 34.2%). On tweet_topic it lands **within 4 points of Jev**
222
- (75.5% vs 79.3%). Macro-F1 leads over English Laya are +13.8 points (59.9% vs 46.1%) and +8.9 points (45.2% vs 36.2%).
223
- With its shipped temperature, its probabilities are also better calibrated than Jev's on both sets: ECE 0.027 vs
224
- 0.063 on tweet_topic and 0.046 vs 0.166 on fin_topic.
225
-
226
- <sub>Zero-shot for every system: neither set is in v3's training pool; accuracy over every row of the pinned test files (n = 1,693 and 4,117). Jev (1.13, API) and English Laya: numbers published by the [elcronos jev-vs-open-decision-models study](https://github.com/elcronos/jev-vs-open-decision-models) with its own prompt (results/cross_dataset_summary.json @ a1901bc), not re-run by us. v3: scored by us on the identical rows, label sets and instruction, in v3's own input format; tweet_topic accuracy 95% CI 73.4–77.5%. ECE: 15 equal-width bins as in the study; v3 with its shipped global temperature (0.880, fitted on v3's own calibration split, never on these sets), Jev's ECE as published (raw API probabilities).</sub>
227
-
228
  ### Speed: many questions, one read
229
 
230
  ![Warm p50 latency: v3 vs a Laya-architecture engine, 1K to 8K-token states](figures/latency.png)
@@ -431,7 +470,8 @@ Judge each option:
431
  [multilingual](figures/multilingual.data.json), [calibration](figures/calibration.data.json),
432
  [long_context](figures/long_context.json), [jevbench](figures/jevbench.data.json),
433
  [zeroshot](figures/zeroshot.json), [latency](figures/latency.data.json),
434
- [quantization](figures/quantization.data.json), [design_table](figures/design_table.data.json).
 
435
  - Every v3 and re-run Laya number comes from prediction files that were each scored once. Paired differences use
436
  a case-cluster bootstrap within suites (2,000 resamples). A win is only claimed when the 95% CI excludes zero,
437
  except where a chart or note says otherwise (JevBench leads over Laya and M. Ghafiri, and per-language MASSIVE
 
44
 
45
  **Jev-style decisions on your laptop.** Give it any text and a question; it returns a calibrated probability for every option in one forward pass. 0.8B parameters, open weights, Apache-2.0.
46
 
47
+ ![Jev-Style 0.8B Decision v3: the whole model is 0.53 GB in 4-bit; beyond its training data it is ahead of the best official Laya checkpoint on Banking77, 37 held-out MASSIVE languages and tweet_topic](https://huggingface.co/chaoliangUNSW/Jev-Style-0.8B-Decision-v3/resolve/main/figures/banner.png)
48
 
49
+ | Beyond its training data | **Jev-Style v3 · 0.8B** | Best official Laya |
50
+ |---|:---:|:---:|
51
+ | Banking77, 77 intents (never trained) ↑ | **68.2%** | 49.2% |
52
+ | MASSIVE intent, 37 held-out languages ↑ | **65.5%** | 36.1% |
53
+ | tweet_topic, zero-shot ↑ | **75.5%** | 63.2%¹ |
54
+ | JevBench v1.4.1, 231 public items, zero-shot ↑ | **64.1%** | 58.4%² |
55
+ | Runs on your own machine | **Yes, 0.53 GB (4-bit GGUF)** | Yes |
56
+ | Longest input per call | **25,600 tokens** | 1,024 by default³ |
57
 
58
+ <sub>Laya: the best of its three official checkpoints, re-run by us on identical rows with their shipped temperatures; paired 95% CIs exclude zero for Banking77 and MASSIVE. ¹ English Laya, as published by the elcronos study. ² Laya's score as published on the JevBench board; it lies inside v3's 95% CI, so this lead is a point estimate. ³ Default input budget in the Laya README: 1,024 tokens for the multilingual and typed checkpoints, 512 for English. Jev (API) has higher accuracy than v3 on each of these sets where its accuracy is published. Details: [Results](#results).</sub>
59
 
60
  **Reads long documents in one call.** Up to 25,600 tokens of input, 25× Laya's 1,024-token default and 25× our 2B v2's prompt. On 1,280 real 24K-token items v3 answers **98.3%** correctly, and accuracy stays flat from 1K to 24K tokens (preregistered claim, passed).
61
 
62
+ **Also:** ahead of Laya multilingual in 51 of 51 languages · +2.6 points over Laya's typed checkpoint on typed decisions, trained on the same split ([how to read that number](#reading-the-typed-number)).
63
 
64
  **[Try it in your browser →](https://huggingface.co/spaces/chaoliangUNSW/jev-style-v3)**
65
 
 
86
 
87
  ## What's new in v3
88
 
89
+ - **Smaller and stronger.** 0.8B instead of 2B, and 79.2% vs 73.5% teacher agreement for our 2B v2 on the same 2,000 typed decisions.
90
  - **No letter cap.** v1 and v2 read one option-letter token, so a question could have at most 26 options. v3
91
  scores a verdict slot per option, so the options are whatever you pass. Banking77 was run with all 77 intents
92
  in one pass.
 
130
 
131
  ## Results
132
 
133
+ <a id="beyond-its-training-data"></a>
134
 
135
+ ### Beyond its training data
136
 
137
+ Results on rows that v3 never trained on (the MASSIVE chart also shows the 14 locales it did train on; each note
138
+ states the protocol). Laya numbers are its official checkpoints re-run by us on
139
+ identical rows unless a note says otherwise. Jev (API) has higher accuracy than v3 on each of these sets where its
140
+ accuracy is published.
141
 
142
+ #### Beyond Laya: up to +30 points
143
+
144
+ ![v3 vs the best official Laya checkpoint on five decision tasks](figures/beyond_laya.png)
145
+
146
+ **On five decision tasks scored on identical rows, the 0.8B v3 beats the best official Laya checkpoint on every
147
+ one:** +19.0 points on 77-way Banking77, +7.2 balanced accuracy on jailbreak detection, +29.5 macro-F1 on
148
+ toxicity, +30.3 on model routing and +29.4 across the 37 locales held out of MASSIVE training.
149
+
150
+ <sub>Laya numbers: official checkpoints (English, typed-decisions, multilingual) re-run by us on identical rows with their shipped temperatures and default token budgets; the best of the three is shown per task. v3 trained on tasks of the same kind from other datasets, never on these evaluation rows: intent (CLINC150/HWU64; Banking77 never trained), jailbreak (other permissive sets plus teacher data), toxicity (civil_comments plus teacher data; toxic-chat is evaluation-only), routing (teacher-written; the gsm8k/mbpp/AG rows are evaluation-only), MASSIVE in 14 other locales (no MASSIVE rows in these 37). n = 400 / 400 / 400 / 399 / 3,700 (37 × 100). Every gap's paired 95% bootstrap CI excludes zero.</sub>
151
+
152
+ #### Zero-shot topics: +12 points over English Laya
153
+
154
+ ![Zero-shot tweet_topic and fin_topic accuracy: v3 vs English Laya, with Jev on tweet_topic](figures/zeroshot.png)
155
+
156
+ **On two topic sets it never trained on, v3 leads English Laya by +12.3 points on tweet_topic** (75.5% vs 63.2%)
157
+ **and +12.5 points on the 20-way fin_topic** (46.7% vs 34.2%). On tweet_topic it lands **within 4 points of Jev**
158
+ (75.5% vs 79.3%). Macro-F1 leads over English Laya are +13.8 points (59.9% vs 46.1%) and +8.9 points (45.2% vs 36.2%).
159
+ With its shipped temperature, its probabilities are also better calibrated than Jev's on both sets: ECE 0.027 vs
160
+ 0.063 on tweet_topic and 0.046 vs 0.166 on fin_topic.
161
+
162
+ <sub>Zero-shot for every system: neither set is in v3's training pool; accuracy over every row of the pinned test files (n = 1,693 and 4,117). Jev (1.13, API) and English Laya: numbers published by the [elcronos jev-vs-open-decision-models study](https://github.com/elcronos/jev-vs-open-decision-models) with its own prompt (results/cross_dataset_summary.json @ a1901bc), not re-run by us. v3: scored by us on the identical rows, label sets and instruction, in v3's own input format; tweet_topic accuracy 95% CI 73.4–77.5%. ECE: 15 equal-width bins as in the study; v3 with its shipped global temperature (0.880, fitted on v3's own calibration split, never on these sets), Jev's ECE as published (raw API probabilities).</sub>
163
+
164
+ #### JevBench: ahead of Laya and every Qwen3.5-0.8B-based system
165
+
166
+ ![JevBench v1.4.1 public accuracy: v3 vs Laya and the Qwen3.5-0.8B-based systems](figures/jevbench.png)
167
+
168
+ **On the 231 public JevBench v1.4.1 items, v3 scores 64.1% zero-shot**: 5.6 points above Laya, and ahead of every
169
+ Qwen3.5-0.8B-based system on the board, including a dedicated 0.8B decision fine-tune (+4.8 points) and
170
+ SimpleJev on the same base (+9.5 points). Every answer is a valid option (231 of 231), because v3 can only score
171
+ the options it is given.
172
+
173
+ <sub>JevBench v1.4.1, public items only (231). v3: self-run zero-shot with the vendored official harness (commit 24b9b5c), 148 / 231 correct, 95% CI 57.7–70.0% (Wilson); training-pool contamination scan: 0 hits; not an official leaderboard entry. Other rows: public accuracy as published in the board's [v1.4.1 results file](https://github.com/fstandhartinger/jevbench). Shown: Laya plus every Qwen3.5-0.8B-based system on the board; other board systems are not shown. Laya's and M. Ghafiri's scores lie inside v3's 95% CI, so those two leads are point estimates, not significant at n = 231.</sub>
174
+
175
+ #### 51 languages, 51 wins
176
+
177
+ ![Per-language MASSIVE intent accuracy, v3 vs Laya multilingual, 51 languages](figures/multilingual.png)
178
+
179
+ **One 0.8B model, 51 languages, 51 wins over Laya.** On MASSIVE intent (20 options per question) v3 averages
180
+ **71.7%** across 51 languages, against 40.1% for the official Laya multilingual checkpoint (+31.7 points). It
181
+ beats Laya multilingual in every one of the 51 languages, by at least 11 points, and stays above 3× chance in all
182
+ of them. That includes the 37 locales held out of MASSIVE training (65.5% vs 36.1%), 32 of them outside the 19
183
+ fine-tuning languages.
184
+
185
+ <sub>MASSIVE intent (mteb/amazon_massive_intent) test rows, 100 per language, 20 candidate intents per row (chance 5%, 3× chance 15%); accuracy = top-scored option. v3: in-domain for the 14 trained locales, held out for the other 37 (vi/th/el/ur had about 1.3K translated-NLI training rows each; zh-TW shares Chinese with zh-CN; a 69-row multilingual jailbreak set in training may include a few prompts in other held-out languages). Laya: official multilingual checkpoint re-run by us on identical rows with its shipped temperature and default token budget (held-out for Laya). v3 is also ahead of the best of the three official Laya checkpoints in all 51 languages (per-language point estimates on 100 rows each, smallest gap 10 points). Paired 95% CI of the 51-language macro difference: +30.1 to +33.1 points.</sub>
186
+
187
+ <a id="typed-decisions"></a>
188
+
189
+ ### Typed decisions: ahead of Laya's typed checkpoint on the same train split
190
+
191
+ ![Typed-decisions accuracy of v3 and Laya's typed checkpoint with teacher-noise reference lines, and v3 vs Laya typed by teacher margin](figures/headline_typed.png)
192
+
193
+ v3 scores **79.2%** on the 2,000 typed decisions, **+2.6 points over Laya's typed checkpoint**, which was trained on
194
+ the same split, and +5.7 over our 2B v2.
195
+
196
+ <sub>Typed-decisions test set (LocalLLaMA/typed-decisions), 2,000 decisions from 400 states. In-domain for v3 and Laya typed (both trained on its train split). Laya: official typed-decisions checkpoint re-run by us on identical rows with its shipped temperature. 2B v2: teacher agreement as reported on its card (same 2,000 decisions, scored by that card's harness; its training pool included typed workflow decisions). v3: 1,583 / 2,000 correct, 95% CI 77.3–80.9% (Wilson); v3 minus Laya typed, paired bootstrap 95% CI +1.0 to +4.2 points. Jev (zero-shot, dataset card) scores 72.7%; that is not a like-for-like comparison (see below).</sub>
197
 
198
  **Head-to-head against Laya's typed checkpoint.** Both models trained on this dataset's train split, and v3 wins on all four
199
  metrics, each with a paired 95% CI that excludes zero:
 
207
 
208
  <sub>Both in-domain; v3 also trained on 27,300 synthetic typed items from other workflows. Laya: official checkpoint re-run by us on identical rows with its shipped temperature. Paired case-cluster bootstrap within suites, 2,000 resamples.</sub>
209
 
 
 
 
 
210
 
211
+ <a id="reading-the-typed-number"></a>
212
 
213
+ **Reading the typed number.** The gold labels come from one ~4B teacher: each is the argmax of the mean of three
214
+ teacher samples. In-domain accuracy therefore measures agreement with that teacher, and the dataset card notes that
215
+ scores well above ~0.75 partly reflect the teacher's quirks rather than the task. The test split does not release the
216
+ individual samples, so the card's teacher self-agreement (73.5%, measured on its 1,600-case set) cannot be recomputed
217
+ here. What the released data does allow:
218
 
219
+ - One draw from the teacher's mean distribution matches the gold argmax **65.9%** of the time on the test split.
220
+ - Split by how sure the teacher was, v3 and Laya typed **tie on the teacher's near-ties**; v3's lead comes from cases
221
+ where the teacher is clear.
222
 
223
+ | Teacher margin (top-1 − top-2 probability) | Decisions | Laya typed | **v3** | Difference, paired 95% CI |
224
+ |---|---:|---:|---:|---|
225
+ | Below 0.1 (near-ties) | 315 | 51.7% | 51.7% | +0.0 pts [−4.4, +4.5] |
226
+ | 0.1 to 0.3 | 555 | 66.1% | **69.7%** | +3.6 pts [+0.2, +7.3] |
227
+ | 0.3 and above (teacher is clear) | 1,130 | 88.7% | **91.4%** | +2.7 pts [+0.8, +4.6] |
228
 
229
+ This rules out fitting the teacher's noise as the source of the gap to Laya typed. It cannot separate task skill from
230
+ the teacher's consistent biases, because both models trained on its labels; the [results beyond the training
231
+ data](#beyond-its-training-data) are the check for that. Jev's 72.7% on this set is zero-shot, so it is not a
232
+ like-for-like comparison with either in-domain model.
233
 
234
+ <sub>Gold labels and label_agreement flags: LocalLLaMA/typed-decisions test split. One teacher draw = mean over decisions of the gold's top probability. Paired bootstrap resampling the 400 states, 2,000 resamples per row. Values and script: [typed_teacher_noise.json](figures/typed_teacher_noise.json).</sub>
 
 
 
 
235
 
236
+ <details>
237
+ <summary><strong>More results:</strong> calibration · 24K-token documents · speed · 4-bit parity</summary>
238
 
239
  ### Probabilities you can act on
240
 
 
264
 
265
  <sub>v3 only. Laya's default input budget is 512 tokens (English) / 1,024 (multilingual, typed) per the Laya README, so Laya is not plotted. Suite long_grid_plus, English and Chinese documents: preregistered 2026-09-24 and amended before any model was scored (+96 items per 24K depth decile, thresholds unchanged); 320 items per bin, 1,280 at 24K. Controlled accuracy = the real item is correct AND its question-only and state-swap controls pass; both controls are at chance in every length bin. 25K claim rule: |24K − 2K–4K reference| ≤ 5 points and every 24K evidence-depth decile within 10 points of it.</sub>
266
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
267
  ### Speed: many questions, one read
268
 
269
  ![Warm p50 latency: v3 vs a Laya-architecture engine, 1K to 8K-token states](figures/latency.png)
 
470
  [multilingual](figures/multilingual.data.json), [calibration](figures/calibration.data.json),
471
  [long_context](figures/long_context.json), [jevbench](figures/jevbench.data.json),
472
  [zeroshot](figures/zeroshot.json), [latency](figures/latency.data.json),
473
+ [quantization](figures/quantization.data.json), [design_table](figures/design_table.data.json),
474
+ [typed_teacher_noise](figures/typed_teacher_noise.json), [banner](figures/banner.data.json).
475
  - Every v3 and re-run Laya number comes from prediction files that were each scored once. Paired differences use
476
  a case-cluster bootstrap within suites (2,000 resamples). A win is only claimed when the 95% CI excludes zero,
477
  except where a chart or note says otherwise (JevBench leads over Laya and M. Ghafiri, and per-language MASSIVE
figures/banner.data.json ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "rows": [
3
+ {
4
+ "task": "Banking77",
5
+ "subset": "77 intents",
6
+ "v3_pct": 68.2,
7
+ "laya_pct": 49.2,
8
+ "laya": "best Laya"
9
+ },
10
+ {
11
+ "task": "MASSIVE intent",
12
+ "subset": "37 held-out languages",
13
+ "v3_pct": 65.5,
14
+ "laya_pct": 36.1,
15
+ "laya": "best Laya"
16
+ },
17
+ {
18
+ "task": "tweet_topic",
19
+ "subset": "zero-shot",
20
+ "v3_pct": 75.5,
21
+ "laya_pct": 63.2,
22
+ "laya": "Laya English"
23
+ }
24
+ ],
25
+ "note": "Banking77 and tweet_topic are not in v3's training data; MASSIVE shows the 37 locales held out of its training. Laya: best official checkpoint re-run by us on identical rows; tweet_topic: English Laya as published by the elcronos study.",
26
+ "sources": [
27
+ "figures/beyond_laya.json",
28
+ "figures/zeroshot.json"
29
+ ]
30
+ }
figures/banner.png CHANGED

Git LFS Details

  • SHA256: 86aebc939a802a8c931cc85f7aab2c0b89e05a1d5cd0c6928aafd3289d396bec
  • Pointer size: 131 Bytes
  • Size of remote file: 345 kB

Git LFS Details

  • SHA256: 43d8ca0c7ce6fc140558e57a0196c8d69d6dc61ce597f9186a0a286ba8cace58
  • Pointer size: 131 Bytes
  • Size of remote file: 386 kB
figures/headline_typed.data.json CHANGED
@@ -2,30 +2,6 @@
2
  "figure": "headline_typed",
3
  "panels": {
4
  "accuracy_pct": [
5
- {
6
- "label": "Jev-Style 2B v1",
7
- "entry": "typed.teacher_agreement.v1",
8
- "raw": 0.5335,
9
- "plotted": 53.4,
10
- "source": "https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2/raw/main/README.md",
11
- "field": "card text: 'The separate typed-decisions group contains 2,000 teacher-reference decisions from 400 states. Teacher agreement is 53.35% for v1, 37.55% for English Laya and 73.45% for v2'"
12
- },
13
- {
14
- "label": "Jev",
15
- "entry": "typed.accuracy.jev",
16
- "raw": 0.727,
17
- "plotted": 72.7,
18
- "source": "docs/round2_audit/track_jev_results.json",
19
- "field": "string 'Jev typed-decisions (zero-shot)': '0.727 acc; KL 1.442; Brier 0.148; ECE 0.144; soft-acc 0.580' (from LocalLLaMA/typed-decisions dataset card / Laya HF card)"
20
- },
21
- {
22
- "label": "Jev-Style 2B v2",
23
- "entry": "typed.teacher_agreement.v2",
24
- "raw": 0.7345,
25
- "plotted": 73.5,
26
- "source": "https://huggingface.co/chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-v2/raw/main/README.md",
27
- "field": "card text: 'The separate typed-decisions group contains 2,000 teacher-reference decisions from 400 states. Teacher agreement is 53.35% for v1, 37.55% for English Laya and 73.45% for v2'"
28
- },
29
  {
30
  "label": "Laya (typed ckpt)",
31
  "entry": "typed.accuracy.laya_typed",
@@ -43,39 +19,46 @@
43
  "field": "metrics['typed.accuracy'].ours_value"
44
  }
45
  ],
46
- "brier_vs_soft": [
 
 
 
 
47
  {
48
- "label": "Jev",
49
- "entry": "typed.Brier_vs_soft_labels.jev",
50
- "raw": 0.148,
51
- "plotted": 0.148,
52
- "source": "docs/round2_audit/track_jev_results.json",
53
- "field": "string 'Jev typed-decisions (zero-shot)': '0.727 acc; KL 1.442; Brier 0.148; ECE 0.144; soft-acc 0.580' (from LocalLLaMA/typed-decisions dataset card / Laya HF card)"
 
 
54
  },
55
  {
56
- "label": "Laya (typed ckpt)",
57
- "entry": "typed.brier_vs_soft.laya_typed",
58
- "raw": 0.06146485330135357,
59
- "plotted": 0.061,
60
- "source": "runs/macjev/report_r2/scoreboard/scoreboard.json",
61
- "field": "metrics['typed.brier_vs_soft'].laya.typed.value"
 
 
62
  },
63
  {
64
- "label": "Jev-Style 0.8B v3",
65
- "entry": "typed.brier_vs_soft.v3",
66
- "raw": 0.04583493309263009,
67
- "plotted": 0.046,
68
- "source": "runs/macjev/report_r2/scoreboard/scoreboard.json",
69
- "field": "metrics['typed.brier_vs_soft'].ours_value"
 
 
70
  }
71
  ]
72
  },
73
  "annotations": {
74
- "delta_vs_jev_pts": 6.4,
75
- "delta_vs_v2_pts": 5.7,
76
  "delta_vs_laya_typed_pts": 2.6,
77
- "brier_ratio_jev_over_v3": 3.229,
78
- "brier_pct_lower_than_laya_typed": 25.4,
79
  "v3_acc_ci95_wilson": [
80
  0.7731,
81
  0.8087
@@ -83,8 +66,13 @@
83
  "v3_minus_laya_typed_paired_ci95": [
84
  0.010499999999999954,
85
  0.04249999999999998
86
- ]
 
 
 
 
 
 
87
  },
88
- "protocol": "in-domain for v3 and Laya typed; zero-shot for Jev (dataset card); 2B v1/v2 as reported on the v2 card",
89
- "footnote": "Typed-decisions test set (LocalLLaMA/typed-decisions), 2,000 decisions from 400 states. In-domain for v3 and Laya typed (both trained on its\ntrain split); zero-shot for Jev (dataset-card numbers, Jev API, all 2,000 decisions). Laya: official typed-decisions checkpoint re-run by us on\nidentical rows with its shipped temperature. 2B v1/v2: teacher agreement as reported on the v2 card (same 2,000 decisions, that card's harness;\nv1 was not trained on typed decisions, v2's pool included typed workflow decisions). v3 95% CI 77.3\u201380.9% (Wilson);\nv3 minus Laya typed, paired bootstrap 95% CI +1.0 to +4.2 pts. Plotted values: figures/headline_typed.data.json."
90
  }
 
2
  "figure": "headline_typed",
3
  "panels": {
4
  "accuracy_pct": [
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5
  {
6
  "label": "Laya (typed ckpt)",
7
  "entry": "typed.accuracy.laya_typed",
 
19
  "field": "metrics['typed.accuracy'].ours_value"
20
  }
21
  ],
22
+ "reference_lines_pct": {
23
+ "teacher_self_agreement_dataset_card": 73.5,
24
+ "one_teacher_draw_vs_gold_test_split": 65.9
25
+ },
26
+ "by_teacher_margin_pct": [
27
  {
28
+ "bin": "near-ties margin < 0.1",
29
+ "n": 315,
30
+ "v3": 51.7,
31
+ "laya_typed": 51.7,
32
+ "diff_ci95_pts": [
33
+ -4.36,
34
+ 4.52
35
+ ]
36
  },
37
  {
38
+ "bin": "0.1 \u2013 0.3",
39
+ "n": 555,
40
+ "v3": 69.7,
41
+ "laya_typed": 66.1,
42
+ "diff_ci95_pts": [
43
+ 0.18,
44
+ 7.27
45
+ ]
46
  },
47
  {
48
+ "bin": "clear margin \u2265 0.3",
49
+ "n": 1130,
50
+ "v3": 91.4,
51
+ "laya_typed": 88.7,
52
+ "diff_ci95_pts": [
53
+ 0.82,
54
+ 4.61
55
+ ]
56
  }
57
  ]
58
  },
59
  "annotations": {
 
 
60
  "delta_vs_laya_typed_pts": 2.6,
61
+ "delta_vs_v2_pts": 5.7,
 
62
  "v3_acc_ci95_wilson": [
63
  0.7731,
64
  0.8087
 
66
  "v3_minus_laya_typed_paired_ci95": [
67
  0.010499999999999954,
68
  0.04249999999999998
69
+ ],
70
+ "jev_zero_shot_pct_footnote_only": 72.7
71
+ },
72
+ "protocol": "in-domain for v3 and Laya typed; 2B v1/v2 as reported on the v2 card; Jev named in the footnote only",
73
+ "sources": {
74
+ "chart_data": "v3_card/chart_data.json",
75
+ "teacher_noise": "figures/typed_teacher_noise.json"
76
  },
77
+ "footnote": "Typed-decisions test set (LocalLLaMA/typed-decisions), 2,000 decisions from 400 states; gold = argmax of the mean of three samples from one\nteacher. In-domain for v3 and Laya typed (both trained on its train split). Laya: official typed-decisions checkpoint re-run by us on identical rows.\nOur 2B v2: 73.5% (teacher agreement, as reported on its card). Teacher self-agreement: dataset card, measured on its 1,600-case set.\nOne teacher draw: expected agreement of a draw from the mean distribution, test split. Jev (zero-shot, dataset card) scores 72.7%; not like-for-like.\nv3 95% CI 77.3\u201380.9% (Wilson); v3 minus Laya typed +1.0 to +4.2 pts; margin = top-1 minus top-2 gold probability;\nper-bin paired 95% CIs in figures/typed_teacher_noise.json. Plotted values: figures/headline_typed.data.json."
 
78
  }
figures/headline_typed.png CHANGED

Git LFS Details

  • SHA256: 9e763e4cbd86a25299785b94a89945f2c3a7cab4b1c7d5000f9456c4d220a1f4
  • Pointer size: 131 Bytes
  • Size of remote file: 176 kB

Git LFS Details

  • SHA256: 0f4a50b91df02593f177b5154f57621ff5d7450b2812f303488545d2d55255be
  • Pointer size: 131 Bytes
  • Size of remote file: 217 kB
figures/headline_typed.svg CHANGED
figures/typed_teacher_noise.json ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "what": "typed-decisions test split, 2,000 decisions / 400 states; v3 vs Laya typed (official checkpoint, re-run by us); both trained on the train split",
3
+ "gold": "argmax of the mean of three teacher samples; individual samples not released",
4
+ "single_draw_reference": 0.6589,
5
+ "dataset_card_references": {
6
+ "teacher_self_agreement": 0.735,
7
+ "factor_model": 0.704,
8
+ "majority": 0.52,
9
+ "note": "from the LocalLLaMA/typed-decisions card, measured on its 1,600-case set"
10
+ },
11
+ "argmax_agree_rate": 0.594,
12
+ "bootstrap": "paired, resampling states (clusters), 2000 resamples per subset, one generator seeded 0",
13
+ "by_margin": [
14
+ {
15
+ "subset": "all",
16
+ "n": 2000,
17
+ "v3": 0.7915,
18
+ "laya_typed": 0.766,
19
+ "diff_pts": 2.55,
20
+ "diff_ci95_pts": [
21
+ 1.05,
22
+ 4.25
23
+ ],
24
+ "mean_gold_max_prob": 0.6589
25
+ },
26
+ {
27
+ "subset": "margin < 0.1",
28
+ "n": 315,
29
+ "v3": 0.51746,
30
+ "laya_typed": 0.51746,
31
+ "diff_pts": 0.0,
32
+ "diff_ci95_pts": [
33
+ -4.36,
34
+ 4.52
35
+ ],
36
+ "mean_gold_max_prob": 0.4481
37
+ },
38
+ {
39
+ "subset": "0.1 <= margin < 0.3",
40
+ "n": 555,
41
+ "v3": 0.697297,
42
+ "laya_typed": 0.661261,
43
+ "diff_pts": 3.6,
44
+ "diff_ci95_pts": [
45
+ 0.18,
46
+ 7.27
47
+ ],
48
+ "mean_gold_max_prob": 0.5337
49
+ },
50
+ {
51
+ "subset": "margin >= 0.3",
52
+ "n": 1130,
53
+ "v3": 0.914159,
54
+ "laya_typed": 0.886726,
55
+ "diff_pts": 2.74,
56
+ "diff_ci95_pts": [
57
+ 0.82,
58
+ 4.61
59
+ ],
60
+ "mean_gold_max_prob": 0.7792
61
+ }
62
+ ],
63
+ "by_argmax_agree": [
64
+ {
65
+ "subset": "argmax_agree",
66
+ "n": 1188,
67
+ "v3": 0.906566,
68
+ "laya_typed": 0.877104,
69
+ "diff_pts": 2.95,
70
+ "diff_ci95_pts": [
71
+ 1.18,
72
+ 4.71
73
+ ],
74
+ "mean_gold_max_prob": 0.7581
75
+ },
76
+ {
77
+ "subset": "not argmax_agree",
78
+ "n": 812,
79
+ "v3": 0.623153,
80
+ "laya_typed": 0.603448,
81
+ "diff_pts": 1.97,
82
+ "diff_ci95_pts": [
83
+ -0.99,
84
+ 4.83
85
+ ],
86
+ "mean_gold_max_prob": 0.5139
87
+ }
88
+ ],
89
+ "sources": {
90
+ "v3": "runs/macjev/received/20260924-0313/evals/main/test/typed_test.jsonl",
91
+ "laya_typed": "runs/macjev/laya_baselines/typed/typed_test.jsonl"
92
+ }
93
+ }
manifest.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
  "format": "jev-style-manifest-v1",
3
  "repo": "chaoliangUNSW/Jev-Style-0.8B-Decision-v3",
4
- "created_unix": 1790264851.973563,
5
  "files": {
6
  "LICENSE": {
7
  "sha256": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
@@ -12,8 +12,12 @@
12
  "bytes": 1966
13
  },
14
  "README.md": {
15
- "sha256": "5fb2e4d23bf7844512d7c825d7fbd482533126fe93c1c19aabfa0290dd607df7",
16
- "bytes": 32989
 
 
 
 
17
  },
18
  "chat_template.jinja": {
19
  "sha256": "273d8e0e683b885071fb17e08d71e5f2a5ddfb5309756181681de4f5a1822d80",
@@ -23,9 +27,13 @@
23
  "sha256": "3f56b6210db3c52e0bafb50da005eb52b6d340ab4ad980091879b397743675fc",
24
  "bytes": 1790
25
  },
 
 
 
 
26
  "figures/banner.png": {
27
- "sha256": "86aebc939a802a8c931cc85f7aab2c0b89e05a1d5cd0c6928aafd3289d396bec",
28
- "bytes": 344736
29
  },
30
  "figures/beyond_laya.json": {
31
  "sha256": "a6dfd582fe3aca32bcd67cf01b85190d92084a0790f7744dd03180df5923f946",
@@ -64,16 +72,16 @@
64
  "bytes": 26655
65
  },
66
  "figures/headline_typed.data.json": {
67
- "sha256": "04380aa3823f80489ae37f893de0bda67ae2dae38fd07e37f88700f7c173f0ec",
68
- "bytes": 3819
69
  },
70
  "figures/headline_typed.png": {
71
- "sha256": "9e763e4cbd86a25299785b94a89945f2c3a7cab4b1c7d5000f9456c4d220a1f4",
72
- "bytes": 175884
73
  },
74
  "figures/headline_typed.svg": {
75
- "sha256": "a49ec30b28319d682d9465145bf6435ac4ec4ad11c46a28c552c9b851fb27f37",
76
- "bytes": 20359
77
  },
78
  "figures/jevbench.data.json": {
79
  "sha256": "f9173e65f5cd1fe9fadad8c93a8d00dbe5181e7314d49f18c6f547a0ed168543",
@@ -135,6 +143,10 @@
135
  "sha256": "70720bb46ec1e718dcbfac9a83d3b5d11d50b3a5cc3356915465d70d95f3d43b",
136
  "bytes": 28756
137
  },
 
 
 
 
138
  "figures/zeroshot.json": {
139
  "sha256": "003c087074b876efab789fadf94f51d134c604d5f2efce417557732de2272e9d",
140
  "bytes": 1867
 
1
  {
2
  "format": "jev-style-manifest-v1",
3
  "repo": "chaoliangUNSW/Jev-Style-0.8B-Decision-v3",
4
+ "created_unix": 1790293524.81474,
5
  "files": {
6
  "LICENSE": {
7
  "sha256": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
 
12
  "bytes": 1966
13
  },
14
  "README.md": {
15
+ "sha256": "de01bf84f05dcfe340e9bd08ade9fd904553fa5eb0db736d109fcbb39a868f77",
16
+ "bytes": 35644
17
+ },
18
+ "__pycache__/jev_style_decision.cpython-312.pyc": {
19
+ "sha256": "4e68d51a447aa64877128431a50b7fdca85f57eb7558754cb76f65858339e528",
20
+ "bytes": 35081
21
  },
22
  "chat_template.jinja": {
23
  "sha256": "273d8e0e683b885071fb17e08d71e5f2a5ddfb5309756181681de4f5a1822d80",
 
27
  "sha256": "3f56b6210db3c52e0bafb50da005eb52b6d340ab4ad980091879b397743675fc",
28
  "bytes": 1790
29
  },
30
+ "figures/banner.data.json": {
31
+ "sha256": "b9d992327c8d9db777cef1d53a49226f19d45c4ce5c16c3e284f525405c41e20",
32
+ "bytes": 728
33
+ },
34
  "figures/banner.png": {
35
+ "sha256": "43d8ca0c7ce6fc140558e57a0196c8d69d6dc61ce597f9186a0a286ba8cace58",
36
+ "bytes": 385876
37
  },
38
  "figures/beyond_laya.json": {
39
  "sha256": "a6dfd582fe3aca32bcd67cf01b85190d92084a0790f7744dd03180df5923f946",
 
72
  "bytes": 26655
73
  },
74
  "figures/headline_typed.data.json": {
75
+ "sha256": "13f18eb866d87d21057478cc40d8180b054258d733996b9b0a98fa846bcc0ae1",
76
+ "bytes": 2492
77
  },
78
  "figures/headline_typed.png": {
79
+ "sha256": "0f4a50b91df02593f177b5154f57621ff5d7450b2812f303488545d2d55255be",
80
+ "bytes": 216728
81
  },
82
  "figures/headline_typed.svg": {
83
+ "sha256": "345d8bf7fefdbdf26890ce9c30d74b3df2a4642d19768168f10d0fc1821d7f6b",
84
+ "bytes": 23730
85
  },
86
  "figures/jevbench.data.json": {
87
  "sha256": "f9173e65f5cd1fe9fadad8c93a8d00dbe5181e7314d49f18c6f547a0ed168543",
 
143
  "sha256": "70720bb46ec1e718dcbfac9a83d3b5d11d50b3a5cc3356915465d70d95f3d43b",
144
  "bytes": 28756
145
  },
146
+ "figures/typed_teacher_noise.json": {
147
+ "sha256": "08ac501f990c889d72cba4b7909de5129860159f28a95fee19ce82cc0bd36dbe",
148
+ "bytes": 2004
149
+ },
150
  "figures/zeroshot.json": {
151
  "sha256": "003c087074b876efab789fadf94f51d134c604d5f2efce417557732de2272e9d",
152
  "bytes": 1867