Raymond1122 commited on
Commit
94c654d
·
verified ·
1 Parent(s): 7775e0b

Laya-style model card: benchmark figures, 13-subset table w/ CI, honest limits

Browse files
Files changed (2) hide show
  1. README.md +42 -19
  2. eval/figs/fig8_laya_compare.png +0 -0
README.md CHANGED
@@ -21,6 +21,29 @@ pipeline_tag: text-classification
21
 
22
  A calibrated **typed-decision model**: give it a state (text, ticket, policy, JSON) and a typed question — `choice`, `boolean`, or rubric `score` — and it returns a probability for every option in a **single forward pass (~24 ms)**. No generation, no parsing, nothing to hallucinate.
23
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24
  | | metask-jev-4b | Bespoke Nimble-9B | Jev 1.13.0 |
25
  |---|---:|---:|---:|
26
  | 13 human-labeled subsets (3,880 items), macro | **79.6%** | 74.8% | 76.0% |
@@ -55,18 +78,28 @@ Wins: verification-style noul (civil +21.0, paws +11.2) and consistency scoring
55
 
56
  <img src="eval/figs/fig1_subsets.png" width="620" alt="13-subset comparison">
57
 
58
- ## JevBench v1.2 — public 231 decisions
59
 
60
- Scored under the official protocol (422 = wrong), at **4096-token context** — the same configuration used for every system below. The hard tier contains long policy documents: at the 9B pipeline's 2048-token limit 36 of 111 items are rejected; this model natively handles 4096 and answers 88% of them correctly. **Context length, not capability, was the bottleneck.**
61
 
62
- | tier | items | metask-jev-4b |
63
- |---|---:|---:|
64
- | judge (original) | 72 | 98.6% |
65
- | easy | 48 | 100.0% |
66
- | hard | 111 | 59.5% |
67
- | **total** | 231 | **80.1%** |
 
 
 
 
 
 
 
 
68
 
69
- <img src="eval/figs/fig2_jevbench_h2h.png" width="620" alt="JevBench head-to-head">
 
 
70
 
71
  ## Calibration
72
 
@@ -99,16 +132,6 @@ Single forward pass over the prompt, one softmax over ≤26 candidate logits.
99
 
100
  Objective and prompt format are unchanged from the official Nimble protocol; the recipe card with reproduction commands lives in the [GitHub repo](https://github.com/metask-ai/metask-jev).
101
 
102
- ## On the JevBench board
103
-
104
- Self-measured axes inserted into the published v1.2.7 ranking (16 official entrants + this model). Official run pending — axes here use our 231-decision protocol for Intelligence, val-fit temperature for Calibration, self-hosted 4090 for Speed/Cost.
105
-
106
- <img src="eval/figs/fig6_board_style.png" width="660" alt="JevBench board with metask-jev-4b">
107
-
108
- Would rank **#5** — ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash, behind djev — with the top-right quadrant of the Intelligence×Speed plane to itself among open weights:
109
-
110
- <img src="eval/figs/fig7_scatter.png" width="660" alt="Intelligence vs Speed scatter">
111
-
112
  ## Honest limits
113
 
114
  - **summeval-relevance (26.7%)** is the one clear regression vs 9B (49.2%): a 5-level rubric with a systematic 3↔4 boundary shift. NLL and expected-score error are actually *better* than 9B — the argmax metric amplifies the boundary shift. If your use case is fine-grained relevance scoring, evaluate this subset yourself first.
 
21
 
22
  A calibrated **typed-decision model**: give it a state (text, ticket, policy, JSON) and a typed question — `choice`, `boolean`, or rubric `score` — and it returns a probability for every option in a **single forward pass (~24 ms)**. No generation, no parsing, nothing to hallucinate.
23
 
24
+ ## On the JevBench board
25
+
26
+ Self-measured axes inserted into the published v1.2.7 ranking (16 official entrants + this model). Official run pending — axes here use our 231-decision protocol for Intelligence, val-fit temperature for Calibration, self-hosted 4090 for Speed/Cost.
27
+
28
+ <img src="eval/figs/fig6_board_style.png" width="660" alt="JevBench board with metask-jev-4b">
29
+
30
+ Would rank **#5** — ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash, behind djev — with the top-right quadrant of the Intelligence×Speed plane to itself among open weights:
31
+
32
+ <img src="eval/figs/fig7_scatter.png" width="660" alt="Intelligence vs Speed scatter">
33
+
34
+ **JevBench v1.2 — public 231 decisions, tier split** (422-as-wrong protocol, @4096 ctx):
35
+
36
+ | tier | items | metask-jev-4b |
37
+ |---|---:|---:|
38
+ | judge (original) | 72 | 98.6% |
39
+ | easy | 48 | 100.0% |
40
+ | hard | 111 | 59.5% |
41
+ | **total** | 231 | **80.1%** |
42
+
43
+ The hard tier contains long policy documents: at the 9B pipeline's 2048-token limit 36 of 111 items are rejected; this model natively handles 4096 and answers 88% of them correctly. **Context length, not capability, was the bottleneck.**
44
+
45
+ ## Head-to-head summary
46
+
47
  | | metask-jev-4b | Bespoke Nimble-9B | Jev 1.13.0 |
48
  |---|---:|---:|---:|
49
  | 13 human-labeled subsets (3,880 items), macro | **79.6%** | 74.8% | 76.0% |
 
78
 
79
  <img src="eval/figs/fig1_subsets.png" width="620" alt="13-subset comparison">
80
 
81
+ ## vs Laya (421M, the strongest open small-model baseline)
82
 
83
+ [Laya](https://huggingface.co/convaiinnovations/laya) trains a 25M marker head on ModernBERT-large with RLCD (pure RL, no cross-entropy) over ~30k human-labeled decisions; its typed-decisions checkpoint reports 0.766 acc / 0.062 Brier on its own 400-case suite. Different architectures, different suites — the comparison below is indicative, not apples-to-apples.
84
 
85
+ | | metask-jev-4b | laya |
86
+ |---|---|---|
87
+ | backbone | Qwen3.5-4B (decoder, LoRA merged) | ModernBERT-large (encoder + 25M head) |
88
+ | params | 4.54B | 421M |
89
+ | context | **4096** (native 32k) | 512 (root) / 1024 (typed-decisions ckpt) |
90
+ | training | SFT, candidate CE, 44.8k decisions | RLCD (proper-scoring reward), ~30k |
91
+ | raw ECE | **0.100** | 0.466 |
92
+ | ECE after temp | **0.028** | 0.081 |
93
+ | long documents (JevBench hard, ≤4096 tok) | **59.5%** | not run (512–1024 ctx) |
94
+ | high-cardinality choice (77 options) | n/a (26-option cap, same as Jev) | 0.425 without tuning |
95
+ | multilingual | en only | **100+ languages** (separate ckpt) |
96
+ | generative capability retained | yes (base LM) | no |
97
+
98
+ **Where we win**: calibration out of the box (raw ECE 0.100 is below laya's *post*-temperature 0.081; after our own temperature fit it is 0.028, ~3× lower), long-context hard items (59.5% on JevBench hard — laya's 512–1024 budget cannot run that tier), and 12/13 over Nimble-9B on human-labeled data.
99
 
100
+ **Where laya wins**: parameter efficiency (421M vs 4.5B), 100+ languages via its multilingual checkpoint, a mature packaging story (PyPI, Router, demo Space), and the RLCD training methodology is fully documented (arXiv:2510.01237).
101
+
102
+ <img src="eval/figs/fig8_laya_compare.png" width="660" alt="Laya comparison">
103
 
104
  ## Calibration
105
 
 
132
 
133
  Objective and prompt format are unchanged from the official Nimble protocol; the recipe card with reproduction commands lives in the [GitHub repo](https://github.com/metask-ai/metask-jev).
134
 
 
 
 
 
 
 
 
 
 
 
135
  ## Honest limits
136
 
137
  - **summeval-relevance (26.7%)** is the one clear regression vs 9B (49.2%): a 5-level rubric with a systematic 3↔4 boundary shift. NLL and expected-score error are actually *better* than 9B — the argmax metric amplifies the boundary shift. If your use case is fine-grained relevance scoring, evaluate this subset yourself first.
eval/figs/fig8_laya_compare.png ADDED