Raymond1122 commited on
Commit
7775e0b
·
verified ·
1 Parent(s): b0ac64d

Laya-style model card: benchmark figures, 13-subset table w/ CI, honest limits

Browse files
.gitattributes CHANGED
@@ -34,3 +34,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ eval/figs/fig6_board_style.png filter=lfs diff=lfs merge=lfs -text
38
+ eval/figs/fig7_scatter.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -99,6 +99,16 @@ Single forward pass over the prompt, one softmax over ≤26 candidate logits.
99
 
100
  Objective and prompt format are unchanged from the official Nimble protocol; the recipe card with reproduction commands lives in the [GitHub repo](https://github.com/metask-ai/metask-jev).
101
 
 
 
 
 
 
 
 
 
 
 
102
  ## Honest limits
103
 
104
  - **summeval-relevance (26.7%)** is the one clear regression vs 9B (49.2%): a 5-level rubric with a systematic 3↔4 boundary shift. NLL and expected-score error are actually *better* than 9B — the argmax metric amplifies the boundary shift. If your use case is fine-grained relevance scoring, evaluate this subset yourself first.
 
99
 
100
  Objective and prompt format are unchanged from the official Nimble protocol; the recipe card with reproduction commands lives in the [GitHub repo](https://github.com/metask-ai/metask-jev).
101
 
102
+ ## On the JevBench board
103
+
104
+ Self-measured axes inserted into the published v1.2.7 ranking (16 official entrants + this model). Official run pending — axes here use our 231-decision protocol for Intelligence, val-fit temperature for Calibration, self-hosted 4090 for Speed/Cost.
105
+
106
+ <img src="eval/figs/fig6_board_style.png" width="660" alt="JevBench board with metask-jev-4b">
107
+
108
+ Would rank **#5** — ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash, behind djev — with the top-right quadrant of the Intelligence×Speed plane to itself among open weights:
109
+
110
+ <img src="eval/figs/fig7_scatter.png" width="660" alt="Intelligence vs Speed scatter">
111
+
112
  ## Honest limits
113
 
114
  - **summeval-relevance (26.7%)** is the one clear regression vs 9B (49.2%): a 5-level rubric with a systematic 3↔4 boundary shift. NLL and expected-score error are actually *better* than 9B — the argmax metric amplifies the boundary shift. If your use case is fine-grained relevance scoring, evaluate this subset yourself first.
eval/figs/fig6_board_style.png ADDED

Git LFS Details

  • SHA256: 1051da518c48c62758737c8b59982151172663b7a84f2f289093c5ed60c64909
  • Pointer size: 131 Bytes
  • Size of remote file: 250 kB
eval/figs/fig7_scatter.png ADDED

Git LFS Details

  • SHA256: 8f289e09c5f9201b669f14b532be641d00c1f70b840fb1383b0dfe45b3d50ade
  • Pointer size: 131 Bytes
  • Size of remote file: 119 kB