kirp commited on
Commit
0ac3499
·
verified ·
1 Parent(s): 13f4884

Card: v11 results, NeoHorse head-to-head (JevBench / Kev / OpenJev), image rows; GIFs re-recorded with v11

Browse files
Files changed (4) hide show
  1. README.md +79 -23
  2. assets/2048.gif +2 -2
  3. assets/snake.gif +2 -2
  4. assets/tetris.gif +2 -2
README.md CHANGED
@@ -40,7 +40,7 @@ its official 0.853.
40
 
41
  | System | Params | Public accuracy (231) | Sealed accuracy (308) | v1.4 score |
42
  |---|---:|---:|---:|---:|
43
- | **JPT-4B** | 4B | **0.840** | pending | pending |
44
  | Jev 1.13.0 (TypeSafe AI, API) | closed | 0.866 | 0.367 | 63.3 |
45
  | JevK5 v0.2.0 | 27B | 0.853 | 0.331 | 62.0 |
46
  | Winnow-12B Q8 | 12B | 0.857 | 0.331 | 55.6 |
@@ -58,22 +58,74 @@ Other systems' rows are copied from `results/v1.4/jevbench-v1.4-results.json` at
58
  ¹ Not in the v1.4 results. The number is its own card's report on the 231 public items with the benchmark's
59
  harness (hard 0.622, standard 0.986, easy 1.000).
60
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
61
  ### Other benchmarks
62
 
63
  | Benchmark | What it tests | JPT-4B | Qwen3.5-4B (same prompt, zero-shot) |
64
  |---|---|---:|---:|
65
- | JevBench public hard tier (111) | hardest general decisions | **0.703** (ECE 0.076, Brier 0.389) | 0.595 |
66
- | Typed decisions test (2,000) | in-distribution typed decisions | **0.792** (ECE 0.157, Brier 0.335) | 0.596 (ECE 0.171, Brier 0.559) |
67
- | ANLI r1 / r3 | adversarial natural-language inference | **0.683 / 0.617** | 0.660 / 0.513 |
68
- | Banking77 / MASSIVE (en / de / zh) | intent classification, incl. multilingual | **0.807 / 0.867 / 0.830 / 0.830** | 0.663 / 0.733 / 0.670 / 0.703 |
69
- | EnvBench v0.1 public / held-out (skill, 0–100) | sequential decisions in game/puzzle envs (2048, Snake, chess, …) | **45.4 / 46.7** | — |
70
- | ScreenSpot-v2 set-of-marks (300, images) | GUI element grounding from a screenshot | **0.923** (ECE 0.025) | 0.903 (ECE 0.058) |
71
- | Screen2Words match (300, images) | screenshot summarization | **0.923** (ECE 0.028) | 0.887 (ECE 0.030) |
72
- | ERQA (400, images) | embodied/robotics visual reasoning | 0.435 (ECE 0.178) | **0.463** (ECE 0.100) |
73
-
74
- Numbers are accuracy, with ECE (10 bins) and Brier where shown, at the fitted temperature T = 1.087. The
75
  banking/MASSIVE/typed rows are in-distribution: their train splits are in the training mix, their test items are not.
76
- ERQA is the one set where fine-tuning hurt (−2.8 points, more overconfident); no image or robotics data was trained on.
77
 
78
  JevBench public hard tier by family (n):
79
 
@@ -81,12 +133,15 @@ JevBench public hard tier by family (n):
81
  |---|---:|---:|
82
  | long_policy (19) | **0.74** | 0.47 |
83
  | judge_hard (17) | **0.76** | 0.71 |
84
- | multi_hop (18) | **0.67** | 0.56 |
85
- | temporal_numeric (15) | 0.27 | **0.33** |
86
- | probability (10) | **0.60** | 0.50 |
87
  | trap / adversarial / routing_hard (19) | **1.00** | 0.95 |
88
  | ambiguous (7) | **0.86** | 0.71 |
89
- | tradeoff (6) | **0.67** | 0.33 |
 
 
 
90
 
91
  ## Quick start
92
 
@@ -96,18 +151,18 @@ the running engine. Two things any real model server is — an engine, and a cli
96
  ```bash
97
  python -m sglang.launch_server --model-path kirp/jpt-4b --port 30000 \
98
  --context-length 32768 --mamba-scheduler-strategy extra_buffer & # Qwen3.5's DeltaNet layers need this flag
99
- llm2jev --model kirp/jpt-4b --backend sglang --url http://127.0.0.1:30000 --port 8080 --temperature 1.087
100
  ```
101
 
102
  vLLM instead: `vllm serve kirp/jpt-4b --max-logprobs 256 --return-tokens-as-token-ids --enable-scale-out --port 8000`
103
- then `llm2jev --model kirp/jpt-4b --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.087` — the three vLLM flags
104
  are required, not optional: without them every request comes back a plain HTTP 400 with no hint why.
105
 
106
  No GPU / just trying it out, no engine, no clone — llm2jev runs the model itself:
107
 
108
  ```bash
109
  pip install "llm2jev[hf,vision]"
110
- llm2jev --model kirp/jpt-4b --backend hf --port 8080 --temperature 1.087
111
  ```
112
 
113
  This path serializes requests (one forward at a time in-process — concurrent calls are safe, just not parallel);
@@ -115,7 +170,7 @@ This path serializes requests (one forward at a time in-process — concurrent c
115
  use — `hf` is for a quick check, not for traffic.
116
 
117
  Tested with SGLang 0.5.9. Its torch 2.9.1 pins cuDNN 9.10, which SGLang refuses to run on, so install cuDNN 9.15+ over
118
- it: `pip install "sglang==0.5.9" && pip install "nvidia-cudnn-cu12>=9.15"`. `--temperature 1.087` is the temperature
119
  fitted on the calibration split; leaving it at 1.0 changes calibration slightly, never the ranking.
120
 
121
  ```python
@@ -146,7 +201,7 @@ questions = {"click": {"type": "choice", "instructions": "Which box should be cl
146
  - **Method:** LoRA (r=16) on every attention, DeltaNet and MLP projection of the language model, merged into full
147
  weights. The vision tower is untouched. The loss is the multi-class Brier score over the option labels, on
148
  llm2jev's chat prompt with thinking disabled: the same prompt the model is served with.
149
- - **Data** (48,010 questions in 32,427 records, all converted to typed decisions; one epoch over two option-shuffled copies):
150
  - public classification, NLI, QA, preference and safety datasets;
151
  - long legal and contract documents (ContractNLI, MAUD, LegalBench, ConditionalQA, ShARC);
152
  - table and numeric reasoning (TAT-QA, MultiHiertt);
@@ -155,8 +210,9 @@ questions = {"click": {"type": "choice", "instructions": "Which box should be cl
155
  - community typed-decision sets;
156
  - oracle-labelled rollouts from 20 small game and puzzle environments;
157
  - programmatically generated rule-arithmetic items (dates, time zones, day counts, caps; labels computed by code);
158
- - 991 long policy / contract / regulation documents (11–17k characters, 3,452 questions) with amendments,
159
- exceptions and precedence rules, written by an LLM (GPT-6 Luna) with no JevBench item shown to it.
 
160
  - **Held out:** no item from JevBench, EnvBench's held-out seeds, the Decision Index frozen suite or our typed test
161
  split was used in training. Documents were checked for 8-gram overlap with JevBench.
162
 
 
40
 
41
  | System | Params | Public accuracy (231) | Sealed accuracy (308) | v1.4 score |
42
  |---|---:|---:|---:|---:|
43
+ | **JPT-4B** | 4B | **0.879** | pending | pending |
44
  | Jev 1.13.0 (TypeSafe AI, API) | closed | 0.866 | 0.367 | 63.3 |
45
  | JevK5 v0.2.0 | 27B | 0.853 | 0.331 | 62.0 |
46
  | Winnow-12B Q8 | 12B | 0.857 | 0.331 | 55.6 |
 
58
  ¹ Not in the v1.4 results. The number is its own card's report on the 231 public items with the benchmark's
59
  harness (hard 0.622, standard 0.986, easy 1.000).
60
 
61
+ ### Head-to-head with the NeoHorse-Jev-4B comparison
62
+
63
+ The same items, selection rules and metrics as the text table on the
64
+ [NeoHorse-Jev-4B card](https://huggingface.co/TokenRhythm/NeoHorse-Jev-4B) (results dated 2026-09-24). Only JPT-4B was
65
+ run by us; every other number is copied from that card.
66
+
67
+ | Model | Params | JevBench (family-macro) | Kev (clean acc.) | OpenJev text, 18 tasks | Mean of the three |
68
+ |---|---:|---:|---:|---:|---:|
69
+ | **JPT-4B** | 4B | **88.84** | 79.69 | **69.20** | **79.25** |
70
+ | NeoHorse-Jev-4B | 4B | 75.73 | **81.92** | 56.75 | 71.47 |
71
+ | Open-Jev-9B | 9B | 77.13 | 77.87 | 63.75 | 72.92 |
72
+ | Kev-4B | 4B | 73.71 | 81.47 | 52.82 | 69.33 |
73
+ | Laya English | — | 55.82 | 61.30 | 37.24 | 51.45 |
74
+
75
+ <details>
76
+ <summary>What differs from their table, and per-task numbers</summary>
77
+
78
+ - **OpenJev is 18 tasks, not 19.** GSM8K `nli_rerank@4` needs NeoHorse's frozen candidate pool, which is not
79
+ published, so it is left out. The other models' 18-task means are recomputed here from the per-task numbers on their
80
+ card; they are not on the card.
81
+ - **Nimble, VitaminC and MASSIVE are not run.** Their frozen subset IDs are not published, so we cannot guarantee the
82
+ same items. For the same reason there is no six-group AVG for JPT-4B.
83
+ - **GPQA.** OpenJev lists the correct answer first. That doesn't matter for per-option scorers, but JPT-4B reads all
84
+ four options at once, so we shuffle each item's options with a fixed seed. This can only lower our score.
85
+ - **Kev.** 177 transcripts whose turns use the role `customer` are sent as `{"conversation": [...]}`, because Qwen's
86
+ chat template rejects that role.
87
+ - Kev comes from jaredpalmer/kev at `30c619b`. OpenJev's task loaders are imported unchanged from
88
+ AlexWortega/openjev at `552759d`.
89
+
90
+ | OpenJev task | JPT-4B | NeoHorse-Jev-4B | Open-Jev-9B | Kev-4B |
91
+ |---|---:|---:|---:|---:|
92
+ | scitail | 79.21 | **87.02** | 79.16 | 84.81 |
93
+ | anli r1 / r2 / r3 | 71.90 / 60.00 / **63.08** | 67.20 / 56.20 / 53.42 | **74.00 / 66.30** / 59.42 | 65.60 / 54.30 / 52.25 |
94
+ | wanli | 65.10 | 65.74 | **67.10** | 63.50 |
95
+ | control | 67.20 | 65.96 | **67.58** | 64.35 |
96
+ | MNLI matched / mismatched | 86.52 / 86.71 | 88.95 / **89.46** | 80.64 / 80.54 | **89.17** / 89.35 |
97
+ | arc_easy / arc_challenge | **97.47 / 92.83** | 88.93 / 78.50 | 95.71 / 87.29 | 75.42 / 66.89 |
98
+ | winogrande | **68.59** | 60.69 | 66.30 | 58.33 |
99
+ | gsm8k_mc4 / mc10 | **65.50 / 68.84** | 41.77 / 21.83 | 52.69 / 35.71 | 38.59 / 20.77 |
100
+ | gpqa / gpqa_fewshot | **40.40** / 38.89 | 34.34 / 34.34 | 37.37 / **39.90** | 33.84 / 34.85 |
101
+ | chess | 51.80 | 22.00 | **52.40** | 19.20 |
102
+ | hellaswag | **84.66** | 34.95 | 54.09 | 18.92 |
103
+ | mmlu / mmlu_fewshot | **72.37 / 71.08** | 59.19 / 60.20 | 66.80 / 65.05 | 52.29 / 57.50 |
104
+
105
+ | Kev suite | JPT-4B | NeoHorse-Jev-4B | Kev-4B |
106
+ |---|---:|---:|---:|
107
+ | decision-v7 dev / test | 82.67 / 80.58 | 86.23 / 86.58 | **87.18 / 87.08** |
108
+ | transfer-v4 dev / test | 80.18 / 84.15 | **81.71 / 84.60** | 79.73 / 83.69 |
109
+ | transfer-v9 dev / test | 74.86 / 75.72 | **75.53 / 76.86** | 74.76 / 76.39 |
110
+
111
+ </details>
112
+
113
  ### Other benchmarks
114
 
115
  | Benchmark | What it tests | JPT-4B | Qwen3.5-4B (same prompt, zero-shot) |
116
  |---|---|---:|---:|
117
+ | JevBench public hard tier (111) | hardest general decisions | **0.784** (ECE 0.068, Brier 0.318) | 0.595 |
118
+ | Typed decisions test (2,000) | in-distribution typed decisions | **0.796** (ECE 0.160, Brier 0.332) | 0.596 (ECE 0.171, Brier 0.559) |
119
+ | ANLI r1 / r3 | adversarial natural-language inference | **0.697 / 0.613** | 0.660 / 0.513 |
120
+ | Banking77 / MASSIVE (en / de / zh) | intent classification, incl. multilingual | **0.757 / 0.857 / 0.833 / 0.837** | 0.663 / 0.733 / 0.670 / 0.703 |
121
+ | EnvBench v0.1 public / held-out (skill, 0–100) | sequential decisions in game/puzzle envs (2048, Snake, chess, …) | **47.7 / 47.0** | — |
122
+ | ScreenSpot-v2 set-of-marks (300, images) | GUI element grounding from a screenshot | **0.923** (ECE 0.029) | 0.903 (ECE 0.058) |
123
+ | Screen2Words match (300, images) | screenshot summarization | **0.927** (ECE 0.037) | 0.887 (ECE 0.030) |
124
+ | ERQA (400, images) | embodied/robotics visual reasoning | 0.455 (ECE 0.150) | **0.463** (ECE 0.100) |
125
+
126
+ Numbers are accuracy, with ECE (10 bins) and Brier where shown, at the fitted temperature T = 1.036. The
127
  banking/MASSIVE/typed rows are in-distribution: their train splits are in the training mix, their test items are not.
128
+ Image rows are zero-shot: no image or robotics data was trained on (ERQA is 0.8 points below the base model).
129
 
130
  JevBench public hard tier by family (n):
131
 
 
133
  |---|---:|---:|
134
  | long_policy (19) | **0.74** | 0.47 |
135
  | judge_hard (17) | **0.76** | 0.71 |
136
+ | multi_hop (18) | **0.83** | 0.56 |
137
+ | temporal_numeric (15) | **0.40** | 0.33 |
138
+ | probability (10) | **0.90** | 0.50 |
139
  | trap / adversarial / routing_hard (19) | **1.00** | 0.95 |
140
  | ambiguous (7) | **0.86** | 0.71 |
141
+ | tradeoff (6) | **0.83** | 0.33 |
142
+
143
+ Option order: reversing the options of the 139 public `choice` questions changes 8 verdicts; accuracy is 0.892
144
+ either way.
145
 
146
  ## Quick start
147
 
 
151
  ```bash
152
  python -m sglang.launch_server --model-path kirp/jpt-4b --port 30000 \
153
  --context-length 32768 --mamba-scheduler-strategy extra_buffer & # Qwen3.5's DeltaNet layers need this flag
154
+ llm2jev --model kirp/jpt-4b --backend sglang --url http://127.0.0.1:30000 --port 8080 --temperature 1.036
155
  ```
156
 
157
  vLLM instead: `vllm serve kirp/jpt-4b --max-logprobs 256 --return-tokens-as-token-ids --enable-scale-out --port 8000`
158
+ then `llm2jev --model kirp/jpt-4b --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.036` — the three vLLM flags
159
  are required, not optional: without them every request comes back a plain HTTP 400 with no hint why.
160
 
161
  No GPU / just trying it out, no engine, no clone — llm2jev runs the model itself:
162
 
163
  ```bash
164
  pip install "llm2jev[hf,vision]"
165
+ llm2jev --model kirp/jpt-4b --backend hf --port 8080 --temperature 1.036
166
  ```
167
 
168
  This path serializes requests (one forward at a time in-process — concurrent calls are safe, just not parallel);
 
170
  use — `hf` is for a quick check, not for traffic.
171
 
172
  Tested with SGLang 0.5.9. Its torch 2.9.1 pins cuDNN 9.10, which SGLang refuses to run on, so install cuDNN 9.15+ over
173
+ it: `pip install "sglang==0.5.9" && pip install "nvidia-cudnn-cu12>=9.15"`. `--temperature 1.036` is the temperature
174
  fitted on the calibration split; leaving it at 1.0 changes calibration slightly, never the ranking.
175
 
176
  ```python
 
201
  - **Method:** LoRA (r=16) on every attention, DeltaNet and MLP projection of the language model, merged into full
202
  weights. The vision tower is untouched. The loss is the multi-class Brier score over the option labels, on
203
  llm2jev's chat prompt with thinking disabled: the same prompt the model is served with.
204
+ - **Data** (49,221 questions in 32,835 records, all converted to typed decisions; one epoch over two option-shuffled copies):
205
  - public classification, NLI, QA, preference and safety datasets;
206
  - long legal and contract documents (ContractNLI, MAUD, LegalBench, ConditionalQA, ShARC);
207
  - table and numeric reasoning (TAT-QA, MultiHiertt);
 
210
  - community typed-decision sets;
211
  - oracle-labelled rollouts from 20 small game and puzzle environments;
212
  - programmatically generated rule-arithmetic items (dates, time zones, day counts, caps; labels computed by code);
213
+ - 1,289 long policy / contract / regulation documents (11–17k characters, 4,471 questions) with amendments,
214
+ exceptions, precedence rules and constrained trade-offs (ranked rules, scarce-resource allocation, authority
215
+ limits), written by an LLM (GPT-6 Luna) with no JevBench item shown to it.
216
  - **Held out:** no item from JevBench, EnvBench's held-out seeds, the Decision Index frozen suite or our typed test
217
  split was used in training. Documents were checked for 8-gram overlap with JevBench.
218
 
assets/2048.gif CHANGED

Git LFS Details

  • SHA256: 47a9dba83c485933e032eede29d799432578cb0111ca022992854635ef19e38c
  • Pointer size: 131 Bytes
  • Size of remote file: 138 kB

Git LFS Details

  • SHA256: 195f8aafb9dd13689104972ae215f742b5587572087211309eee66b02cfea6d9
  • Pointer size: 131 Bytes
  • Size of remote file: 132 kB
assets/snake.gif CHANGED

Git LFS Details

  • SHA256: 27673ba822e63cb2eda4e2797ad38cb71ef4db1911cfd446ea36b0c0d0024a58
  • Pointer size: 131 Bytes
  • Size of remote file: 131 kB

Git LFS Details

  • SHA256: 757ae7733ba359cc2c7a9531ea7c7b768494a27102a005d63e08054ddc63b7bc
  • Pointer size: 131 Bytes
  • Size of remote file: 130 kB
assets/tetris.gif CHANGED

Git LFS Details

  • SHA256: 7d50c279e20c13ac713cee1b0d0228f4ac2c9b2dbc1bd3f7aaa4b56a6ad9a8e0
  • Pointer size: 131 Bytes
  • Size of remote file: 147 kB

Git LFS Details

  • SHA256: 101f15f1a56abbc157227327b21e72b9320d67b739d3641f421d792468044168
  • Pointer size: 131 Bytes
  • Size of remote file: 158 kB