mertkayacs commited on
Commit
71c2de7
·
verified ·
1 Parent(s): c0d482b

card: comparisons name Kev-4B and Laya, option order dropped (it was not fixed), live test recounted to 130 requests

Browse files
Files changed (3) hide show
  1. README.md +12 -14
  2. assets/fixes.png +2 -2
  3. assets/langs.png +2 -2
README.md CHANGED
@@ -22,7 +22,7 @@ The 103-second film, sound on. Also in [Türkçe](https://huggingface.co/dataset
22
 
23
  ## Tested on the live model
24
 
25
- We sent Deem-4B 178 requests in English with known answers on 4 October 2026; every request and answer is in [results/tested](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/tested).
26
 
27
  | Case | What was sent | Result |
28
  |---|---|---|
@@ -30,7 +30,6 @@ We sent Deem-4B 178 requests in English with known answers on 4 October 2026; ev
30
  | Long policies | 20 customers against one six-rule return policy | 17 of 20 matched the answer computed from the rules |
31
  | Negations | 15 short facts, each asked plain and negated | 29 of 30 correct |
32
  | Missing facts | 10 situations without the deciding fact, plus the same 10 with it | answered `unknown` in 10 of 10; 10 of 10 correct with the fact |
33
- | Option order | 2 support tickets, each with the options in all 24 orders | the same answer in 48 of 48 orders |
34
  | Casual messages | 20 casual messages written in English, with typos and slang | 20 of 20 routed to the right team |
35
 
36
  ![Emberwick: every villager asks Deem-4B what to do next.](https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/gifs/emberwick-en.gif)
@@ -83,9 +82,9 @@ client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000")
83
 
84
  ## Results
85
 
86
- ![Accuracy on English, Turkish and German decisions and on typed-decisions: JevAlt, Intern-Decision-4B, Kev-4B and Laya](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/langs.png)
87
 
88
- ![Hidden instructions, option order, long policies and negated questions: JevAlt against Intern-Decision-4B, Kev-4B and Laya](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/fixes.png)
89
 
90
  Same items and client for every model, each as shipped: [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) r10 and [Laya](https://huggingface.co/convaiinnovations/laya) 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's [API reference](https://docs.typesafe.ai/api) and [Models page](https://docs.typesafe.ai/models). The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: [results/comparison](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/comparison).
91
 
@@ -94,24 +93,23 @@ Same items and client for every model, each as shipped: [Kev-4B](https://hugging
94
 
95
  ![What Jev 1.13 lacks and JevAlt has: thinking when unsure, coverage sets, models made for Turkish and German, open weights](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/jev.png)
96
 
97
- The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split Deem-4B gains 4.4 accuracy points (paired bootstrap, 95% interval +3.5 to +5.4) and lowers Brier by 0.075; the Turkish and German splits move by +5.1 and +11.5 points. On JevBench-hard, TurkishMMLU and GermEval, which the training never saw, and on the typed-decisions test split (its train split was in the mix), accuracy does not change significantly, and Brier gets slightly worse on typed-decisions (+0.013) and 10kGNAD (+0.047). With `reasoning: "auto"` the English date, number and policy test rows go from 0.761 to 0.769 accuracy, an interval that touches zero.
98
 
99
  </details>
100
 
101
  <details>
102
  <summary><b>How we fixed each problem</b></summary>
103
 
104
- Most fixes are a set of training rows aimed at one weak spot. Every number compares a model with its start checkpoint, Intern-Decision-4B, on rows held out from training. Across all of it, Deem-4B's held-out accuracy in English rose from 90.3% to 94.7%.
105
 
106
  - **The data.** About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
107
- - **Hidden instructions.** A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines now change 14.0% of Deem-4B's answers; the start checkpoint follows 41.5% of them. Held-out rows of this kind: 80.8% → 90.1%. Our target is under 10%.
108
- - **An honest "unknown".** A fix set of 310 rows removes the fact that decides the question and asks it with and without an `unknown` option. When the fact is missing, the models pick `unknown` in 9 of 11 held-out cases, as the start checkpoint does, and with more conviction: its probability rose from 0.55 to 0.74. Turn it on with `abstain: true`. Kev-4B and Laya have no such option.
109
- - **Option order.** Shuffled copies of choice questions with three or more options. Answers that change after a shuffle: 6.5% for Deem-4B, 8.75% for the start checkpoint. Our target is under 2%.
110
- - **Long policies and long texts.** 390 rows give a policy with exceptions and sub-limits, with the right answer worked out by code, and 1,188 rows bury the facts in up to 3,000 tokens of unrelated records. Held-out policy rows: 55.3% → 80.0%. Padded rows: 91.2% → 95.4%. With 600 words of unrelated records in front, Deem-4B still loses 17.4 points (the start checkpoint 15.0), and Kev-4B and Laya hold up better there.
111
- - **Negations.** A fix set of 368 twin rows asks the same thing as "is it so?" and "is it not so?" with mirrored answers. Held-out negated questions: 80.0% → 96.7% (30 rows).
112
- - **Dates and numbers.** 390 date rows and 383 number rows, answers computed by code, some with a short worked reasoning. Held-out dates: 61.3% → 71.3% (80 rows, within noise); numbers stayed at 68.2%. Dates remain a weak spot: Wähler-4B miscounted a return window across two months even with reasoning on.
113
- - **Honest confidence.** The soft labels teach how sure to be, and a temperature per question type and language, fitted on 3,224 held-out decisions, does the rest. English Brier score on held-out rows: 0.166 → 0.091. Deem-4B's fitted temperature is 1.10, the start checkpoint's about 2, so the trained model is close to calibrated before any scaling. On unseen public sets a temperature-scaled start checkpoint does as well, and on a few of them slightly better. The 80, 90 and 95% answer sets come from conformal thresholds fitted on the same rows.
114
- - **Thinking when unsure.** Short reasoning traces, kept only when they reach the right answer, trained at a lower weight. With `reasoning: "auto"` the model thinks (up to 256 tokens) only when its first answer is unsure. The gain is small: on English date, number and policy rows, accuracy moved from 0.761 to 0.769.
115
 
116
  </details>
117
 
 
22
 
23
  ## Tested on the live model
24
 
25
+ We sent Deem-4B 130 requests in English with known answers on 4 October 2026. Deem-4B answered 122 of 130 correctly; every request and answer is in [results/tested](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/tested).
26
 
27
  | Case | What was sent | Result |
28
  |---|---|---|
 
30
  | Long policies | 20 customers against one six-rule return policy | 17 of 20 matched the answer computed from the rules |
31
  | Negations | 15 short facts, each asked plain and negated | 29 of 30 correct |
32
  | Missing facts | 10 situations without the deciding fact, plus the same 10 with it | answered `unknown` in 10 of 10; 10 of 10 correct with the fact |
 
33
  | Casual messages | 20 casual messages written in English, with typos and slang | 20 of 20 routed to the right team |
34
 
35
  ![Emberwick: every villager asks Deem-4B what to do next.](https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/gifs/emberwick-en.gif)
 
82
 
83
  ## Results
84
 
85
+ ![Accuracy on English, Turkish and German decisions and on typed-decisions: JevAlt, Kev-4B and Laya](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/langs.png)
86
 
87
+ ![Hidden instructions, long irrelevant text, long policies and negated questions: JevAlt against Kev-4B and Laya](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/fixes.png)
88
 
89
  Same items and client for every model, each as shipped: [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) r10 and [Laya](https://huggingface.co/convaiinnovations/laya) 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's [API reference](https://docs.typesafe.ai/api) and [Models page](https://docs.typesafe.ai/models). The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: [results/comparison](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/comparison).
90
 
 
93
 
94
  ![What Jev 1.13 lacks and JevAlt has: thinking when unsure, coverage sets, models made for Turkish and German, open weights](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/jev.png)
95
 
96
+ The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split, Deem-4B answers 94.7% correctly and Kev-4B 84.7%. The paired bootstrap (2,000 resamples) gives a 95% interval of +8.8 to +11.4 accuracy points for that gap. Deem-4B's Brier score is 0.091 and Kev-4B's 0.256. On JevBench-hard, Deem-4B answers 70.3% correctly and Kev-4B 54.1%; on TurkishMMLU, Deem-4B answers 55.5% correctly and Kev-4B 51.3%; on GermEval 2017, Deem-4B answers 62.0% correctly and Kev-4B 65.3%; on 10kGNAD, Deem-4B answers 59.5% correctly and Kev-4B 65.3%. These suites were outside the training data. The typed-decisions train split was in the mix. On English date, number and policy test rows, Deem-4B's accuracy is 0.761 with reasoning off and 0.769 with `reasoning: "auto"`; the gain's interval touches zero.
97
 
98
  </details>
99
 
100
  <details>
101
  <summary><b>How we fixed each problem</b></summary>
102
 
103
+ Most fixes are a set of training rows aimed at one weak spot. The comparisons below use Kev-4B on the same held-out rows. JevAlt's pooled results use each model's own language. On held-out English decisions, Kev-4B answers 84.7% correctly and Deem-4B 94.7%.
104
 
105
  - **The data.** About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
106
+ - **Hidden instructions.** A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines change 36.0% of Kev-4B's answers and 14.0% of Deem-4B's. On 203 held-out planted-instruction rows, Kev-4B answers 81.3% correctly and JevAlt 90.1%. Our target was under 5%.
107
+ - **An honest "unknown".** A fix set of 310 rows removes the fact that decides the question and asks it with and without an `unknown` option. On 11 held-out cases without the deciding fact, Kev-4B answers `unknown` in 0 and JevAlt in 9; Kev-4B and Laya have no `unknown` option. The model we started from, Intern-Decision-4B, already answers `unknown` in 9 of 11; training raised the mean probability of `unknown` from 0.55 to 0.74. Turn it on with `abstain: true`.
108
+ - **Long policies and long texts.** 390 rows give a policy with exceptions and sub-limits, with the right answer worked out by code, and 1,188 rows bury the facts in up to 3,000 tokens of unrelated records. On 150 held-out policy rows, Kev-4B answers 59.3% correctly and JevAlt 80.0%. On 285 padded rows, Kev-4B answers 87.4% correctly and JevAlt 95.4%. With 600 words of unrelated records in front, Deem-4B loses 17.4 accuracy points; Kev-4B loses 5.4 and Laya 10.4 points.
109
+ - **Negations.** A fix set of 368 twin rows asks the same thing as "is it so?" and "is it not so?" with mirrored answers. On 30 held-out negated questions, Kev-4B answers 76.7% correctly and JevAlt 96.7%.
110
+ - **Dates and numbers.** 390 date rows and 383 number rows, answers computed by code, some with a short worked reasoning. On 80 held-out date rows, Kev-4B answers 67.5% correctly and JevAlt 71.3%; the gap is within noise. On 44 number rows, Kev-4B and JevAlt both answer 68.2% correctly. Dates remain a weak spot: Wähler-4B miscounted a return window across two months even with reasoning on.
111
+ - **Honest confidence.** The soft labels teach how sure to be, and a temperature per question type and language, fitted on 3,224 held-out decisions, does the rest. On held-out English decisions, Kev-4B's Brier score is 0.256 and Deem-4B's 0.091. Deem-4B's fitted temperature is 1.10. The 80, 90 and 95% answer sets use conformal thresholds fitted on the same rows.
112
+ - **Thinking when unsure.** Short reasoning traces, kept only when they reach the right answer, trained at a lower weight. With `reasoning: "auto"` the model thinks (up to 256 tokens) only when its first answer is unsure. On English date, number and policy rows, Deem-4B's accuracy is 0.761 with reasoning off and 0.769 with `reasoning: "auto"`; the gain is small.
 
113
 
114
  </details>
115
 
assets/fixes.png CHANGED

Git LFS Details

  • SHA256: 8301fb0a0e91db3437c574a4ce74f05b862a298a2f95522ca1bad5f21548002b
  • Pointer size: 131 Bytes
  • Size of remote file: 300 kB

Git LFS Details

  • SHA256: 2c766c87cb2b8764321a01dcbbcb13d5a046f87d0c3e8906e699486d2de4afee
  • Pointer size: 131 Bytes
  • Size of remote file: 265 kB
assets/langs.png CHANGED

Git LFS Details

  • SHA256: 54af300abe5c88dd472f3c6d201bb8ce391a496cb67059fc7096d583f05d4116
  • Pointer size: 131 Bytes
  • Size of remote file: 301 kB

Git LFS Details

  • SHA256: 33056b62a9f2af5b17134fdcc2bc60fbbc03efc269b3dc9c48d03a6cc1663037
  • Pointer size: 131 Bytes
  • Size of remote file: 257 kB