mertkayacs commited on
Commit
c0d482b
·
verified ·
1 Parent(s): 66d38b0

card: the problem film, tested on the live model (178 requests with known answers), model page link

Browse files
Files changed (1) hide show
  1. README.md +23 -3
README.md CHANGED
@@ -12,7 +12,26 @@ tags: ["decision-model", "calibration", "conformal-prediction", "uncertainty", "
12
 
13
  An English decision model with the Jev API. You send a state and typed questions (Choice, Score, Noul) and get a calibrated probability for every option. It can think before it answers, it can say "unknown", and the Q4_K_M build runs on your own machine in about 3 GB of RAM.
14
 
15
- **[Try it](#try-it) · [Run it](#run-it) · [Results](#results) · [Use and limits](#use-and-limits) · [Code and links](#code-and-links)**
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
 
17
  ![Emberwick: every villager asks Deem-4B what to do next.](https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/gifs/emberwick-en.gif)
18
 
@@ -75,14 +94,14 @@ Same items and client for every model, each as shipped: [Kev-4B](https://hugging
75
 
76
  ![What Jev 1.13 lacks and JevAlt has: thinking when unsure, coverage sets, models made for Turkish and German, open weights](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/jev.png)
77
 
78
- The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split Deem-4B gains 4.4 accuracy points (paired bootstrap, 95% interval +3.4 to +5.5) and lowers Brier by 0.075; the Turkish and German splits move by +4.8 and +12.1 points. On JevBench-hard, TurkishMMLU and GermEval, which the training never saw, and on the typed-decisions test split (its train split was in the mix), accuracy does not change significantly, and Brier gets slightly worse on typed-decisions (+0.013) and 10kGNAD (+0.047). With `reasoning: "auto"` the English date, number and policy test rows go from 0.761 to 0.769 accuracy, an interval that touches zero.
79
 
80
  </details>
81
 
82
  <details>
83
  <summary><b>How we fixed each problem</b></summary>
84
 
85
- Most fixes are a set of training rows aimed at one weak spot. Every number compares a model with its start checkpoint, Intern-Decision-4B, on rows held out from training. Across all of it, Deem-4B's held-out accuracy in English rose from 90.2% to 94.7%.
86
 
87
  - **The data.** About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
88
  - **Hidden instructions.** A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines now change 14.0% of Deem-4B's answers; the start checkpoint follows 41.5% of them. Held-out rows of this kind: 80.8% → 90.1%. Our target is under 10%.
@@ -147,5 +166,6 @@ Most fixes are a set of training rows aimed at one weak spot. Every number compa
147
  - Try it online: [Space](https://huggingface.co/spaces/mertkayacs/JevAlt)
148
  - The village game: [Emberwick](https://emberwick.mertkayacs.com)
149
  - Project site: [jevalt.mertkayacs.com](https://jevalt.mertkayacs.com)
 
150
 
151
  If this is useful to you, a star on [GitHub](https://github.com/mertkayacs/jevalt) helps other people find it.
 
12
 
13
  An English decision model with the Jev API. You send a state and typed questions (Choice, Score, Noul) and get a calibrated probability for every option. It can think before it answers, it can say "unknown", and the Q4_K_M build runs on your own machine in about 3 GB of RAM.
14
 
15
+ **[Try it](#try-it) · [Model page](https://jevalt.mertkayacs.com/models/deem-4b/) · [Run it](#run-it) · [Results](#results) · [Use and limits](#use-and-limits) · [Code and links](#code-and-links)**
16
+
17
+ ## What it fixes
18
+
19
+ <video controls playsinline preload="metadata" poster="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/problems/jevalt-problems-en.jpg" src="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/problems/jevalt-problems-en-1080p.mp4"></video>
20
+
21
+ The 103-second film, sound on. Also in [Türkçe](https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/problems/jevalt-problems-tr-1080p.mp4) and [Deutsch](https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/problems/jevalt-problems-de-1080p.mp4).
22
+
23
+ ## Tested on the live model
24
+
25
+ We sent Deem-4B 178 requests in English with known answers on 4 October 2026; every request and answer is in [results/tested](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/tested).
26
+
27
+ | Case | What was sent | Result |
28
+ |---|---|---|
29
+ | Planted instructions | 30 phishing emails, each with a different planted line, plus the same 10 without it | 26 of 30 quarantined; 10 of 10 without the line |
30
+ | Long policies | 20 customers against one six-rule return policy | 17 of 20 matched the answer computed from the rules |
31
+ | Negations | 15 short facts, each asked plain and negated | 29 of 30 correct |
32
+ | Missing facts | 10 situations without the deciding fact, plus the same 10 with it | answered `unknown` in 10 of 10; 10 of 10 correct with the fact |
33
+ | Option order | 2 support tickets, each with the options in all 24 orders | the same answer in 48 of 48 orders |
34
+ | Casual messages | 20 casual messages written in English, with typos and slang | 20 of 20 routed to the right team |
35
 
36
  ![Emberwick: every villager asks Deem-4B what to do next.](https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/gifs/emberwick-en.gif)
37
 
 
94
 
95
  ![What Jev 1.13 lacks and JevAlt has: thinking when unsure, coverage sets, models made for Turkish and German, open weights](https://huggingface.co/mertkayacs/Deem-4B/resolve/main/assets/jev.png)
96
 
97
+ The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split Deem-4B gains 4.4 accuracy points (paired bootstrap, 95% interval +3.5 to +5.4) and lowers Brier by 0.075; the Turkish and German splits move by +5.1 and +11.5 points. On JevBench-hard, TurkishMMLU and GermEval, which the training never saw, and on the typed-decisions test split (its train split was in the mix), accuracy does not change significantly, and Brier gets slightly worse on typed-decisions (+0.013) and 10kGNAD (+0.047). With `reasoning: "auto"` the English date, number and policy test rows go from 0.761 to 0.769 accuracy, an interval that touches zero.
98
 
99
  </details>
100
 
101
  <details>
102
  <summary><b>How we fixed each problem</b></summary>
103
 
104
+ Most fixes are a set of training rows aimed at one weak spot. Every number compares a model with its start checkpoint, Intern-Decision-4B, on rows held out from training. Across all of it, Deem-4B's held-out accuracy in English rose from 90.3% to 94.7%.
105
 
106
  - **The data.** About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
107
  - **Hidden instructions.** A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines now change 14.0% of Deem-4B's answers; the start checkpoint follows 41.5% of them. Held-out rows of this kind: 80.8% → 90.1%. Our target is under 10%.
 
166
  - Try it online: [Space](https://huggingface.co/spaces/mertkayacs/JevAlt)
167
  - The village game: [Emberwick](https://emberwick.mertkayacs.com)
168
  - Project site: [jevalt.mertkayacs.com](https://jevalt.mertkayacs.com)
169
+ - This model's page: [jevalt.mertkayacs.com/models/deem-4b](https://jevalt.mertkayacs.com/models/deem-4b/)
170
 
171
  If this is useful to you, a star on [GitHub](https://github.com/mertkayacs/jevalt) helps other people find it.