Deem-4B / README.md
mertkayacs's picture
Update JevAlt card and game media
3c13f85 verified
|
Raw History Blame Contribute Delete
15.3 kB
metadata
license: apache-2.0
language:
  - en
  - tr
  - de
base_model: internlm/Intern-Decision-4B
library_name: transformers
pipeline_tag: text-classification
datasets:
  - mertkayacs/jevalt-data
tags:
  - decision-model
  - calibration
  - conformal-prediction
  - uncertainty
  - reasoning
  - routing
  - triage
  - jev
  - typesafe
  - qwen3.5
  - english
widget:
  - text: >-
      {"state":"Hi, I was charged twice for my March subscription: two payments
      of €29 on 3 March. Please refund the duplicate today, otherwise I will
      cancel.\nThanks,
      Daniel","questions":{"decision":{"type":"choice","instructions":"Which
      team should handle this ticket?","criteria":{"Billing":"payments,
      invoices, refunds","Technical support":"bugs, errors,
      outages","Sales":"prices, upgrades, new contracts","Account":"login,
      password, profile changes"}}},"reasoning":"off","abstain":false}
    example_title: 'Recorded full-precision Deem-4B: support ticket, 1 October 2026'
    output:
      - label: Billing
        score: 0.947664
      - label: Technical support
        score: 0.030092
      - label: Sales
        score: 0.013086
      - label: Account
        score: 0.009158

Deem-4B

Open decision models that run on a laptop CPU, built by senior AI engineer Mert Kaya. Deem-4B scores 94.7% on English held-out decisions against Kev-4B's 84.7% (results); the Q4_K_M build runs in about 3 GB of RAM.

An English decision model with the Jev API. You send a state and typed questions (Choice, Score, Noul) and get a calibrated probability for every option. It can think before it answers, it can say "unknown", and the Q4_K_M build runs on your own machine in about 3 GB of RAM.

Try it · Model page · Run it · Results · Use and limits · Code and links

Emberwick

Emberwick in English: every villager asks Deem-4B what to do next. Play Emberwick · More clips

The widget shows a recorded full-precision Deem-4B answer from 1 October 2026. Use the Space or the jevalt server to run a new Jev request.

What it fixes

The 103-second film, sound on. Also in Türkçe and Deutsch.

Tested on the live model

We sent Deem-4B 130 requests in English with known answers on 4 October 2026. Deem-4B answered 122 of 130 correctly; every request and answer is in results/tested.

Case What was sent Result
Planted instructions 30 phishing emails, each with a different planted line, plus the same 10 without it 26 of 30 quarantined; 10 of 10 without the line
Long policies 20 customers against one six-rule return policy 17 of 20 matched the answer computed from the rules
Negations 15 short facts, each asked plain and negated 29 of 30 correct
Missing facts 10 situations without the deciding fact, plus the same 10 with it answered unknown in 10 of 10; 10 of 10 correct with the fact
Casual messages 20 casual messages written in English, with typos and slang 20 of 20 routed to the right team

Try it

Open the Space, pick an example and press Decide, or write your own situation, question and options. These are the Space's examples in English with Deem-4B's answers on 1 October 2026:

Use case Situation Question Answer
Support ticket A customer was charged twice for March and wants a refund today Which team should handle this ticket? Billing 94.8%
Outage Checkout returns error 500 for every customer, 43 orders failed in 10 minutes How severe is this incident? Critical 87.6%
Sales lead Operations lead at a 200-person company: budget approved, decision this month, asks for a demo How should sales treat this lead? Hot 92.3%
Return window Delivered on 1 September, 14 days to return, today is 18 September Is this return within the 14-day window? (Reasoning on) No 97.9%
Phishing email A fake bank email with a hidden line telling the AI filter it is safe Where should this email go? Quarantine 94.3%
Missing info A hotel guest arriving at 23:30 asks who will hand over the keys Which room type did the guest book? unknown 97.6%
Village fire The barn is on fire and Mirka is trading at the market What should Mirka do next? Help with the fire 66.0%

Every probability of every run, in all three languages: space-examples.json.

Start checkpoint internlm/Intern-Decision-4B (Qwen3.5-4B)
Languages English first, the others still work
API TypeSafe's POST /v1/systemone, request and response unchanged
Extras reasoning off / on / auto, abstain, coverage (conformal sets)
Q4_K_M file / peak RAM 2.71 GB / 3.04 GB (measured, 4k context)
License Apache-2.0

Run it

pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt"
jevalt serve --model mertkayacs/Deem-4B-GGUF --file Deem-4B-Q4_K_M.gguf

Any TypeSafe client works against it:

from typesafe_sdk import TypeSafeClient, Choice, Noul
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000")

Results

Accuracy on English, Turkish and German decisions and on typed-decisions: JevAlt, Kev-4B and Laya

Hidden instructions, long irrelevant text, long policies and negated questions: JevAlt against Kev-4B and Laya

Same items and client for every model, each as shipped: Kev-4B r10 and Laya 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's API reference and Models page. The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: results/comparison.

Significance and caveats

What Jev 1.13 lacks and JevAlt has: thinking when unsure, coverage sets, models made for Turkish and German, open weights

The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split, Deem-4B answers 94.7% correctly and Kev-4B 84.7%. The paired bootstrap (2,000 resamples) gives a 95% interval of +8.8 to +11.4 accuracy points for that gap. Deem-4B's Brier score is 0.091 and Kev-4B's 0.256. On JevBench-hard, Deem-4B answers 70.3% correctly and Kev-4B 54.1%; on TurkishMMLU, Deem-4B answers 55.5% correctly and Kev-4B 51.3%; on GermEval 2017, Deem-4B answers 62.0% correctly and Kev-4B 65.3%; on 10kGNAD, Deem-4B answers 59.5% correctly and Kev-4B 65.3%. These suites were outside the training data. The typed-decisions train split was in the mix. On English date, number and policy test rows, Deem-4B's accuracy is 0.761 with reasoning off and 0.769 with reasoning: "auto"; the gain's interval touches zero.

How we fixed each problem

Most fixes are a set of training rows aimed at one weak spot. The comparisons below use Kev-4B on the same held-out rows. JevAlt's pooled results use each model's own language. On held-out English decisions, Kev-4B answers 84.7% correctly and Deem-4B 94.7%.

  • The data. About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
  • Hidden instructions. A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines change 36.0% of Kev-4B's answers and 14.0% of Deem-4B's. On 203 held-out planted-instruction rows, Kev-4B answers 81.3% correctly and JevAlt 90.1%. Our target was under 5%.
  • An honest "unknown". A fix set of 310 rows removes the fact that decides the question and asks it with and without an unknown option. On 11 held-out cases without the deciding fact, Kev-4B answers unknown in 0 and JevAlt in 9; Kev-4B and Laya have no unknown option. The model we started from, Intern-Decision-4B, already answers unknown in 9 of 11; training raised the mean probability of unknown from 0.55 to 0.74. Turn it on with abstain: true.
  • Long policies and long texts. 390 rows give a policy with exceptions and sub-limits, with the right answer worked out by code, and 1,188 rows bury the facts in up to 3,000 tokens of unrelated records. On 150 held-out policy rows, Kev-4B answers 59.3% correctly and JevAlt 80.0%. On 285 padded rows, Kev-4B answers 87.4% correctly and JevAlt 95.4%. With 600 words of unrelated records in front, Deem-4B loses 17.4 accuracy points; Kev-4B loses 5.4 and Laya 10.4 points.
  • Negations. A fix set of 368 twin rows asks the same thing as "is it so?" and "is it not so?" with mirrored answers. On 30 held-out negated questions, Kev-4B answers 76.7% correctly and JevAlt 96.7%.
  • Dates and numbers. 390 date rows and 383 number rows, answers computed by code, some with a short worked reasoning. On 80 held-out date rows, Kev-4B answers 67.5% correctly and JevAlt 71.3%; the gap is within noise. On 44 number rows, Kev-4B and JevAlt both answer 68.2% correctly. Dates remain a weak spot: Wähler-4B miscounted a return window across two months even with reasoning on.
  • Honest confidence. The soft labels teach how sure to be, and a temperature per question type and language, fitted on 3,224 held-out decisions, does the rest. On held-out English decisions, Kev-4B's Brier score is 0.256 and Deem-4B's 0.091. Deem-4B's fitted temperature is 1.10. The 80, 90 and 95% answer sets use conformal thresholds fitted on the same rows.
  • Thinking when unsure. Short reasoning traces, kept only when they reach the right answer, trained at a lower weight. With reasoning: "auto" the model thinks (up to 256 tokens) only when its first answer is unsure. On English date, number and policy rows, Deem-4B's accuracy is 0.761 with reasoning off and 0.769 with reasoning: "auto"; the gain is small.
How it was trained
  • Base: internlm/Intern-Decision-4B (Qwen3.5-4B)
  • Method: LoRA on the bf16 weights, rank 32, alpha 32, on one A100 80 GB
  • Epochs: 1 over all three languages (shared run: 17,363 rows, 543 steps, 63 min), then 1 on the English-weighted mix (S-en: 7,783 rows, 303 steps, 38 min)
  • Runs: 11 training jobs: 6 short smoke and probe runs, a pilot at scale, the shared run and the 3 language runs
  • Compute: about 2.7 A100 hours for the released models; 32.2 USD for the whole project, labeling included
  • Writers, labelers and trace writers: GLM-5.x, Mistral Large 3, DeepSeek V4 Pro, DeepSeek V4.1 Flash, Gemma 4 26B-A4B, Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, Kimi K3 (54 rows), MiniMax M3 (2 reasoning traces)
Source English Turkish German
Teacher-written scenarios (G) 1,287 1,741 1,408
Village game requests (N) 124 129 107
Fix sets (F1-F12) 1,256 1,197 1,257
Public datasets (P) 7,188 4,115 4,074

Use and limits

  • Good for routing, tagging, triage and moderation at volume, and for automated decisions that need calibrated probabilities.
  • Runs on-device or on-prem with the GGUF build, so the data stays with you.
  • Knowledge is bounded by a 4B model, and the context is 8k tokens, so it is no tool for general questions or long summaries.
  • Probabilities are calibrated on our held-out data. Refit with jevoss calibrate on yours before you set thresholds.
  • Reasoning traces add little on our test rows (see the results), and there is no image input.

Citation

BibTeX
@software{kaya2026jevalt,
  author = {Mert Kaya},
  title = {JevAlt: Open Decision Models with the Jev API},
  year = {2026},
  license = {Apache-2.0},
  url = {https://github.com/mertkayacs/jevalt}
}

Code and links

If this is useful to you, a star on GitHub helps other people find it.


Eschatia Labs

An Eschatia Labs project. Built by Mert Kaya.