--- license: apache-2.0 language: - en - tr - de base_model: internlm/Intern-Decision-4B library_name: transformers pipeline_tag: text-classification datasets: - mertkayacs/jevalt-data tags: - decision-model - calibration - conformal-prediction - uncertainty - reasoning - routing - triage - jev - typesafe - qwen3.5 - english - small-language-model - local-llm - on-device - text-classification - english-llm widget: - text: '{"state":"Hi, I was charged twice for my March subscription: two payments of €29 on 3 March. Please refund the duplicate today, otherwise I will cancel.\nThanks, Daniel","questions":{"decision":{"type":"choice","instructions":"Which team should handle this ticket?","criteria":{"Billing":"payments, invoices, refunds","Technical support":"bugs, errors, outages","Sales":"prices, upgrades, new contracts","Account":"login, password, profile changes"}}},"reasoning":"off","abstain":false}' example_title: 'Recorded full-precision Deem-4B: support ticket, 1 October 2026' output: - label: Billing score: 0.947664 - label: Technical support score: 0.030092 - label: Sales score: 0.013086 - label: Account score: 0.009158 base_model_relation: finetune --- # Deem-4B: small English language model for local decisions Deem-4B is one of the best open English decision models at 4B parameters: a small English language model for text classification, routing and decisions with probabilities. **Deem-4B: 94.7% accuracy; Kev-4B: 84.7%** on held-out English decisions. The test set comes from the same data pipeline as the training data. [Try a decision in the browser](https://huggingface.co/spaces/mertkayacs/JevAlt), or run the GGUF build locally below. ## Run locally The Q4_K_M build used **3.04 GB RAM** at a 4k context. The JevAlt server uses llama.cpp to read decision probabilities and load the matching calibration. ```sh pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serve --model mertkayacs/Deem-4B-GGUF --file Deem-4B-Q4_K_M.gguf ``` Send a situation, question and options to the local server: ```python import requests request = { "state": "I was charged twice for my subscription. Please refund the duplicate.", "questions": { "team": { "type": "choice", "instructions": "Which team should handle this ticket?", "criteria": {"billing": "payments, refunds", "technical": "bugs, outages"}, } }, "reasoning": "off", "abstain": False, } response = requests.post("http://127.0.0.1:8000/v1/systemone", json=request) response.raise_for_status() print(response.json()) ``` Write the situation, question and options in English. The API also supports yes/no questions, ordered scores and an `unknown` option with `abstain: true`. The card widget shows a recorded answer from the full-precision model.
Use the full-precision weights with Transformers This repository holds the full-precision weights. The JevAlt Transformers backend preserves the decision-token format and calibration: ```sh pip install "jevalt[hf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serve --backend hf --model mertkayacs/Deem-4B ``` See the [local run guide](https://github.com/mertkayacs/jevalt/blob/main/docs/local.md) for memory and client settings.
## Results and limits On the same held-out English decisions, Deem-4B's probability error (Brier score, lower is better) is **0.091**, against **0.256 for Kev-4B**. A separate live test on 4 October 2026 scored **122/130** for Deem-4B. [Comparison files](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/comparison) and [live requests, including misses](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/tested) record both tests. The models were fine-tuned with LoRA from [Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B), using [JevAlt training data](https://huggingface.co/datasets/mertkayacs/jevalt-data). English is this model's focus. - Refit calibration on your own data before using confidence thresholds. See [JevOss](https://github.com/mertkayacs/jevoss). - The model accepts text, with an 8k context. Long irrelevant text and date arithmetic remain weak spots. - Planted instructions still change some answers. Keep authorization checks outside the model and review consequential decisions. ## Downloads and project [GGUF files and measured agreement](https://huggingface.co/mertkayacs/Deem-4B-GGUF) | [Code and training](https://github.com/mertkayacs/jevalt) | [Deem-4B project page](https://jevalt.mertkayacs.com/models/deem-4b/) | [Training and evaluation details](https://github.com/mertkayacs/jevalt/blob/main/docs/how-it-was-built.md). Apache-2.0. Cite the [JevAlt repository](https://github.com/mertkayacs/jevalt) for this release. An [Eschatia Labs](https://eschatialabs.com) project. [Mert Kaya](https://mertkayacs.com).