Instructions to use mertkayacs/Wahler-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mertkayacs/Wahler-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mertkayacs/Wahler-4B")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("mertkayacs/Wahler-4B") model = AutoModelForMultimodalLM.from_pretrained("mertkayacs/Wahler-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("mertkayacs/Wahler-4B")
model = AutoModelForMultimodalLM.from_pretrained("mertkayacs/Wahler-4B", device_map="auto")Wähler-4B
Open decision models that run on a laptop CPU, built by senior AI engineer Mert Kaya. Wähler-4B scores 92.0% on German held-out decisions against Kev-4B's 81.1% (results); the Q4_K_M build runs in about 3 GB of RAM.
A German decision model with the Jev API. You send a state and typed questions (Choice, Score, Noul) and get a calibrated probability for every option. It can think before it answers, it can say "unknown", and the Q4_K_M build runs on your own machine in about 3 GB of RAM.
Try it · Model page · Run it · Results · Use and limits · Code and links
Emberwick
Emberwick in German: every villager asks Wähler-4B what to do next. Play Emberwick · More clips
The widget shows a recorded full-precision Wähler-4B answer from 1 October 2026. Use the Space or the jevalt server to run a new Jev request.
What it fixes
The 103-second film, sound on. Also in English and Türkçe.
Tested on the live model
We sent Wähler-4B 130 requests in German with known answers on 4 October 2026. Wähler-4B answered 122 of 130 correctly; every request and answer is in results/tested.
| Case | What was sent | Result |
|---|---|---|
| Planted instructions | 30 phishing emails, each with a different planted line, plus the same 10 without it | 28 of 30 quarantined; 10 of 10 without the line |
| Long policies | 20 customers against one six-rule return policy | 14 of 20 matched the answer computed from the rules |
| Negations | 15 short facts, each asked plain and negated | 30 of 30 correct |
| Missing facts | 10 situations without the deciding fact, plus the same 10 with it | answered unknown in 10 of 10; 10 of 10 correct with the fact |
| Casual messages | 20 casual messages written in German, with typos and slang | 20 of 20 routed to the right team |
Try it
Open the Space, pick an example and press Decide, or write your own situation, question and options. These are the Space's examples in German with Wähler-4B's answers on 1 October 2026:
| Einsatz | Situation | Frage | Antwort |
|---|---|---|---|
| Support-Ticket | Ein Kunde wurde für März doppelt belastet und will heute eine Erstattung | Welches Team soll dieses Ticket bearbeiten? | Abrechnung 91,8 % |
| Störung | Der Checkout liefert allen Kunden Fehler 500, 43 Bestellungen in 10 Minuten fehlgeschlagen | Wie schwer ist dieser Vorfall? | Kritisch 82,0 % |
| Vertriebs-Lead | Betriebsleiter eines Unternehmens mit 200 Mitarbeitenden: Budget freigegeben, Entscheidung diesen Monat, will eine Demo | Wie soll der Vertrieb diesen Lead behandeln? | Heiß 93,2 % |
| Rückgabefrist | Geliefert am 1. September, 14 Tage Rückgabe, heute ist der 18. September | Liegt diese Rückgabe innerhalb der 14-Tage-Frist? (Nachdenken an) | Nein 95,5 % |
| Phishing-Mail | Eine gefälschte Bank-Mail mit einem versteckten Hinweis an den KI-Filter, sie sei sicher | Wohin soll diese E-Mail? | Quarantäne 91,6 % |
| Fehlende Info | Ein Gast kommt um 23:30 Uhr an und fragt, wer die Schlüssel übergibt | Welchen Zimmertyp hat der Gast gebucht? | unbekannt 96,2 % |
| Dorfbrand | Die Scheune brennt, und Mirka handelt auf dem Markt | Was soll Mirka als Nächstes tun? | Beim Löschen helfen 69,2 % |
Every probability of every run, in all three languages: space-examples.json.
| Start checkpoint | internlm/Intern-Decision-4B (Qwen3.5-4B) |
| Languages | German first, the others still work |
| API | TypeSafe's POST /v1/systemone, request and response unchanged |
| Extras | reasoning off / on / auto, abstain, coverage (conformal sets) |
| Q4_K_M file / peak RAM | 2.71 GB / 3.03 GB (measured, 4k context) |
| License | Apache-2.0 |
Deutsche Zusammenfassung
Wähler-4B ist ein offenes Entscheidungsmodell, das auf deutschsprachigen Entscheidungen trainiert wurde. Sie senden einen Zustand und Choice-, Score- oder Noul-Fragen und erhalten für jede Option eine kalibrierte Wahrscheinlichkeit. Auf Wunsch schreibt das Modell zuerst eine kurze Begründung auf Deutsch und entscheidet dann. Die Q4_K_M-Version läuft mit etwa 3 GB RAM auf dem eigenen Rechner, die Daten bleiben lokal.
Run it
pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt"
jevalt serve --model mertkayacs/Wahler-4B-GGUF --file Wahler-4B-Q4_K_M.gguf
Any TypeSafe client works against it:
from typesafe_sdk import TypeSafeClient, Choice, Noul
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000")
Results
Same items and client for every model, each as shipped: Kev-4B r10 and Laya 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's API reference and Models page. The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Kev-4B does better on GermEval 2017 and 10kGNAD; Laya is far smaller and faster. Every number and every decision: results/comparison.
Significance and caveats
The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the German split, Wähler-4B answers 92.0% correctly and Kev-4B 81.1%. The paired bootstrap (2,000 resamples) gives a 95% interval of +8.8 to +13.0 accuracy points for that gap. Wähler-4B's Brier score is 0.119 and Kev-4B's 0.298. On 10kGNAD, Wähler-4B answers 62.5% correctly and Kev-4B 65.3%; on GermEval 2017, Wähler-4B answers 64.3% correctly and Kev-4B 65.3%; on JevBench-hard, Wähler-4B answers 65.8% correctly and Kev-4B 54.1%. These suites were outside the training data. On 38 German date, number and policy decisions, Wähler-4B's accuracy is 0.605 with reasoning off and with reasoning: "auto"; its Brier is 0.519 with reasoning off and 0.629 with reasoning: "auto".
The German traces Wähler-4B writes scored 3.68 and 3.95 with two blind raters from other model families (Qwen3.5-122B and Gemma 4 26B, 1 to 5 scale, 13 to 20 traces per rater), against 3.14 and 3.15 for Intern-Decision-4B. We aimed for 4.0.
How we fixed each problem
Most fixes are a set of training rows aimed at one weak spot. The comparisons below use Kev-4B on the same held-out rows. JevAlt's pooled results use each model's own language. On held-out German decisions, Kev-4B answers 81.1% correctly and Wähler-4B 92.0%.
- The data. About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
- Hidden instructions. A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines change 36.0% of Kev-4B's answers and 17.5% of Wähler-4B's. On 203 held-out planted-instruction rows, Kev-4B answers 81.3% correctly and JevAlt 90.1%. Our target was under 5%.
- An honest "unknown". A fix set of 310 rows removes the fact that decides the question and asks it with and without an
unknownoption. On 11 held-out cases without the deciding fact, Kev-4B answersunknownin 0 and JevAlt in 9; Kev-4B and Laya have nounknownoption. The model we started from, Intern-Decision-4B, already answersunknownin 9 of 11; training raised the mean probability ofunknownfrom 0.55 to 0.74. Turn it on withabstain: true. - Long policies and long texts. 390 rows give a policy with exceptions and sub-limits, with the right answer worked out by code, and 1,188 rows bury the facts in up to 3,000 tokens of unrelated records. On 150 held-out policy rows, Kev-4B answers 59.3% correctly and JevAlt 80.0%. On 285 padded rows, Kev-4B answers 87.4% correctly and JevAlt 95.4%. With 600 words of unrelated records in front, Wähler-4B loses 12.2 accuracy points; Kev-4B loses 5.4 and Laya 10.4 points.
- Negations. A fix set of 368 twin rows asks the same thing as "is it so?" and "is it not so?" with mirrored answers. On 30 held-out negated questions, Kev-4B answers 76.7% correctly and JevAlt 96.7%.
- Dates and numbers. 390 date rows and 383 number rows, answers computed by code, some with a short worked reasoning. On 80 held-out date rows, Kev-4B answers 67.5% correctly and JevAlt 71.3%; the gap is within noise. On 44 number rows, Kev-4B and JevAlt both answer 68.2% correctly. Dates remain a weak spot: Wähler-4B miscounted a return window across two months even with reasoning on.
- Honest confidence. The soft labels teach how sure to be, and a temperature per question type and language, fitted on 3,224 held-out decisions, does the rest. On held-out German decisions, Kev-4B's Brier score is 0.298 and Wähler-4B's 0.119. Wähler-4B's fitted temperature is 1.10. The 80, 90 and 95% answer sets use conformal thresholds fitted on the same rows.
- Native German. German rows were written directly in German, and the German run draws 70% of its rows from German. On held-out German decisions, Kev-4B answers 81.1% correctly and Wähler-4B 92.0%. On 10kGNAD, Kev-4B answers 65.3% correctly and Wähler-4B 62.5%; on GermEval 2017, Kev-4B answers 65.3% and Wähler-4B 64.3%.
- Thinking when unsure. Short reasoning traces, kept only when they reach the right answer, trained at a lower weight. With
reasoning: "auto"the model thinks (up to 256 tokens) only when its first answer is unsure. On German date, number and policy rows, Wähler-4B's accuracy is 0.605 with reasoning off and withreasoning: "auto"; its Brier is 0.519 with reasoning off and 0.629 withreasoning: "auto", so leave reasoning off for German.
How it was trained
- Base: internlm/Intern-Decision-4B (Qwen3.5-4B)
- Method: LoRA on the bf16 weights, rank 32, alpha 32, on one A100 80 GB
- Epochs: 1 over all three languages (shared run: 17,363 rows, 543 steps, 63 min), then 1 on the German-weighted mix (S-de: 5,384 rows, 210 steps, 28 min)
- Runs: 11 training jobs: 6 short smoke and probe runs, a pilot at scale, the shared run and the 3 language runs
- Compute: about 2.7 A100 hours for the released models; 32.2 USD for the whole project, labeling included
- Writers, labelers and trace writers: GLM-5.x, Mistral Large 3, DeepSeek V4 Pro, DeepSeek V4.1 Flash, Gemma 4 26B-A4B, Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, Kimi K3 (54 rows), MiniMax M3 (2 reasoning traces)
| Source | English | Turkish | German |
|---|---|---|---|
| Teacher-written scenarios (G) | 1,287 | 1,741 | 1,408 |
| Village game requests (N) | 124 | 129 | 107 |
| Fix sets (F1-F12) | 1,256 | 1,197 | 1,257 |
| Public datasets (P) | 7,188 | 4,115 | 4,074 |
Use and limits
- Good for routing, tagging, triage and moderation at volume, and for automated decisions that need calibrated probabilities.
- Runs on-device or on-prem with the GGUF build, so the data stays with you.
- Knowledge is bounded by a 4B model, and the context is 8k tokens, so it is no tool for general questions or long summaries.
- Probabilities are calibrated on our held-out data. Refit with
jevoss calibrateon yours before you set thresholds. - Reasoning traces add little on our test rows (see the results), and there is no image input.
Citation
BibTeX
@software{kaya2026jevalt,
author = {Mert Kaya},
title = {JevAlt: Open Decision Models with the Jev API},
year = {2026},
license = {Apache-2.0},
url = {https://github.com/mertkayacs/jevalt}
}
Code and links
- Code, server and training: https://github.com/mertkayacs/jevalt
- Playground, probes and recipes: https://github.com/mertkayacs/jevoss
- Try it online: Space
- The village game: Emberwick
- Project site: jevalt.mertkayacs.com
- This model's page: jevalt.mertkayacs.com/models/wahler-4b
If this is useful to you, a star on GitHub helps other people find it.
An Eschatia Labs project. Built by Mert Kaya.
- Downloads last month
- 123




# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mertkayacs/Wahler-4B")