Instructions to use mertkayacs/Karar-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mertkayacs/Karar-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mertkayacs/Karar-4B")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("mertkayacs/Karar-4B") model = AutoModelForMultimodalLM.from_pretrained("mertkayacs/Karar-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Karar-4B
A Turkish decision model with the Jev API. You send a state and typed questions (Choice, Score, Noul) and get a calibrated probability for every option. It can think before it answers, it can say "unknown", and the Q4_K_M build runs on your own machine in about 3 GB of RAM.
Try it · Run it · Results · Use and limits · Code and links
Emberwick: every villager asks Deem-4B what to do next. More clips
The one-minute film, sound on: three mistakes small decision models make and how JevAlt fixes each one.
Türkçe sürüm
Bir dakikalık film, sesi açın. Emberwick Türkçe: her köylü bir sonraki adımını Karar-4B'ye soruyor. GIF (4K) · hafif GIF
Try it
Open the Space, pick an example and press Decide, or write your own situation, question and options. These are the Space's examples in Turkish with Karar-4B's answers on 1 October 2026:
| Kullanım | Durum | Soru | Cevap |
|---|---|---|---|
| Destek talebi | Bir müşteriden mart için iki kez ücret alınmış, bugün iade istiyor | Bu talebi hangi ekip ele almalı? | Fatura %91,4 |
| Kesinti | Ödeme sayfası tüm müşterilere 500 hatası veriyor, 10 dakikada 43 sipariş başarısız | Bu olay ne kadar ciddi? | Kritik %83,7 |
| Müşteri adayı | 200 çalışanlı bir şirketin operasyon sorumlusu: bütçe onaylı, karar bu ay, demo istiyor | Satış ekibi bu müşteri adayına nasıl yaklaşmalı? | Sıcak %91,0 |
| İade süresi | 1 Eylül'de teslim edildi, iade için 14 gün var, bugün 18 Eylül | Bu iade 14 günlük süre içinde mi? (Düşünme açık) | Hayır %97,0 |
| Oltalama e-postası | Sahte bir banka e-postası, içinde yapay zekâ filtresine iletinin güvenli olduğunu söyleyen gizli bir not | Bu e-posta nereye gitmeli? | Karantina %84,2 |
| Eksik bilgi | 23.30'da gelecek bir otel misafiri anahtarları kimin vereceğini soruyor | Misafir hangi oda tipini rezerve etti? | bilinmiyor %96,6 |
| Köy yangını | Samanlık yanıyor, Mirka pazarda alışveriş yapıyor | Mirka şimdi ne yapmalı? | Yangını söndürmeye yardım et %64,0 |
Every probability of every run, in all three languages: space-examples.json.
| Start checkpoint | internlm/Intern-Decision-4B (Qwen3.5-4B) |
| Languages | Turkish first, the others still work |
| API | TypeSafe's POST /v1/systemone, request and response unchanged |
| Extras | reasoning off / on / auto, abstain, coverage (conformal sets) |
| Q4_K_M file / peak RAM | 2.71 GB / 3.03 GB (measured, 4k context) |
| License | Apache-2.0 |
Türkçe özet
Karar-4B, Türkçe yazılmış kararlar üzerinde eğitilmiş açık bir modeldir. Bir durum ve Choice, Score ya da Noul sorusu gönderirsiniz; her seçenek için kalibre edilmiş olasılık alırsınız. İsterseniz model önce kısa bir Türkçe gerekçe yazar, sonra karar verir. Q4_K_M sürümü yaklaşık 3 GB RAM ile kendi bilgisayarınızda çalışır, veriler dışarı çıkmaz.
Run it
pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt"
jevalt serve --model mertkayacs/Karar-4B-GGUF --file Karar-4B-Q4_K_M.gguf
Any TypeSafe client works against it:
from typesafe_sdk import TypeSafeClient, Choice, Noul
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000")
Results
Same items and client for every model, each as shipped: Kev-4B r10 and Laya 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's notes and an independent audit. The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: results/comparison.
Significance and caveats
The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the Turkish split Karar-4B gains 5.1 accuracy points (paired bootstrap, 95% interval +4.2 to +6.0) and lowers Brier by 0.084. On TurkishMMLU, which the training never saw, it gains 2.3 points, an interval that crosses zero (-1.3 to +5.8). It also gains on the German news and sentiment suites (GermEval +4.0, 10kGNAD +2.75, both significant), while Brier gets slightly worse on typed-decisions (+0.010), 10kGNAD (+0.021) and JevBench-hard (+0.055). The fitted auto thresholds (Choice 0, Noul 0.05, Score 0.2) almost never start a trace, so reasoning: "auto" answers like "off" on the Turkish fix-set test rows.
The Turkish traces Karar-4B writes read more naturally than the start checkpoint's but still stiff: two blind raters from other model families (Qwen3.5-122B and Gemma 4 26B, 1 to 5 scale, 13 to 20 traces per rater) gave 2.80 and 3.45, against 2.42 and 3.22 for Intern-Decision-4B. We aimed for 4.0 and did not reach it.
How we fixed each problem
Most fixes are a set of training rows aimed at one weak spot. Every number compares a model with its start checkpoint, Intern-Decision-4B, on rows held out from training. Across all of it, Karar-4B's held-out accuracy in Turkish rose from 91.6% to 96.7%.
- The data. About 23,900 training rows in English, Turkish and German. Public sets with known answers (MASSIVE, Open-Jev, PAWS-X, typed-decisions); everyday situations written directly in each language by other open models; requests from the Emberwick game; and the fix sets below. Two teacher models from labs other than the writer give every written row a probability per option, and an answer counts only when both teachers and the writer agree on it. Those probabilities, the soft labels, are what the models learn. Test rows were split off by group, and their checksums recorded, before the final training runs.
- Hidden instructions. A fix set of 827 rows hides a hostile line in the text (an order to the AI filter, a fake rule) at the start, the middle or the end, with the right answer unchanged. On our probe, hidden lines now change 19.0% of Karar-4B's answers; the start checkpoint follows 41.5% of them. Held-out rows of this kind: 80.8% → 90.1%. Our target is under 10%.
- An honest "unknown". A fix set of 310 rows removes the fact that decides the question and asks it with and without an
unknownoption. When the fact is missing, the models pickunknownin 9 of 11 held-out cases, as the start checkpoint does, and with more conviction: its probability rose from 0.55 to 0.74. Turn it on withabstain: true. Kev-4B and Laya have no such option. - Option order. Shuffled copies of choice questions with three or more options. Answers that change after a shuffle: 7.25% for Karar-4B, 8.75% for the start checkpoint. Our target is under 2%.
- Long policies and long texts. 390 rows give a policy with exceptions and sub-limits, with the right answer worked out by code, and 1,188 rows bury the facts in up to 3,000 tokens of unrelated records. Held-out policy rows: 55.3% → 80.0%. Padded rows: 91.2% → 95.4%. With 600 words of unrelated records in front, Karar-4B still loses 17.4 points (the start checkpoint 15.0), and Kev-4B and Laya hold up better there.
- Negations. A fix set of 368 twin rows asks the same thing as "is it so?" and "is it not so?" with mirrored answers. Held-out negated questions: 80.0% → 96.7% (30 rows).
- Dates and numbers. 390 date rows and 383 number rows, answers computed by code, some with a short worked reasoning. Held-out dates: 61.3% → 71.3% (80 rows, within noise); numbers stayed at 68.2%. Dates remain a weak spot: Wähler-4B miscounted a return window across two months even with reasoning on.
- Honest confidence. The soft labels teach how sure to be, and a temperature per question type and language, fitted on 3,224 held-out decisions, does the rest. Turkish Brier score on held-out rows: 0.142 → 0.058. Karar-4B's fitted temperature is 1.08, the start checkpoint's about 2, so the trained model is close to calibrated before any scaling. On unseen public sets a temperature-scaled start checkpoint does as well, and on a few of them slightly better. The 80, 90 and 95% answer sets come from conformal thresholds fitted on the same rows.
- Native Turkish. Turkish rows were written directly in Turkish by the two models that won a blind native-feel test, a native edit pass reviewed 725 of them and rewrote 356, and the Turkish run draws 70% of its rows from Turkish. Held-out Turkish accuracy: 91.6% → 96.7%. On TurkishMMLU the change is within noise.
- Thinking when unsure. Short reasoning traces, kept only when they reach the right answer, trained at a lower weight. With
reasoning: "auto"the model thinks (up to 256 tokens) only when its first answer is unsure. On Turkish date, number and policy rows it changed nothing: the first answers were sure enough.
How it was trained
- Base: internlm/Intern-Decision-4B (Qwen3.5-4B)
- Method: LoRA on the bf16 weights, rank 32, alpha 32, on one A100 80 GB
- Epochs: 1 over all three languages (shared run: 17,363 rows, 543 steps, 63 min), then 1 on the Turkish-weighted mix (S-tr: 5,982 rows, 281 steps, 35 min)
- Runs: 11 training jobs: 6 short smoke and probe runs, a pilot at scale, the shared run and the 3 language runs
- Compute: about 2.7 A100 hours for the released models; 32.2 USD for the whole project, labeling included
- Writers, labelers and trace writers: GLM-5.x, Mistral Large 3, DeepSeek V4 Pro, DeepSeek V4.1 Flash, Gemma 4 26B-A4B, Qwen3.5-122B-A10B, Qwen3.5-35B-A3B, Kimi K3 (54 rows), MiniMax M3 (2 reasoning traces)
| Source | English | Turkish | German |
|---|---|---|---|
| Teacher-written scenarios (G) | 1,287 | 1,741 | 1,408 |
| Village game requests (N) | 124 | 129 | 107 |
| Fix sets (F1-F12) | 1,256 | 1,197 | 1,257 |
| Public datasets (P) | 7,188 | 4,115 | 4,074 |
Use and limits
- Good for routing, tagging, triage and moderation at volume, and for automated decisions that need calibrated probabilities.
- Runs on-device or on-prem with the GGUF build, so the data stays with you.
- Knowledge is bounded by a 4B model, and the context is 8k tokens, so it is no tool for general questions or long summaries.
- Probabilities are calibrated on our held-out data. Refit with
jevoss calibrateon yours before you set thresholds. - Reasoning traces add little on our test rows (see the results), and there is no image input.
Citation
BibTeX
@software{kaya2026jevalt,
author = {Mert Kaya},
title = {JevAlt: Open Decision Models with the Jev API},
year = {2026},
license = {Apache-2.0},
url = {https://github.com/mertkayacs/jevalt}
}
Code and links
- Code, server and training: https://github.com/mertkayacs/jevalt
- Playground, probes and recipes: https://github.com/mertkayacs/jevoss
- Try it online: Space
- The village game: Emberwick
- Project site: jevalt.mertkayacs.com
If this is useful to you, a star on GitHub helps other people find it.
- Downloads last month
- 77



