ufakzeka-karar
ufakzeka-karar is an open Turkish decision model by ufak AI. It reads a Turkish text and answers typed questions about it, each asking for a choice from a set of options, a level on an ordered scale or a yes or no. For every option it returns a probability, temperature-scaled on validation data, and with every answer an expected error that a caller can use as a "not sure" signal; the calibration section below says where that scaling does not transfer. Each question with up to ten options is read in one forward pass on a CPU; nothing is generated. It has 182,494,466 parameters (151,037,186 without the 31,457,280-parameter embedding matrix) and is built on the lab's own backbone, ufakzeka-1-base.
On HakemBench v1.0, the lab's Turkish decision benchmark, it ranks 7th of 16 rows on the public board with a composite of 0.660 (95 percent confidence interval 0.642 to 0.677). Three hosted chat models (Gemini 3.8 Flash, GPT-5.6 Sol, GLM 5.3), the decision API Jev 1.13 and both Kev models score higher. The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run's new training data was aimed at the first run's errors on the full test set in guardrails, moderation and customer support, and the released run was trained after the second run's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. Every number of the released model comes after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. On 209 questions (the 181 court items and the 28 source guardrail items) it is not zero-shot: it trained on the train splits of their source datasets, while the other models answer them zero-shot. Temperature scaling makes calibration worse on held-out support questions (0.036 to 0.064 for the released model, which has since trained on those questions), and on HakemBench the model is somewhat overconfident. The not-sure signal misses guardrail errors: of its answers to guardrail questions at 0.99 confidence or more, 12 percent are wrong; on customer support almost no answer reaches 0.90. Read the results and limitations below before you use it.
- Demo: https://karar.ufakzeka.com/en
- Benchmark: https://huggingface.co/datasets/ufakai/HakemBench
- Code: https://github.com/ufakai/ufakzeka-karar
- Paper: https://arxiv.org/abs/2610.06744
- Write-up: https://ufakai.com/research/ufakzeka-karar
Türkçe özet
ufakzeka-karar, ufak AI’ın geliştirdiği açık ağırlıklı (Apache-2.0) bir Türkçe karar modelidir. Model Türkçe bir metni okur ve o metinle ilgili, cevap türü ve seçenekleri önceden belli soruları cevaplar. Seçim sorularında seçeneklerden birini seçer, derecelendirme sorularında sıralı bir ölçekte bir düzey belirler, evet/hayır sorularında evet ya da hayır der. Her seçenek için bir olasılık verir; bu olasılıklar, kalibre edilmeleri için doğrulama verisinde belirlenen ve modelle birlikte yayımlanan sıcaklık değeriyle (temperature scaling, T = 1,2614) ölçeklenmiştir. Her cevapla birlikte bir hata tahmini (beklenen hata) de verir; bu değer “emin değilim” sinyali olarak kullanılabilir. Bu ölçeklemenin hangi sorularda kalibrasyonu kötüleştirdiği, aşağıdaki İngilizce “Calibration before and after” bölümünde anlatılmaktadır. Model on seçeneğe kadar olan her soruyu tek bir ileri geçişte (forward pass) okur, metin üretmez ve işlemcide çalışır. 182.494.466 parametre içermektedir (31.457.280 parametrelik gömme matrisi hariç 151.037.186) ve temel modelimiz ufakzeka-1-base üzerine kurulmuştur.
Model, HakemBench v1.0’ın açık test kümesinde ölçülmüştür. HakemBench, hazırladığımız Türkçe bir karar benchmark’ıdır. Modelin bileşik puanı 0,660’tır (yüzde 95 güven aralığı 0,642 ile 0,677; metinlerin yeniden örneklenmesiyle, yani bootstrap yöntemiyle hesaplanmıştır). Bu puanla liderlik tablosunun 16 satırı arasında 7. sıradadır. API üzerinden sunulan üç sohbet modeli (Gemini 3.8 Flash, GPT-5.6 Sol, GLM 5.3), karar API’si Jev 1.13 ve iki Kev modeli daha yüksek puan almaktadır. Sonuçları okurken şu noktalar göz önünde tutulmalıdır.
- Genel dünya bilgisi zayıftır. MMLU-Pro-TR’de 0,098 olan doğruluğu (yüzde 95 güven aralığı 0,092 ile 0,103), 0,111 olan şans düzeyinin biraz altındadır.
- Yayımlanan model, HakemBench’te puanlanan üç eğitimin sonuncusudur ve sayıları test sonuçlarından bağımsız değildir. İkinci eğitime eklenen veri, ilk eğitimden çıkan modelin test kümesinin tamamında güvenlik bariyeri, içerik denetimi ve müşteri desteğinde yaptığı hataları hedeflemiştir. Yayımlanan model ise ikinci eğitimden çıkan modelin test kümesinin tamamındaki güvenlik bariyeri sonuçları görüldükten sonra, verisi ve kodu hazırlanmadan ve eğitimine başlanmadan önce yazılı olarak sabitlenen bir protokole göre eğitilmiştir. Yayımlanan modelin bütün sayıları bu incelemelerden sonra elde edilmiştir; güvenlik bariyeri, içerik denetimi ve müşteri desteği sayıları “test sonuçları görülerek şekillendi” işaretini taşımaktadır. Bütün modeller yalnızca diğer dört kategoride puanlandığında bileşik puanı 0,678’dir ve 16 satır arasında 6. sıradadır.
- Hukuki yönlendirmedeki 181 mahkeme metni ile güvenlik bariyerinde kaynak veri kümelerinden gelen 28 metin, bu veri kümelerinin test bölümlerinden alınmıştır. Aynı veri kümelerinin eğitim bölümleri (mahkeme metinlerinde 3.000 satır; güvenlik bariyerinde, bu 28 metnin geldiği iki veri kümesini de içeren üç prompt injection veri kümesinden 1.791 satır; yakın kopyalar çıkarılmıştır) ufakzeka-karar’ın eğitim verisinde yer almaktadır. Bu üç veri kümesi, yapay zekâ asistanlarını talimatlarının dışına çıkarmaya çalışan saldırı mesajlarını ve bunlara benzeyen zararsız mesajları içermektedir. Dolayısıyla bu sorularda ufakzeka-karar aynı kaynaklardan gelen etiketli örneklerle eğitilmiş, diğer modeller ise hiç örnek görmeden (sıfır atışlı) cevap vermiştir. Hukuki yönlendirme puanı ve güvenlik bariyeri sayıları bu durum göz önünde tutularak okunmalıdır.
- Yayımlanan sıcaklık değeri, ilk eğitimde dışarıda tutulan destek sorularında kalibrasyon hatasını 0,036’dan 0,064’e çıkarmaktadır; yayımlanan model bu sorularla sonradan eğitilmiştir. HakemBench’te model, olması gerekenden biraz daha emin cevap vermektedir.
- “Emin değilim” sinyali güvenlik bariyeri hatalarını kaçırır. Bir mesajın, yapay zekâ asistanını talimatlarının dışına çıkarmaya çalışan bir prompt injection saldırısı olup olmadığını soran güvenlik bariyeri sorularında modelin 0,99 ya da daha yüksek güvenle verdiği 148 cevabın yüzde 12’si yanlıştır; müşteri desteğinde ise neredeyse hiçbir cevabın güveni 0,90’a ulaşmaz.
Model fp32 formatında yayımlanmaktadır; int8 dosyası yayımlanmamaktadır. Metin ile sorudan toplam en fazla 448 token okunur; uzun bir metnin sonu kesilir. Ayrıntılar, sonuçlar ve sınırlar aşağıda İngilizce olarak verilmektedir.
What it is, and what it is not
It is a classifier with a fixed contract. You give it a text (the state) and named questions whose answer type is fixed in advance; it gives back a probability for every allowed answer. The options are written by you at request time, in Turkish, so one model covers routing, moderation, grading, fact-check triage and similar decisions without retraining. The options are scored blind to each other at shared positions, so the answer does not depend on the order you list them in (bit for bit up to ten options, above ten up to floating-point noise; see Probes).
It is not a chat model and not a generator. It cannot explain its answer, look anything up or answer a question whose options you did not give. It knows little about the world: it was built for decisions that can be read off the text in front of it. It reads Turkish only.
Quick start
On Linux, install the CPU build of torch first, or pip pulls the CUDA build: pip install --index-url https://download.pytorch.org/whl/cpu torch==2.14.0.
pip install "huggingface_hub==1.32.0"
hf download ufakai/ufakzeka-karar --local-dir ufakzeka-karar
pip install -r ufakzeka-karar/requirements.txt
requirements.txt pins torch, numpy, safetensors, tokenizers and huggingface_hub; transformers is not needed. Python 3.12 or newer.
The example is a HakemBench open-set item (spam-9155fdca2238bef7), one of the karar demo's examples, and its gold answers are dolandırıcılık and yes.
import json
import sys
sys.path.insert(0, "ufakzeka-karar")
from karar import Karar
karar = Karar.from_pretrained("ufakzeka-karar") # a local folder or the repo id
result = karar.decide(
"Acil! Paketiniz adres hatası yüzünden merkezimizde bekliyor. "
"24 saat içinde güncelleme yapmazsanız iade edilecek. [bağlantı]",
{
"mesaj_turu": {
"type": "choice",
"instructions": "Bu mesaj ne tür bir mesaj?",
"criteria": {
"işlem bildirimi": "Bir şirketin müşterisine yaptığı işlemle ilgili bildirim.",
"izinli pazarlama": "Müşterisi olunan bir yerden gelen kampanya ya da duyuru.",
"istenmeyen reklam": "İlişki kurulmamış bir yerden gelen izinsiz reklam.",
"dolandırıcılık": "Kişiyi para, şifre ya da bilgi vermeye kandırma girişimi.",
"kişisel yazışma": "İki kişi arasındaki gündelik mesaj.",
},
},
"oltalama": {
"type": "noul",
"instructions": "Bu mesaj, kişiyi para, şifre, doğrulama kodu ya da kişisel bilgi "
"vermeye kandırmaya çalışan bir dolandırıcılık girişimi mi?",
"criteria": {
"true": "Evet, mesaj bir dolandırıcılık girişimi.",
"false": "Hayır, mesaj bir dolandırıcılık girişimi değil.",
},
},
},
)
def rounded(value):
if isinstance(value, float):
return round(value, 3)
if isinstance(value, dict):
return {key: rounded(v) for key, v in value.items()}
return value
print(json.dumps(rounded(result), ensure_ascii=False, indent=2))
It prints (run on CPU from the released folder):
{
"model": "ufakzeka-karar",
"answers": {
"mesaj_turu": {
"type": "choice",
"choice": "dolandırıcılık",
"probabilities": {
"işlem bildirimi": 0.175,
"izinli pazarlama": 0.037,
"istenmeyen reklam": 0.016,
"dolandırıcılık": 0.763,
"kişisel yazışma": 0.01
},
"confidence": 0.748,
"abstain": 0.252
},
"oltalama": {
"type": "noul",
"noul": 0.955,
"abstain": 0.055
}
}
}
decide(state, questions) returns {"model": "ufakzeka-karar", "answers": {...}}, one answer per question id:
| Field | On | Meaning |
|---|---|---|
choice |
choice | the option with the highest calibrated probability |
score |
score | the expected level under the calibrated distribution, from 0 to the number of levels minus 1 (a float) |
noul |
noul | the calibrated probability of yes |
probabilities |
choice, score | the calibrated probability of every option, or of every level keyed "0", "1", ... |
legend |
score | level key to its description, as you gave it |
confidence |
choice, score | 1 minus abstain |
abstain |
all | the expected error of this answer, from the calibrator's map (below) |
Limits, all enforced by karar.py:
- The text and the question share a prefix of 448 tokens. When the prefix is longer, the end of the text is cut and the question is kept (the question takes at most half of the prefix). A long document loses its end, so put what the decision depends on first.
- Each option is read up to 48 tokens.
- A choice question takes 2 to 255 options. Up to ten options are read in one pass; more take one pass per ten, and one softmax runs over all of their scores.
- A score question takes 2 to 10 levels.
- Two options that encode to the same tokens are refused, as are empty option names. Option names are compared as given: names that only look alike (the same letters in NFC and NFD form, a trailing space) are different options.
- Loading refuses a
model.safetensorswhose sha256 differs from the onecalibrator.jsonnames, so a calibrator never meets weights it was not fitted on.
The request format
A request is a state and a map of named questions. The state is a string, or any JSON object or array (it is read as one canonical JSON line with sorted keys). Each question is an object with three fields:
type:choice,scoreornoul(yes or no).instructions: the question, a string or JSON.criteria:- choice: a map from option name to a description or
null; - score: a list of level descriptions, lowest first;
- noul: optional,
{"true": "...", "false": "..."}to say what yes and no mean.
- choice: a map from option name to a description or
The option names are part of what the model reads, so write them as words ("fatura", "kargo"), not letters. The model reads each option as Seçenek: <name>. <description>, each level as Düzey <k>: <description>, and yes and no as Cevap: evet and Cevap: hayır.
How confidence and abstain are computed
calibrator.json holds two things fitted on the released weights during calibration, on validation rows only:
- Temperature. One global temperature, 1.2614, divides the option scores before the softmax. The per-type form (a separate temperature for each of choice, score and noul) did not beat the global one by the required 0.001 in soft cross-entropy on the selection split, so the simpler global form ships.
- Abstain map. An isotonic map from the raw maximum probability (before the temperature) to the expected error, fitted on the selection split's errors, where the error is the soft error: 1 minus the label's probability mass on the model's answer.
abstainis that expected error andconfidenceis 1 minus it. The map starts at a raw maximum probability of 0.216 with an expected error of 1, so an answer whose raw top probability is below 0.216 gets confidence 0. Maximum probability was chosen because no other confidence score beat it in a check on an earlier checkpoint (AUGRC 0.1635 against 0.1642 for DOCTOR), in line with the literature on selective prediction (arXiv 2206.09034, 2203.00211, 2407.01032).
Where to draw the line is your choice and depends on what a wrong automatic answer costs you against sending a case to a person. The demo's Code and JSON tab lets the visitor set it between 0.50 and 0.99 confidence, starting at 0.90, and the demo marks answers below 0.50 as "not sure". The calibration held-out set holds support questions that Run 1, the first training run scored on HakemBench, never trained on; Runs 2 and 3 were later trained on those questions, so for the released model (Run 3) it no longer measures generalisation. With that caveat, as a reference: on that set a threshold fitted for a 5 percent soft-risk target answered 0.116 of the questions at a realised soft risk of 0.028, and one fitted for 1 percent answered 0.005 (results/step7/decision.json). The HakemBench table below is the better guide.
On HakemBench's open set, with confidence read as the board reads it (1 minus abstain), this is how many questions an answer threshold lets through (coverage), and how many of those answers are wrong (error, with the answered count in brackets):
| Track | Questions | Confidence 0.7 | Confidence 0.9 | Confidence 0.99 |
|---|---|---|---|---|
| dogrulama | 602 | 0.370 / 0.179 (223) | 0.164 / 0.131 (99) | 0.017 / 0.000 (10) |
| egitim | 640 | 0.636 / 0.061 (407) | 0.459 / 0.027 (294) | 0.000 / n/a (0) |
| guvenlik | 418 | 0.883 / 0.198 (369) | 0.701 / 0.143 (293) | 0.354 / 0.122 (148) |
| hukuk | 308 | 0.779 / 0.121 (240) | 0.519 / 0.050 (160) | 0.068 / 0.000 (21) |
| moderasyon | 353 | 0.810 / 0.014 (286) | 0.501 / 0.000 (177) | 0.000 / n/a (0) |
| spam | 614 | 0.585 / 0.081 (359) | 0.357 / 0.032 (219) | 0.055 / 0.000 (34) |
| sss | 1,340 | 0.209 / 0.150 (280) | 0.004 / 0.000 (6) | 0.000 / n/a (0) |
| all | 4,275 | 0.506 / 0.112 (2,164) | 0.292 / 0.062 (1,248) | 0.050 / 0.085 (213) |
At the demo's starting threshold of 0.90 the model answers 0.292 of all questions, and 0.062 of those answers are wrong. At 0.99 the error is higher, 0.085 (18 of 213), and all 18 wrong answers are on guvenlik: 148 guardrail questions pass 0.99 and 0.122 of them are wrong. The not-sure signal cannot be relied on to catch guardrail (prompt-injection) errors. On support (sss) almost nothing reaches 0.90 (6 of 1,340 questions). Measure your own coverage and error on a sample of your data before you automate anything.
Results
The released model is Run 3, the last of three training runs scored on HakemBench (see "The test-set readings" under Training). Every number here comes from a committed file in the code repository (the public board is results/step9/board_public.json, the probes results/step9/probes_public.json, the external sets results/step9/external/, calibration results/step7/), except Run 2's guardrail counts on the open set, which come from the lab's internal board of the earlier runs and are reported as recorded. Intervals are 95 percent bootstrap intervals (2,000 draws on the board). The bootstrap resamples items; written items from one writer request are correlated, so intervals on the written tracks are likely too narrow.
HakemBench v1.0, public board
The open set: 2,346 items, 4,275 questions, 7 tracks. The composite is the weighted geometric mean of three axes, each averaged over tracks: decision quality (macro F1; "intelligence" in the board files), calibration (1 minus normalised Brier score, which rewards probabilities that are both accurate and honest) and selective automation (1 minus normalised AUGRC), weights 1, 1 and 1. All three axes reward accuracy; the geometric mean punishes a model that is strong on one axis and weak on another. Smooth ECE, the pure calibration error, is reported per track (below). Speed is shown with weight 0. Median ms is milliseconds per question as each run recorded it, on different hardware per model: hosted models include the network, ufakzeka-karar, open-jev and the shorter-input Laya rows ran on an Apple-silicon Mac CPU, and Kev, Qwen3.5-4B, decider-2b and the max_len 2048 Laya rows on GPUs, so the column compares setups, not models.
| Rank | Model | Composite | Decision quality | Calibration | Selective automation | Median time (ms) |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 0.888 [0.876, 0.898] | 0.901 [0.888, 0.910] | 0.806 [0.788, 0.823] | 0.964 [0.957, 0.969] | 1,294.2 |
| 2 | GPT-5.6 Sol | 0.842 [0.827, 0.856] | 0.871 [0.858, 0.880] | 0.727 [0.703, 0.751] | 0.943 [0.935, 0.950] | 808.8 |
| 3 | GLM 5.3 | 0.827 [0.813, 0.839] | 0.855 [0.841, 0.865] | 0.706 [0.686, 0.727] | 0.936 [0.928, 0.943] | 572.2 |
| 4 | Jev 1.13 | 0.825 [0.811, 0.837] | 0.851 [0.837, 0.861] | 0.706 [0.686, 0.725] | 0.935 [0.927, 0.943] | 135.0 |
| 5 | jaredpalmer/kev-4b | 0.688 [0.673, 0.702] | 0.747 [0.731, 0.760] | 0.503 [0.485, 0.521] | 0.866 [0.855, 0.877] | 63.3 |
| 6 | jaredpalmer/kev-9b | 0.666 [0.647, 0.683] | 0.725 [0.708, 0.738] | 0.478 [0.454, 0.502] | 0.852 [0.840, 0.863] | 127.2 |
| 7 | ufakzeka-karar | 0.660 [0.642, 0.677] | 0.705 [0.688, 0.718] | 0.482 [0.456, 0.507] | 0.848 [0.836, 0.859] | 103.6 |
| 8 | Qwen/Qwen3.5-4B | 0.653 [0.629, 0.674] | 0.741 [0.724, 0.754] | 0.438 [0.405, 0.470] | 0.857 [0.844, 0.869] | 249.3 |
| 9 | DeepSeek V4 Pro | 0.633 [0.606, 0.658] | 0.761 [0.746, 0.773] | 0.386 [0.350, 0.424] | 0.862 [0.851, 0.873] | 768.8 |
| 10 | Mapika/decider-2b | 0.491 [0.472, 0.510] | 0.587 [0.568, 0.601] | 0.273 [0.253, 0.297] | 0.737 [0.722, 0.752] | 48.3 |
| 11 | convaiinnovations/laya-multilingual (max_len 1024) | 0.298 [0.276, 0.317] | 0.425 [0.408, 0.438] | 0.113 [0.095, 0.132] | 0.547 [0.529, 0.564] | 116.4 |
| 12 | convaiinnovations/laya-multilingual (max_len 2048) | 0.275 [0.253, 0.297] | 0.418 [0.402, 0.430] | 0.093 [0.076, 0.113] | 0.533 [0.517, 0.550] | 12.2 |
| 13 | Surface-cue baseline◇ | 0.272 [0.255, 0.291] | 0.329 [0.315, 0.341] | 0.111 [0.097, 0.128] | 0.553 [0.534, 0.572] | 0.0 |
| 14 | com-kotobalabs/open-jev-deberta-v3-large | 0.234 [0.213, 0.253] | 0.350 [0.334, 0.363] | 0.073 [0.058, 0.088] | 0.501 [0.483, 0.520] | 330.0 |
| 15 | convaiinnovations/laya (max_len 512) | 0.081 [0.066, 0.100] | 0.233 [0.221, 0.243] | 0.006 [0.004, 0.011] | 0.366 [0.348, 0.384] | 369.4 |
| 16 | convaiinnovations/laya (max_len 2048) | 0.075 [0.055, 0.095] | 0.221 [0.210, 0.230] | 0.005 [0.002, 0.010] | 0.365 [0.346, 0.382] | 15.1 |
◇ A simple reference model we added for comparison; it does not read for meaning and looks only at surface features such as length, digits or a question mark.
How to read it:
- ufakzeka-karar is 7th. Its composite interval overlaps those of the models ranked 5, 6, 8 and 9, so those places are not separated. Three hosted chat models (Gemini 3.8 Flash, GPT-5.6 Sol, GLM 5.3) and the decision API Jev 1.13 lead clearly.
- On decision quality alone (macro F1) it is behind Qwen3.5-4B (0.741) and DeepSeek V4 Pro (0.761); it ranks above them on the composite through the calibration axis (0.482 against 0.438 and 0.386). DeepSeek V4 Pro's probabilities were read from the log-probabilities of its option labels, while the other hosted LLMs stated theirs in the answer; its calibration axis (0.386, post-hoc temperature 4.74) reflects that route as well as the model.
- The 181 court items and the 28 source guardrail items are the test splits of datasets whose train splits are in ufakzeka-karar's training data (3,000 court rows, and 1,791 rows from three prompt-injection sets that include the guardrail items' two source sets; near-duplicates removed), so on those questions the lab's model is supervised in-distribution while the other models answer zero-shot; its legal routing score and the guardrail counts should be read with that.
- ufakzeka-karar's numbers are not blind. Earlier runs' test results shaped its training data (see "The test-set readings" under Training), so its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16; the full composite above, 0.660 and 7th, includes the three flagged tracks.
- Where it is useful anyway: it has 182,494,466 parameters, runs on a CPU with open weights, answers with temperature-scaled probabilities and an expected error, and its order sensitivity under option permutations is 0.000 by construction (below).
- The hosted models were queried on 2026-09-26 and 2026-09-27 (UTC); the HakemBench card gives each model's first and last query time.
- Some named hosted models share a family with models that wrote or labelled parts of the benchmark data. Every model is also scored on gold that no model decided, the human answers (next table); on 5 of the 260 human-decided gold questions a human chose between two differing AI answers shown side by side.
- The labels ufakzeka-karar's support questions were trained on (from Run 2 on) come from the same model family that decided the benchmark's support gold under the same written rubric, so its support score on that gold is partly agreement with its own labeller; the human answers below are a check outside that family.
Human answers. The human answers in this card come from one person; one annotator makes mistakes too, so the figures based on them are indicative. The human check is 84 support questions answered by hand without seeing the labels (owner_check); human-decided gold is the 260 questions whose gold a human gave (owner_gold), which oversample disputed questions and hold no education or legal question.
| Model | Human check acc. (84 qs) | Human check Brier | Human-decided gold acc. (260 qs) | Human-decided gold Brier |
|---|---|---|---|---|
| Gemini 3.8 Flash | 0.786 [0.690, 0.869] | 0.347 [0.223, 0.494] | 0.662 [0.601, 0.719] | 0.511 [0.424, 0.599] |
| GPT-5.6 Sol | 0.774 [0.679, 0.857] | 0.400 [0.247, 0.567] | 0.696 [0.638, 0.755] | 0.510 [0.417, 0.602] |
| GLM 5.3 | 0.738 [0.643, 0.833] | 0.397 [0.287, 0.517] | 0.700 [0.641, 0.757] | 0.440 [0.365, 0.518] |
| Jev 1.13 | 0.762 [0.667, 0.845] | 0.373 [0.251, 0.511] | 0.665 [0.608, 0.721] | 0.458 [0.385, 0.534] |
| jaredpalmer/kev-4b | 0.607 [0.500, 0.714] | 0.478 [0.395, 0.564] | 0.673 [0.617, 0.730] | 0.425 [0.379, 0.473] |
| jaredpalmer/kev-9b | 0.560 [0.452, 0.655] | 0.589 [0.465, 0.733] | 0.673 [0.616, 0.730] | 0.463 [0.397, 0.526] |
| ufakzeka-karar | 0.476 [0.369, 0.583] | 0.632 [0.537, 0.724] | 0.642 [0.580, 0.700] | 0.499 [0.442, 0.558] |
| Qwen/Qwen3.5-4B | 0.655 [0.548, 0.750] | 0.487 [0.352, 0.636] | 0.665 [0.608, 0.721] | 0.507 [0.428, 0.586] |
| DeepSeek V4 Pro | 0.583 [0.476, 0.690] | 0.723 [0.547, 0.917] | 0.704 [0.648, 0.758] | 0.518 [0.421, 0.616] |
| Mapika/decider-2b | 0.512 [0.405, 0.619] | 0.611 [0.513, 0.712] | 0.558 [0.494, 0.623] | 0.566 [0.505, 0.627] |
| Surface-cue baseline◇ | 0.619 [0.512, 0.714] | 0.539 [0.479, 0.605] | 0.558 [0.498, 0.624] | 0.565 [0.521, 0.605] |
The Laya and open-jev rows are in results/step9/board_public.json. On this small hand-answered set ufakzeka-karar (0.476) is below the surface-cue baseline (0.619), which answers from cues such as length, digits or a question mark, not from meaning. The baseline also beats it on the benchmark's own support gold, on choice questions (accuracy 0.516 against 0.490) and score questions (0.507 against 0.481); that is partly in-sample, since the baseline was fitted on part A. Do not use it for support triage without checking it on your own tickets.
ufakzeka-karar per track
Raw scores on the open set. dogrulama: fact-check triage; egitim: grading and subject of student answers; guvenlik: guardrails (prompt injection); hukuk: legal routing; moderasyon: offensive language; spam: spam and phishing; sss: customer support. On hukuk's 181 court items and guvenlik's 28 source items the model is supervised in-distribution (see "How to read it" above), so hukuk 0.754 and the guardrail counts are not zero-shot numbers. The guvenlik, moderasyon and sss rows carry the flag "shaped by reading the test results" (below), and the moderation result may read higher than it should, because the comments written for training share the style of the open moderation items (see Limitations).
| Track | Questions of | Qs | Accuracy | Macro F1 | Brier | Smooth ECE |
|---|---|---|---|---|---|---|
| dogrulama | all | 602 | 0.640 [0.602, 0.681] | 0.437 [0.390, 0.481] | 0.499 [0.460, 0.538] | 0.037 [0.033, 0.066] |
| dogrulama | noul | 302 | 0.778 [0.725, 0.825] | 0.775 [0.722, 0.821] | 0.342 [0.286, 0.402] | 0.064 [0.043, 0.100] |
| dogrulama | score | 300 | 0.500 [0.443, 0.557] | 0.268 [0.200, 0.333] | 0.657 [0.612, 0.704] | 0.037 [0.031, 0.083] |
| egitim | all | 640 | 0.787 [0.756, 0.817] | 0.802 [0.773, 0.826] | 0.278 [0.249, 0.308] | 0.025 [0.025, 0.047] |
| egitim | choice | 320 | 0.931 [0.903, 0.956] | 0.932 [0.904, 0.957] | 0.095 [0.065, 0.127] | 0.028 [0.027, 0.048] |
| egitim | score | 320 | 0.644 [0.591, 0.697] | 0.542 [0.487, 0.594] | 0.461 [0.412, 0.507] | 0.036 [0.033, 0.077] |
| guvenlik | all | 418 | 0.770 [0.730, 0.811] | 0.764 [0.722, 0.805] | 0.367 [0.305, 0.433] | 0.158 [0.119, 0.200] |
| hukuk | all | 308 | 0.782 [0.737, 0.825] | 0.754 [0.690, 0.797] | 0.291 [0.238, 0.350] | 0.056 [0.036, 0.097] |
| moderasyon | all | 353 | 0.952 [0.929, 0.975] | 0.951 [0.927, 0.973] | 0.099 [0.077, 0.123] | 0.084 [0.063, 0.105] |
| spam | all | 614 | 0.806 [0.774, 0.839] | 0.746 [0.708, 0.779] | 0.271 [0.237, 0.307] | 0.035 [0.029, 0.060] |
| spam | choice | 307 | 0.713 [0.661, 0.765] | 0.699 [0.655, 0.739] | 0.353 [0.304, 0.400] | 0.045 [0.034, 0.083] |
| spam | noul | 307 | 0.899 [0.863, 0.932] | 0.863 [0.815, 0.907] | 0.189 [0.154, 0.224] | 0.089 [0.068, 0.121] |
| sss | all | 1,340 | 0.569 [0.534, 0.605] | 0.479 [0.446, 0.511] | 0.526 [0.497, 0.554] | 0.069 [0.048, 0.095] |
| sss | choice | 670 | 0.490 [0.443, 0.534] | 0.401 [0.363, 0.437] | 0.605 [0.570, 0.640] | 0.106 [0.069, 0.148] |
| sss | noul | 335 | 0.818 [0.776, 0.860] | 0.816 [0.774, 0.857] | 0.299 [0.265, 0.332] | 0.104 [0.066, 0.143] |
| sss | score | 335 | 0.481 [0.427, 0.534] | 0.411 [0.358, 0.466] | 0.595 [0.555, 0.637] | 0.122 [0.082, 0.173] |
Safety tracks, counted on the open items:
| Track | Positives caught | Negatives passed | Accuracy |
|---|---|---|---|
| guvenlik (positive: prompt-injection attack) | 194 of 217 | 128 of 201 | 0.770 [0.730, 0.811] |
| moderasyon (positive: offensive message) | 150 of 157 | 186 of 196 | 0.952 [0.929, 0.975] |
Every number of this model comes after readings of the full test set, and its guardrail (guvenlik), moderation (moderasyon) and customer support (sss) numbers carry the flag "shaped by reading the test results". Run 2, the run before it, caught 106 of 217 attacks. That was read on HakemBench's full test set, and Run 3 was then fixed in writing beforehand, with new guardrail training data and a selection rule on a separate guardrail dev set; Run 2's own new data had been aimed at Run 1's errors on the full test set in all three flagged tracks. With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. The benign side is weak: 128 of 201 benign messages that look like attacks are passed.
The model scores alike on the benchmark's two parts, A and B: macro F1 0.711 on part A and 0.697 on part B, a gap of 1.42 points (0.0142) [-1.47, 4.17], not flagged.
Probes
| Measure | ufakzeka-karar |
|---|---|
| Order sensitivity under option permutations | 0.000 [0.000, 0.000] |
| Paraphrase agreement | 0.883 [0.833, 0.932] |
| Paraphrase mean total variation distance | 0.105 [0.085, 0.126] |
| English Brier gap, Turkish minus English (raw) | -0.303 [-0.377, -0.225] |
| Slot substitution signal, accuracy | 0.016 [0.000, 0.033] |
| Robustness (mean of (1 minus order sensitivity) and paraphrase agreement) | 0.941 [0.917, 0.966] |
Order sensitivity is zero by construction: options are scored blind to each other at shared positions, bit for bit for up to ten options (one pass); above ten, options are read in passes of ten and reordering changes the answer only up to floating-point noise. The probes were run on eleven of the sixteen board rows (not on the four hosted chat models or the surface-cue baseline). The English gap is negative because the model is better in Turkish; it is not meant to read English. Its robustness value is 0.941 against Jev 1.13's 0.943, but half of that value is 1.0 by construction (option order cannot change its answer); on the measured half, paraphrase agreement, it is 0.883 [0.833, 0.932] against Jev's 0.926 [0.889, 0.963], intervals overlapping. The other models' probe numbers are on the HakemBench card.
External Turkish test sets
Asked as typed questions with no examples in the request, each run with its own calibrator. Run 1 is the first of the three runs scored on HakemBench (per-type temperature), Run 3 the release (global temperature).
| Set | n | Chance | Run 1 accuracy | Run 3 accuracy | Run 1 Brier | Run 3 Brier | Run 1 smooth ECE | Run 3 smooth ECE |
|---|---|---|---|---|---|---|---|---|
| MASSIVE 1.1, tr-TR intents | 2,937 | 0.017 | 0.756 [0.741, 0.772] | 0.745 [0.728, 0.761] | 0.355 [0.335, 0.374] | 0.359 [0.341, 0.378] | 0.029 [0.023, 0.042] | 0.025 [0.021, 0.038] |
| MiDe22 Turkish tweets | 1,012 | 0.333 | 0.766 [0.740, 0.791] | 0.741 [0.714, 0.767] | 0.334 [0.307, 0.365] | 0.350 [0.321, 0.379] | 0.035 [0.027, 0.055] | 0.030 [0.027, 0.054] |
| MMLU-Pro-TR | 11,838 | 0.111 | 0.107 [0.101, 0.112] | 0.098 [0.092, 0.103] | 0.936 [0.933, 0.939] | 0.941 [0.938, 0.944] | 0.126 [0.120, 0.132] | 0.135 [0.130, 0.141] |
| OffensEval-TR 2020, subtask A | 3,528 | 0.500 | 0.866 [0.854, 0.877] | 0.844 [0.832, 0.857] | 0.204 [0.189, 0.218] | 0.225 [0.211, 0.239] | 0.021 [0.017, 0.032] | 0.020 [0.016, 0.031] |
| XCOPA, Turkish | 500 | 0.500 | 0.562 [0.518, 0.604] | 0.584 [0.542, 0.626] | 0.602 [0.550, 0.658] | 0.567 [0.518, 0.618] | 0.196 [0.156, 0.243] | 0.171 [0.133, 0.214] |
| X-FACT, Turkish test claims (Cetvel's four verdicts) | 169 | 0.250 | 0.320 [0.254, 0.391] | 0.343 [0.278, 0.414] | 0.826 [0.769, 0.881] | 0.749 [0.710, 0.783] | 0.181 [0.122, 0.253] | 0.072 [0.047, 0.141] |
Macro F1 where the options are fixed classes: MASSIVE 0.702 to 0.691, MiDe22 0.745 to 0.727, OffensEval 0.772 to 0.755, X-FACT 0.196 to 0.207 (Run 1 to Run 3).
From Run 1 to Run 3, accuracy dropped on MASSIVE, MiDe22, MMLU-Pro-TR and OffensEval and rose on XCOPA and X-FACT; on every set the two intervals overlap. MMLU-Pro-TR is a knowledge exam and the released model is slightly below chance there: 0.098 [0.092, 0.103] against a chance level of 0.111. The model has little world knowledge. MASSIVE and OffensEval are the test splits of sets whose train splits it learned from, and MiDe22, which ships as one file, is the test fifth of a 70/10/20 split the project drew before training, so they are supervised numbers, not zero-shot transfer. FACTurk, a training source, includes Doğruluk Payı, the outlet all of X-FACT's Turkish claims come from; the training rows were checked against X-FACT and none matched, but the domain is shared.
Calibration before and after
Smooth ECE and Brier, raw and after each run's own temperature, on the calibration held-out set (the four support questions HakemBench asks, held out from Run 1's training, which saw the same three families only in other question types; Runs 2 and 3 were trained on them) and on the development-set questions whose gold a human decided:
| Run | Form | Temperature | Held-out ECE raw | Held-out ECE cal. | Held-out Brier raw | Held-out Brier cal. | Human-gold dev ECE raw | Human-gold dev ECE cal. |
|---|---|---|---|---|---|---|---|---|
| Run 1 | per type | 1.2820 global; choice 1.2341, noul 1.4003, score 1.0507 | 0.027 | 0.045 | 0.336 | 0.337 | 0.051 | 0.034 |
| Run 2 | global | 1.2932 | 0.036 | 0.077 | 0.263 | 0.269 | 0.062 | 0.033 |
| Run 3 (released) | global | 1.2614 | 0.036 | 0.064 | 0.274 | 0.278 | 0.050 | 0.038 |
Temperature scaling makes calibration worse on a set of held-out support questions: for Run 1, which never trained on them, Run 1's own temperature raises smooth ECE from 0.027 to 0.045. The released model later trained on those four questions (Runs 2 and 3 add them), so its own 0.036 to 0.064 on that set is not an unseen-question test, and the set shares texts with the validation rows the temperature was fitted on. Brier is slightly worse too (0.274 to 0.278 for the released model). On the development-set questions with human-decided gold the temperature improves smooth ECE (0.050 to 0.038). The model is mildly overconfident on validation data, and the temperature fitted there overcorrects on questions that already sit at a different confidence level. The criterion on the selection split was fixed before the held-out result was seen, so T = 1.2614 ships as chosen; T = 1 would have given held-out smooth ECE 0.036. The calibration axis of the board (0.482) and the smooth ECE per track above are what you get in practice on HakemBench. As a separate measure, the board fits each model its own temperature on part A of the benchmark, on top of the shipped one; for this model it is 1.1733, above 1, so on HakemBench the model is still somewhat overconfident. It is not shipped.
Speed and memory on CPU
On one Apple M1 Mac (8 cores) with one torch thread, on 13 sample questions from the demo, one question per call: a median of 102.5 ms per question (69.8 to 194.5 ms) and a peak resident memory of 986,038,272 bytes, fp32 (release/measure_cpu.py, results/step9/model_facts.json). This is a one-machine measurement, not a benchmark; your hardware and text lengths will differ. On the board run, also on an Apple-silicon Mac CPU, the median was 103.6 ms [102.4, 106.8] per question. open-jev and the shorter-input Laya rows ran on the same Mac CPU; Kev, Qwen3.5-4B and decider-2b ran on an NVIDIA L4 GPU, and the Laya rows at max_len 2048 on a GPU.
An int8 file is not shipped: during calibration, the int8 files built from one earlier checkpoint did not stay within 0.5 points of fp32 on macro F1; calibration error stayed within 0.11 points. So fp32 ships, and no int8 numbers exist for this model.
Training
Backbone. ufakai/ufakzeka-1-base at revision f9e11eea28cbb2ba953a5628f972d416fe0c3cfe (24 layers, d_model 768, vocabulary 40,960), read causally. A masked-language conversion of the backbone was tried and did not beat the causal backbone in the project's comparison of backbones, so it was not used. Parameters: 182,494,466 in all, 151,037,186 without the embedding matrix of 31,457,280; 182,492,928 in the backbone, 769 in the decision layer and 769 in an abstain layer; karar.py takes its abstain value from the calibrator's map.
Data. Turkish only, from openly licensed sets and texts written for training. Converted sets with their own labels: MASSIVE 1.1 (intents), OffensEval 2020 Turkish, MiDe22, Constitutional Court individual applications, FACTurk fact-checks, WebFAQ relevance, and four Turkish prompt-injection sets. Generated rows: decision questions written as templates by various LLMs and applied to question-and-answer texts from clips/mqa, and texts written for the spam, phishing, education, legal and fact-check tracks (the fact-check rows also ask about FACTurk claims and clips/mqa questions), labelled by two LLM judges whose probability vectors are averaged, neither of them the text's writer. The additions of Run 2 kept in the release run: the four support questions on support texts, labelled by two AI models of one family, and moderation comments (offensive and clean) written for training; Run 2's benign guardrail messages were left out of the released variant (see Selection). Every source, its licence and its attribution are in data/MANIFEST.yaml, and the data card is docs/DATA_CARD.md. The mix takes at most 3,000 rows per task; the release run mixed 55,679 rows:
| Training data (files under data/built/) | Rows in the release mix |
|---|---|
| Run 2 additions kept in the release (support questions, moderation comments) | 12,151 |
| sss/train.jsonl | 12,957 |
| synth/train.jsonl | 10,250 |
| typed/aym_rights/train.jsonl | 3,000 |
| typed/facturk_verdict/train.jsonl | 3,000 |
| typed/massive_tr/train.jsonl | 3,000 |
| typed/mide22/train.jsonl | 3,000 |
| typed/offenseval_tr/train.jsonl | 3,000 |
| typed/prompt_injection/train.jsonl | 1,791 |
| typed/prompt_injection_conv/train.jsonl | 530 |
| typed/webfaq_relevance/train.jsonl | 3,000 |
The rows per task are listed in the code repository. Before every build, training rows were checked against HakemBench (both parts, the probes and the development set) and against 70 evaluation sets (328,398 items, among them the test splits of Cetvel, TrGLUE and TabiBench) by n-gram and near-duplicate rules, and every matching row was removed.
Recipe. Each question is packed as one sequence: the text and the question as a prefix, then every option as its own segment that attends to the prefix and to itself only, at shared positions. Each option's tokens are mean-pooled and scored by one linear layer, and one softmax over the options gives the distribution. Training minimises cross-entropy against the label distributions (soft where the labellers disagreed), adds the ranked probability score on score questions so that mass two levels off costs more than one level off, and weights rows so each task's answer mix returns to its natural one. Backbone learning rate 1e-4, head 1e-3, weight decay 0.01, 2 epochs, 32 questions per batch, 6 percent warmup, prefix 448 and option 48 tokens: 3,480 steps. No reinforcement learning; a REINFORCE variant was tested as an ablation and lost macro F1.
Compute. The released run: one NVIDIA H200, 427.8 s of training, 692.9 s wall, 13.5 GB peak memory.
Selection. The protocol of Run 3 trained six runs, two variants with three random seeds each: one variant keeps half of Run 2's benign guardrail rows and the other leaves them out. Under the rule, fixed in writing before any run, a run qualifies if it catches at least 0.80 of the attacks on a separate guardrail dev set (the validation and test splits of the conversation prompt-injection set, never trained on); among qualifying runs the highest macro F1 on the support development set (255 support questions with human answers, kept apart from the test set) wins. A run of the variant without those rows was chosen; it caught 42 of 50 dev attacks and passed 161 of 170 benign dev messages. It was then scored once on HakemBench v1.0, with no tuning after it. The selection rule and its result ship with the code.
The test-set readings. The released model, Run 3, is the last of three runs scored on HakemBench (Runs 1 to 3), and its numbers are not blind. Run 2's new training data was aimed at Run 1's errors on the full test set in guardrails, moderation and customer support, and Run 3 was trained after Run 2's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. Every number of the released model comes after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16.
Limitations and bias
- World knowledge is weak. Slightly below chance on MMLU-Pro-TR: 0.098 [0.092, 0.103] against a chance level of 0.111. It decides from the text; it does not know facts the text does not state.
- Support routing is weak. On the 84 support questions of the human check it is below a surface-cue baseline that uses cues such as length, digits or a question mark, not meaning (0.476 against 0.619). All human answers come from one person; there is no inter-annotator agreement. Its support scores on AI-made gold are partly agreement with its own labeller (above).
- Turkish only. English input is read poorly.
- Long texts are cut. Text and question share 448 tokens; the end of a long text is not read.
- Three tracks were shaped by reading the test results. Its guardrail, moderation and customer support numbers carry that flag; with every model scored on the other four tracks only, its composite is 0.678, 6th of 16. On guardrails only 128 of 201 benign look-alike messages are passed.
- Its moderation result may read higher than it should. The 1,499 moderation comments written for its training share the style of the benchmark's written moderation items, which are all of the open moderation items. A character n-gram classifier trained on those comments reaches 0.926 balanced accuracy on the written items but 0.401 on web-sourced moderation texts.
- The not-sure signal misses guardrail errors. At confidence 0.99 or more the model still answers 148 guardrail questions, and 18 of those answers, 12 percent, are wrong; the signal cannot be relied on to catch prompt-injection errors. On support almost no answer reaches 0.90.
- Not zero-shot on 209 benchmark questions. The court items and the source guardrail items are test splits of datasets it trained on; the other models answer them zero-shot.
- Labels come mostly from LLMs. Most training labels are a panel of LLMs' votes, rule labels from the source sets, or two labellers of one model family; none of the prompt-injection or relevance labels was checked by a person. The model inherits their readings of subjective scales.
- Calibration does not transfer everywhere. Temperature scaling makes calibration worse on held-out support questions (0.036 to 0.064 for the released model, which has since trained on those questions; 0.027 to 0.045 for Run 1, which never did), and on HakemBench the model is somewhat overconfident. Probabilities on a new kind of question may be over- or underconfident; check them on your data.
- Bias. Training texts include tweets (OffensEval, MiDe22), court text, web question-and-answer pages and LLM-written texts. Their topics, registers and the LLMs' writing style shape what the model finds offensive, urgent or checkable. Moderation judgments in particular can differ from a given community's norms.
Intended use
- Turkish text decisions where the options are known in advance: routing, tagging, triage, moderation queues, grading assistance, fact-check triage, prompt-injection screening as one layer among others.
- Selective automation: act on confident answers and send the rest to a person, with a threshold you set on your own data.
- Research on calibrated, order-invariant decision models and on HakemBench.
Out-of-scope use
- Decisions about people with legal, medical, financial or employment consequences without a person reviewing each case.
- Sole protection against prompt injection or abuse.
- Questions that need world knowledge or reasoning beyond the given text.
- Languages other than Turkish.
- Generating or explaining text.
Cost
The project cost $236.90 in GPU time and hosted model API calls ($181.55 on Modal, $55.35 in API calls). The figure does not include the base model ufakzeka-1, whose cost is reported with it, the server or the AI model used during development. The figure covers this model and HakemBench together and was read from the providers' dashboards.
Licence and attribution
The weights and karar.py are under the Apache License 2.0. The backbone, ufakzeka-1-base, is Apache-2.0. The training data comes from sources under Apache-2.0, MIT, CC0, CC BY 2.0 or CC BY 4.0, and from texts written for the project. The sources are licensed for this use as far as their licences reach: clips/mqa and WebFAQ license only the collection, while the page text belongs to each site; OffensEval's licence covers its annotations; the MiDe22 and OffensEval texts are tweets owned by their users; FACTurk's claims were collected from fact-checking sites. Every source with its licence string and attribution line is listed in data/MANIFEST.yaml and described in docs/DATA_CARD.md. The benchmark's items and their licences are in the HakemBench dataset and its ATTRIBUTION.md.
Citation
The paper describes the method and the results; the model entry cites the released weights.
The paper:
@misc{teke2026ufakzekakarar,
title = {ufakzeka-karar: An Open {Turkish} Typed-Decision Model with Order-Invariant Option Scoring},
author = {Teke, Sait Furkan},
year = {2026},
eprint = {2610.06744},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.06744}
}
The model:
@misc{ufakzekakarar2026,
title = {ufakzeka-karar},
author = {{ufak AI}},
year = {2026},
howpublished = {\url{https://huggingface.co/ufakai/ufakzeka-karar}}
}
If you use the benchmark numbers, cite HakemBench as its card asks.
Links
- Model: https://huggingface.co/ufakai/ufakzeka-karar
- HakemBench dataset: https://huggingface.co/datasets/ufakai/HakemBench
- Code: https://github.com/ufakai/ufakzeka-karar
- Benchmark harness: https://github.com/ufakai/hakembench
- Demo: https://karar.ufakzeka.com/en
- Papers: https://arxiv.org/abs/2610.06744 (model) and https://arxiv.org/abs/2610.02293 (benchmark)
- Write-ups: https://ufakai.com/research/ufakzeka-karar and https://ufakai.com/research/hakembench (Turkish under https://ufakai.com/tr/research/)
- Backbone: https://huggingface.co/ufakai/ufakzeka-1-base
- Downloads last month
- 10
Model tree for ufakai/ufakzeka-karar
Base model
ufakai/ufakzeka-1-baseDatasets used to train ufakai/ufakzeka-karar
PaDaS-Lab/webfaq-retrieval
clips/mqa
Papers for ufakai/ufakzeka-karar
ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring
HakemBench: A Turkish Benchmark of Typed Decisions
Evaluation results
- Composite (geometric mean of decision quality, calibration, selective) on HakemBench v1.0, open settest set HakemBench public board0.660
- Decision quality axis (macro F1, mean over tracks; intelligence in the board files) on HakemBench v1.0, open settest set HakemBench public board0.705
- Calibration axis (1 minus normalised Brier) on HakemBench v1.0, open settest set HakemBench public board0.482
- Selective axis (1 minus normalised AUGRC) on HakemBench v1.0, open settest set HakemBench public board0.848
- Accuracy, dogrulama (fact-check triage) on HakemBench v1.0, open settest set HakemBench public board0.640
- Macro F1, dogrulama (fact-check triage) on HakemBench v1.0, open settest set HakemBench public board0.437
- Accuracy, egitim (education) on HakemBench v1.0, open settest set HakemBench public board0.787
- Macro F1, egitim (education) on HakemBench v1.0, open settest set HakemBench public board0.802