Spaces:
Running on Zero
Download README.md from muratcanlaloglu/TurkishDecisionBenchmark: direct link, hf CLI and curl.
- Browser
- Download file 9.02 kB
-
https://huggingface.co/spaces/muratcanlaloglu/TurkishDecisionBenchmark/resolve/main/README.md
- Command line
-
hf download hf://spaces/muratcanlaloglu/TurkishDecisionBenchmark/README.md
-
curl -L -o README.md https://huggingface.co/spaces/muratcanlaloglu/TurkishDecisionBenchmark/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.29.1
title: TurkishDecisionBenchmark
emoji: 🏆
colorFrom: red
colorTo: indigo
sdk: gradio
sdk_version: 6.29.0
app_file: app.py
pinned: false
short_description: Turkish typed-decision / classification model leaderboard
TurkishDecisionBenchmark
TurkishDecisionBenchmark (TDB) is a benchmark for typed decision /
classification models on difficult Turkish language understanding. Each case is
a Turkish customer message, a typed choice question with described options, and
one gold option.
v0.2 (current)
- 396 cases: 354 scored + 42 deliberately ambiguous, unscored controls
- 8 domains: e-commerce, banking, subscription, telecom, shipping, health (admin), public services, travel — 22 domain tasks + 2 shared tasks (urgency, which of two items)
- 10 phenomena, 32–38 scored cases each (see below)
- 69 contrast groups (minimal edits of one message that flip the answer) covering 195 cases
- Decisive span balanced across the message: start 110, middle 98, end 113, n/a 33
- Difficulty (scored): easy 28, medium 144, hard 182
- Label balance: in every task the most frequent label is at most 2× the least frequent
SHA-256 (both are checked before a result is admitted to the leaderboard):
dataset/v0.2/public.jsonl d2be12688bd957bdce7573440c9d5f3caa04bdb9e6985fab5583e433967e848c
dataset/v0.2/tasks.json 88627dddf653f51ef707d4e704e545422ad4780cfab195fbfb0c13c306662846
Shortcut resistance
Every case set was rewritten until no surface rule can solve it. The acceptance
gate (scripts/baseline_report.py) requires the overlap, last-clause and
first-clause baselines to stay at or below 45% overall and 55% per category.
| Baseline | Accuracy | Group consistency |
|---|---|---|
| majority | 28% | 0% |
| overlap (message ↔ option words) | 30% | 1% |
| last_clause | 30% | 0% |
| first_clause | 29% | 3% |
Categories
| Category | Focus |
|---|---|
| negation | Polarity, double negation, ne … ne |
| correction | Latest intent after self-correction |
| implicit_intent | Indirect requests, stated consequences |
| temporal_reasoning | Resolved past problem vs current one |
| conditional | -sA / -sAydI, conditions met or not by other facts |
| reported_speech | Someone else's words vs the writer's intent |
| numeric_date | Date / amount / count arithmetic |
| distractor | Words pointing to wrong options |
| coreference_scope | Which item, whose request, quantifier scope |
| noisy_turkish | ASCII, typos, colloquial spelling |
| ambiguous | Multi-valid controls; unscored |
Review status
All v0.2 cases are draft: written by the authors, then labelled blind by an
independent LLM annotator (dataset/v0.2/annotations/llm/grok.csv), which agreed
on 352/354 scored cases (κ = 0.99); both disagreements were rewritten. Model labels
never mark a case as reviewed. Human review is pending
(dataset/v0.2/annotations/<name>.csv, see Dataset workflow).
Known limitation: a strong general-purpose LLM solves almost every case. Real runs separate the decision models from each other and from the surface baselines (best model 92.9%, best baseline 29.9%).
Installation
API models + leaderboard:
pip install -r requirements.txt
Local models (optional):
pip install -r requirements-laya.txt # Laya
bash scripts/setup_julia.sh # Julia-1 → models/Julia-1 (Python 3.11+, ~550 MiB)
API keys go in .env (git-ignored; runner.py loads it automatically):
cp .env.example .env
Run benchmark
Presets
| Preset | Provider | Model | Interface |
|---|---|---|---|
liquid-d1-free |
liquid | d1:free |
decisions |
jev-1.13 |
openrouter | typesafe/jev-1.13 |
decisions |
kev-4b |
openrouter | jaredpalmer/kev-4b |
decisions |
solar-decide |
openrouter | upstage/solar-decide |
decisions |
laya-multilingual |
laya | convaiinnovations/laya-multilingual |
local |
julia-1 |
julia | SupersonicLabs/Julia-1 |
local |
baseline-majority / -overlap / -last-clause / -first-clause |
baseline | — | baseline (offline reference) |
python runner.py --preset kev-4b
python runner.py --provider openrouter --model <any-decisions-model>
python scripts/run_all.py # all presets with an available key
python scripts/run_all.py kev-4b solar-decide -- --limit 5 # smoke test, goes to results/v0.2/partial/
The Liquid free endpoint is slow; raise --timeout above 60 if you see read timeouts.
Failed calls count as incorrect. --retry-errors results/v0.2/<file>.json reruns only
the failed cases of a finished run and merges them back in.
Only models that answer choice questions are benchmarked. Respan Span-01 / Span-01 Lite
accept only yes/no (noul) behavior questions and are therefore not included.
Laya options
python runner.py \
--provider laya \
--model convaiinnovations/laya-multilingual \
--laya-max-len 8192
Runner options
| Flag | Meaning |
|---|---|
--version |
Benchmark version (default 0.2) |
--category, --difficulty, --limit N |
Partial run; written to .../partial/ (git-ignored, never on the leaderboard) |
--scored-only |
Skip the ambiguous controls (they run by default so ambiguous metrics are reported) |
--timeout, --max-retries |
HTTP providers; 429 / 5xx / network errors are retried with backoff |
--retry-errors FILE |
Rerun only the failed cases of a result file and merge them back in |
--delay S |
Wait S seconds between cases (rate-limited endpoints) |
--julia-path, --device |
Julia-1 checkpoint path (default models/Julia-1) and device |
--output |
Custom output path |
Results go to results/v0.2/<provider>__<model>.json.
Every provider implements the same adapter interface
(benchmark/adapters/base.py): predict(state, question) -> {choice, probabilities, raw}.
A choice outside the task's classes is recorded as an invalid_prediction error.
Metrics
- Accuracy over all scored cases. API errors and invalid labels stay in the denominator and count as incorrect.
- Group consistency: share of contrast groups where every variant is correct.
- High-confidence error rate: wrong answers with
p >= 0.90, divided by scored cases. - Mean gold p: mean probability assigned to the gold label (answered cases).
- Ambiguous valid-choice rate and ambiguous mean max-probability (overconfidence on cases without a single correct answer).
- Breakdowns by category, domain and difficulty.
Leaderboard Space
The repository is ready to be uploaded directly to a Hugging Face Gradio Space.
app.py reads committed JSON result files and displays:
- 🏆 Overall leaderboard (accuracy → group consistency → high-conf error → mean gold p)
- 🧩 Category, 🗂️ domain and 🔥 difficulty breakdowns
- 🔎 Per-model error explorer
- ℹ️ Methodology and excluded result files
The Space does not expose or consume private API keys for public visitors.
Benchmark runs are executed with runner.py, then the produced result JSON is
committed under results/v0.2/.
Leaderboard rule
A result file is shown only if it matches the benchmark version, the public split, the dataset and tasks SHA-256, is unfiltered, and covers all scored cases. Everything else is listed as "excluded" in the Methodology tab. Results from an older case set are not comparable and must not be committed.
Dataset workflow
Cases live in scripts/v02_cases/<domain>.py; the JSONL is generated — do not edit it.
python scripts/build_v02.py # → dataset/v0.2/public.jsonl
python scripts/validate_dataset.py # schema, labels, groups, balance
python scripts/baseline_report.py # shortcut baselines + acceptance gate
# one module on its own
python scripts/build_v02.py --module travel --out /tmp/travel.jsonl
python scripts/validate_dataset.py --dataset /tmp/travel.jsonl
python scripts/baseline_report.py --dataset /tmp/travel.jsonl
Blind human annotation:
python scripts/annotation.py export --annotator <name> # shuffled sheet, no gold labels
# fill `etiket` (option number or key, `?` if unsure), `dogal_degil`, `not`
python scripts/annotation.py compare --annotator <name> # agreement, kappa, disagreements
python scripts/build_v02.py # merges labels, marks cases reviewed
Changing any case changes the dataset SHA-256; rerun the baselines and models afterwards.
Development roadmap
v0.2 remaining
- Human review of a sample (or all) cases; adjudicate disagreements
- Real model runs to check for a ceiling effect
- Hidden evaluation split (~1/3, git-ignored)
v0.3
- Task gaps found while writing v0.2 (add-on service cancellation, delivered-to-neighbour, grandchild / in-law patients, delay that will cause a missed connection)
- Longer, multi-turn contexts
- Brier score / ECE calibration metrics
- Repeated-run variance