--- title: TurkishDecisionBenchmark emoji: 🏆 colorFrom: red colorTo: indigo sdk: gradio sdk_version: 6.29.0 app_file: app.py pinned: false short_description: Turkish typed-decision / classification model leaderboard --- # TurkishDecisionBenchmark **TurkishDecisionBenchmark (TDB)** is a benchmark for typed decision / classification models on difficult Turkish language understanding. Each case is a Turkish customer message, a typed `choice` question with described options, and one gold option. ## v0.2 (current) - 396 cases: 354 scored + 42 deliberately ambiguous, unscored controls - 8 domains: e-commerce, banking, subscription, telecom, shipping, health (admin), public services, travel — 22 domain tasks + 2 shared tasks (urgency, which of two items) - 10 phenomena, 32–38 scored cases each (see below) - 69 contrast groups (minimal edits of one message that flip the answer) covering 195 cases - Decisive span balanced across the message: start 110, middle 98, end 113, n/a 33 - Difficulty (scored): easy 28, medium 144, hard 182 - Label balance: in every task the most frequent label is at most 2× the least frequent SHA-256 (both are checked before a result is admitted to the leaderboard): ```text dataset/v0.2/public.jsonl d2be12688bd957bdce7573440c9d5f3caa04bdb9e6985fab5583e433967e848c dataset/v0.2/tasks.json 88627dddf653f51ef707d4e704e545422ad4780cfab195fbfb0c13c306662846 ``` ### Shortcut resistance Every case set was rewritten until no surface rule can solve it. The acceptance gate (`scripts/baseline_report.py`) requires the overlap, last-clause and first-clause baselines to stay at or below 45% overall and 55% per category. | Baseline | Accuracy | Group consistency | |---|---|---| | majority | 28% | 0% | | overlap (message ↔ option words) | 30% | 1% | | last_clause | 30% | 0% | | first_clause | 29% | 3% | ### Categories | Category | Focus | |---|---| | negation | Polarity, double negation, *ne â€Ļ ne* | | correction | Latest intent after self-correction | | implicit_intent | Indirect requests, stated consequences | | temporal_reasoning | Resolved past problem vs current one | | conditional | *-sA / -sAydI*, conditions met or not by other facts | | reported_speech | Someone else's words vs the writer's intent | | numeric_date | Date / amount / count arithmetic | | distractor | Words pointing to wrong options | | coreference_scope | Which item, whose request, quantifier scope | | noisy_turkish | ASCII, typos, colloquial spelling | | ambiguous | Multi-valid controls; unscored | ### Review status All v0.2 cases are **draft**: written by the authors, then labelled blind by an independent LLM annotator (`dataset/v0.2/annotations/llm/grok.csv`), which agreed on 352/354 scored cases (Îē = 0.99); both disagreements were rewritten. Model labels never mark a case as reviewed. Human review is pending (`dataset/v0.2/annotations/.csv`, see [Dataset workflow](#dataset-workflow)). Known limitation: a strong general-purpose LLM solves almost every case. Real runs separate the decision models from each other and from the surface baselines (best model 92.9%, best baseline 29.9%). ## Installation API models + leaderboard: ```bash pip install -r requirements.txt ``` Local models (optional): ```bash pip install -r requirements-laya.txt # Laya bash scripts/setup_julia.sh # Julia-1 → models/Julia-1 (Python 3.11+, ~550 MiB) ``` API keys go in `.env` (git-ignored; `runner.py` loads it automatically): ```bash cp .env.example .env ``` ## Run benchmark ### Presets | Preset | Provider | Model | Interface | |---|---|---|---| | `liquid-d1-free` | liquid | `d1:free` | decisions | | `jev-1.13` | openrouter | `typesafe/jev-1.13` | decisions | | `kev-4b` | openrouter | `jaredpalmer/kev-4b` | decisions | | `solar-decide` | openrouter | `upstage/solar-decide` | decisions | | `laya-multilingual` | laya | `convaiinnovations/laya-multilingual` | local | | `julia-1` | julia | `SupersonicLabs/Julia-1` | local | | `baseline-majority` / `-overlap` / `-last-clause` / `-first-clause` | baseline | — | baseline (offline reference) | ```bash python runner.py --preset kev-4b python runner.py --provider openrouter --model python scripts/run_all.py # all presets with an available key python scripts/run_all.py kev-4b solar-decide -- --limit 5 # smoke test, goes to results/v0.2/partial/ ``` The Liquid free endpoint is slow; raise `--timeout` above 60 if you see read timeouts. Failed calls count as incorrect. `--retry-errors results/v0.2/.json` reruns only the failed cases of a finished run and merges them back in. Only models that answer `choice` questions are benchmarked. Respan Span-01 / Span-01 Lite accept only yes/no (`noul`) behavior questions and are therefore not included. ### Laya options ```bash python runner.py \ --provider laya \ --model convaiinnovations/laya-multilingual \ --laya-max-len 8192 ``` ### Runner options | Flag | Meaning | |---|---| | `--version` | Benchmark version (default `0.2`) | | `--category`, `--difficulty`, `--limit N` | Partial run; written to `.../partial/` (git-ignored, never on the leaderboard) | | `--scored-only` | Skip the ambiguous controls (they run by default so ambiguous metrics are reported) | | `--timeout`, `--max-retries` | HTTP providers; 429 / 5xx / network errors are retried with backoff | | `--retry-errors FILE` | Rerun only the failed cases of a result file and merge them back in | | `--delay S` | Wait S seconds between cases (rate-limited endpoints) | | `--julia-path`, `--device` | Julia-1 checkpoint path (default `models/Julia-1`) and device | | `--output` | Custom output path | Results go to `results/v0.2/__.json`. Every provider implements the same adapter interface (`benchmark/adapters/base.py`): `predict(state, question) -> {choice, probabilities, raw}`. A choice outside the task's classes is recorded as an `invalid_prediction` error. ## Metrics - **Accuracy** over all scored cases. API errors and invalid labels stay in the denominator and count as incorrect. - **Group consistency**: share of contrast groups where every variant is correct. - **High-confidence error rate**: wrong answers with `p >= 0.90`, divided by scored cases. - **Mean gold p**: mean probability assigned to the gold label (answered cases). - **Ambiguous valid-choice rate** and **ambiguous mean max-probability** (overconfidence on cases without a single correct answer). - Breakdowns by category, domain and difficulty. ## Leaderboard Space The repository is ready to be uploaded directly to a Hugging Face **Gradio Space**. `app.py` reads committed JSON result files and displays: - 🏆 Overall leaderboard (accuracy → group consistency → high-conf error → mean gold p) - 🧩 Category, đŸ—‚ī¸ domain and đŸ”Ĩ difficulty breakdowns - 🔎 Per-model error explorer - â„šī¸ Methodology and excluded result files The Space does **not** expose or consume private API keys for public visitors. Benchmark runs are executed with `runner.py`, then the produced result JSON is committed under `results/v0.2/`. ### Leaderboard rule A result file is shown only if it matches the benchmark version, the public split, the dataset **and** tasks SHA-256, is unfiltered, and covers all scored cases. Everything else is listed as "excluded" in the Methodology tab. Results from an older case set are not comparable and must not be committed. ## Dataset workflow Cases live in `scripts/v02_cases/.py`; the JSONL is generated — do not edit it. ```bash python scripts/build_v02.py # → dataset/v0.2/public.jsonl python scripts/validate_dataset.py # schema, labels, groups, balance python scripts/baseline_report.py # shortcut baselines + acceptance gate # one module on its own python scripts/build_v02.py --module travel --out /tmp/travel.jsonl python scripts/validate_dataset.py --dataset /tmp/travel.jsonl python scripts/baseline_report.py --dataset /tmp/travel.jsonl ``` Blind human annotation: ```bash python scripts/annotation.py export --annotator # shuffled sheet, no gold labels # fill `etiket` (option number or key, `?` if unsure), `dogal_degil`, `not` python scripts/annotation.py compare --annotator # agreement, kappa, disagreements python scripts/build_v02.py # merges labels, marks cases reviewed ``` Changing any case changes the dataset SHA-256; rerun the baselines and models afterwards. ## Development roadmap ### v0.2 remaining - Human review of a sample (or all) cases; adjudicate disagreements - Real model runs to check for a ceiling effect - Hidden evaluation split (~1/3, git-ignored) ### v0.3 - Task gaps found while writing v0.2 (add-on service cancellation, delivered-to-neighbour, grandchild / in-law patients, delay that will cause a missed connection) - Longer, multi-turn contexts - Brier score / ECE calibration metrics - Repeated-run variance