muratcanlaloglu
Cursor
Use a single emoji in the Space metadata so Hugging Face accepts the push.
2056c3a
|
Raw History Blame Contribute Delete
9.02 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: TurkishDecisionBenchmark
emoji: 🏆
colorFrom: red
colorTo: indigo
sdk: gradio
sdk_version: 6.29.0
app_file: app.py
pinned: false
short_description: Turkish typed-decision / classification model leaderboard

TurkishDecisionBenchmark

TurkishDecisionBenchmark (TDB) is a benchmark for typed decision / classification models on difficult Turkish language understanding. Each case is a Turkish customer message, a typed choice question with described options, and one gold option.

v0.2 (current)

  • 396 cases: 354 scored + 42 deliberately ambiguous, unscored controls
  • 8 domains: e-commerce, banking, subscription, telecom, shipping, health (admin), public services, travel — 22 domain tasks + 2 shared tasks (urgency, which of two items)
  • 10 phenomena, 32–38 scored cases each (see below)
  • 69 contrast groups (minimal edits of one message that flip the answer) covering 195 cases
  • Decisive span balanced across the message: start 110, middle 98, end 113, n/a 33
  • Difficulty (scored): easy 28, medium 144, hard 182
  • Label balance: in every task the most frequent label is at most 2× the least frequent

SHA-256 (both are checked before a result is admitted to the leaderboard):

dataset/v0.2/public.jsonl  d2be12688bd957bdce7573440c9d5f3caa04bdb9e6985fab5583e433967e848c
dataset/v0.2/tasks.json    88627dddf653f51ef707d4e704e545422ad4780cfab195fbfb0c13c306662846

Shortcut resistance

Every case set was rewritten until no surface rule can solve it. The acceptance gate (scripts/baseline_report.py) requires the overlap, last-clause and first-clause baselines to stay at or below 45% overall and 55% per category.

Baseline Accuracy Group consistency
majority 28% 0%
overlap (message ↔ option words) 30% 1%
last_clause 30% 0%
first_clause 29% 3%

Categories

Category Focus
negation Polarity, double negation, ne … ne
correction Latest intent after self-correction
implicit_intent Indirect requests, stated consequences
temporal_reasoning Resolved past problem vs current one
conditional -sA / -sAydI, conditions met or not by other facts
reported_speech Someone else's words vs the writer's intent
numeric_date Date / amount / count arithmetic
distractor Words pointing to wrong options
coreference_scope Which item, whose request, quantifier scope
noisy_turkish ASCII, typos, colloquial spelling
ambiguous Multi-valid controls; unscored

Review status

All v0.2 cases are draft: written by the authors, then labelled blind by an independent LLM annotator (dataset/v0.2/annotations/llm/grok.csv), which agreed on 352/354 scored cases (κ = 0.99); both disagreements were rewritten. Model labels never mark a case as reviewed. Human review is pending (dataset/v0.2/annotations/<name>.csv, see Dataset workflow).

Known limitation: a strong general-purpose LLM solves almost every case. Real runs separate the decision models from each other and from the surface baselines (best model 92.9%, best baseline 29.9%).

Installation

API models + leaderboard:

pip install -r requirements.txt

Local models (optional):

pip install -r requirements-laya.txt   # Laya
bash scripts/setup_julia.sh            # Julia-1 → models/Julia-1 (Python 3.11+, ~550 MiB)

API keys go in .env (git-ignored; runner.py loads it automatically):

cp .env.example .env

Run benchmark

Presets

Preset Provider Model Interface
liquid-d1-free liquid d1:free decisions
jev-1.13 openrouter typesafe/jev-1.13 decisions
kev-4b openrouter jaredpalmer/kev-4b decisions
solar-decide openrouter upstage/solar-decide decisions
laya-multilingual laya convaiinnovations/laya-multilingual local
julia-1 julia SupersonicLabs/Julia-1 local
baseline-majority / -overlap / -last-clause / -first-clause baseline — baseline (offline reference)
python runner.py --preset kev-4b
python runner.py --provider openrouter --model <any-decisions-model>
python scripts/run_all.py                       # all presets with an available key
python scripts/run_all.py kev-4b solar-decide -- --limit 5   # smoke test, goes to results/v0.2/partial/

The Liquid free endpoint is slow; raise --timeout above 60 if you see read timeouts. Failed calls count as incorrect. --retry-errors results/v0.2/<file>.json reruns only the failed cases of a finished run and merges them back in.

Only models that answer choice questions are benchmarked. Respan Span-01 / Span-01 Lite accept only yes/no (noul) behavior questions and are therefore not included.

Laya options

python runner.py \
  --provider laya \
  --model convaiinnovations/laya-multilingual \
  --laya-max-len 8192

Runner options

Flag Meaning
--version Benchmark version (default 0.2)
--category, --difficulty, --limit N Partial run; written to .../partial/ (git-ignored, never on the leaderboard)
--scored-only Skip the ambiguous controls (they run by default so ambiguous metrics are reported)
--timeout, --max-retries HTTP providers; 429 / 5xx / network errors are retried with backoff
--retry-errors FILE Rerun only the failed cases of a result file and merge them back in
--delay S Wait S seconds between cases (rate-limited endpoints)
--julia-path, --device Julia-1 checkpoint path (default models/Julia-1) and device
--output Custom output path

Results go to results/v0.2/<provider>__<model>.json.

Every provider implements the same adapter interface (benchmark/adapters/base.py): predict(state, question) -> {choice, probabilities, raw}. A choice outside the task's classes is recorded as an invalid_prediction error.

Metrics

  • Accuracy over all scored cases. API errors and invalid labels stay in the denominator and count as incorrect.
  • Group consistency: share of contrast groups where every variant is correct.
  • High-confidence error rate: wrong answers with p >= 0.90, divided by scored cases.
  • Mean gold p: mean probability assigned to the gold label (answered cases).
  • Ambiguous valid-choice rate and ambiguous mean max-probability (overconfidence on cases without a single correct answer).
  • Breakdowns by category, domain and difficulty.

Leaderboard Space

The repository is ready to be uploaded directly to a Hugging Face Gradio Space.

app.py reads committed JSON result files and displays:

  • 🏆 Overall leaderboard (accuracy → group consistency → high-conf error → mean gold p)
  • 🧩 Category, 🗂️ domain and 🔥 difficulty breakdowns
  • 🔎 Per-model error explorer
  • ℹ️ Methodology and excluded result files

The Space does not expose or consume private API keys for public visitors. Benchmark runs are executed with runner.py, then the produced result JSON is committed under results/v0.2/.

Leaderboard rule

A result file is shown only if it matches the benchmark version, the public split, the dataset and tasks SHA-256, is unfiltered, and covers all scored cases. Everything else is listed as "excluded" in the Methodology tab. Results from an older case set are not comparable and must not be committed.

Dataset workflow

Cases live in scripts/v02_cases/<domain>.py; the JSONL is generated — do not edit it.

python scripts/build_v02.py                 # → dataset/v0.2/public.jsonl
python scripts/validate_dataset.py          # schema, labels, groups, balance
python scripts/baseline_report.py           # shortcut baselines + acceptance gate

# one module on its own
python scripts/build_v02.py --module travel --out /tmp/travel.jsonl
python scripts/validate_dataset.py --dataset /tmp/travel.jsonl
python scripts/baseline_report.py --dataset /tmp/travel.jsonl

Blind human annotation:

python scripts/annotation.py export --annotator <name>    # shuffled sheet, no gold labels
# fill `etiket` (option number or key, `?` if unsure), `dogal_degil`, `not`
python scripts/annotation.py compare --annotator <name>   # agreement, kappa, disagreements
python scripts/build_v02.py                               # merges labels, marks cases reviewed

Changing any case changes the dataset SHA-256; rerun the baselines and models afterwards.

Development roadmap

v0.2 remaining

  • Human review of a sample (or all) cases; adjudicate disagreements
  • Real model runs to check for a ceiling effect
  • Hidden evaluation split (~1/3, git-ignored)

v0.3

  • Task gaps found while writing v0.2 (add-on service cancellation, delivered-to-neighbour, grandchild / in-law patients, delay that will cause a missed connection)
  • Longer, multi-turn contexts
  • Brier score / ECE calibration metrics
  • Repeated-run variance