mayafree's picture
SEO: rich tags, 21 model cross-links, keyword body (JEV/open-jev/Nimble/CLM/decider)
c77b7a8 verified
|
Raw History Blame Contribute Delete
6.44 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
title: Typed Decision Leaderboard
emoji: 🎯
colorFrom: yellow
colorTo: gray
sdk: gradio
sdk_version: 6.27.0
python_version: '3.10'
app_file: app.py
pinned: false
license: apache-2.0
short_description: JEV & open-Jev answer verifiers, one identical test set
tags:
  - leaderboard
  - answer-verification
  - hallucination-detection
  - llm-evaluation
  - confidence-estimation
  - calibration
  - jev
  - open-jev
  - typed-decisions
  - decision-model
  - verifier
  - system-one
  - llm-judge
  - factuality
  - uncertainty-estimation
  - zero-token
  - AUC
  - ECE
  - benchmark
  - ztc
models:
  - FINAL-Bench/Darwin-397B-ZTC
  - FINAL-Bench/ZTC-Judge-27B
  - FINAL-Bench/ZTC-Judge-9B
  - FINAL-Bench/ZTC-Judge-4B
  - bespokelabs/Bespoke-Nimble-9B
  - Mapika/decider-2b
  - Contrastive-LM/CLM-v0.1-8B
  - pngwn/system-one-qwen3.5-4b-scorer
  - openjev/openjev
  - AlexWortega/openjev
  - jaredpalmer/kev-4b
  - wfzyx/von
  - ZefanCai/Open-Jev-9B
  - apus-ailab/APUS-OpenJev-v1
  - com-kotobalabs/open-jev-deberta-v3-large
  - convaiinnovations/laya
  - convaiinnovations/laya-typed-decisions
  - convaiinnovations/laya-multilingual
  - PatronusAI/Llama-3-Patronus-Lynx-8B-Instruct
  - vectara/hallucination_evaluation_model
  - Qwen/Qwen3-Next-80B-A3B-Instruct

Typed Decision Leaderboard — JEV & open-Jev answer verifiers, one identical test set

An independent, side-by-side benchmark of answer verifiers / typed-decision models / System-1 scorers — the models that read an LLM's answer and decide, with zero generated tokens, whether it can be trusted. Every system is scored on the same 2,018 items with the same labels; the scores, labels and grading code are published, and each competitor is run on its own maker's code.

If you searched for a JEV alternative, an open-Jev leaderboard, Bespoke Nimble vs JEV, CLM-8B benchmark, decider-2b AUC, a hallucination-detection / answer-verification leaderboard, or calibrated confidence (ECE) for LLM answers — this is that table.

Current ranking — weighted AUC (higher is better)

# System AUC Division Notes
1 JEV (TypeSafe AI) 0.7350 API commercial API; three-way tie for first
1 ZTC-Judge-27B (VIDRAFT / FINAL-Bench) 0.7289 local open weights; tie for first
1 Darwin-397B-ZTC (VIDRAFT) 0.7272 local MoE; tie for first
4 ZTC-Judge-9B 0.6506 local open weights
5 ZTC-Judge-4B 0.6360 local runs on a laptop
6 Bespoke-Nimble-9B (Bespoke Labs) 0.6222 local ties the length/format baseline
— length & format baseline 0.6223 — content-blind reference
7 open-jev 4B (pngwn) 0.6101 local below baseline
8 decider-2b (Mapika) 0.6028 local below baseline
9 CLM-8B (Stanford · NVIDIA) 0.5704 local below baseline
10 Patronus Lynx 8B 0.5179 local evidence-grounded design (different axis)
11 Laya-Typed-Decisions (Convai) 0.5144 local English-only checkpoint
12 Laya-Multilingual (Convai) 0.4796 local below chance on this set

The top three sit inside the confidence interval (paired bootstrap; ZTC−JEV 95% CI [−0.034, +0.020]), so they are marked a tie. On a held-out set of answers from a never-seen model (Claude Haiku 4.5, 1,939 items), ZTC-Judge-27B v2 = 0.7752 beats JEV = 0.7521 (Δ +0.023, CI [+0.006, +0.041], excludes zero).

What is measured (the axis)

  • Question: "Is the ANSWER factually correct for the QUESTION?" → a probability of true.
  • Zero generation: verifiers emit a probability without writing a new answer. Token-spending LLM judges (GPT-5.2, Gemini 2.5 Flash-Lite, GPT-4o-mini, Qwen3-Next-80B) are kept in an off-axis reference, not the ranking.
  • A content-blind baseline (answer length & formatting only, AUC 0.6223) sits on the same board. A verifier below it did not read the content.
  • Divisions: open-weight (local, self-hostable) systems are never mixed with commercial APIs.
  • Fairness: every competitor is run on its own published code / head, never marked down by a re-implementation. A label-shuffle negative control reads 0.4996 (chance), validating the harness.
  • Calibration (ECE): does "0.8" mean 80% correct? Reported alongside AUC.
  • The 2,018 item texts stay private per source licenses (KMMLU CC BY-ND, CLIcK, GPQA); scores, labels and grading code are fully open.

The field — the "open-Jev" ecosystem

TypeSafe's JEV ("Decisions, Not Strings") started a category that is now a whole ecosystem: a CMU paper (JEV-as-a-Judge), dozens of open reproductions (open-jev, Bespoke Nimble, CLM-8B, decider, kev, von, ZefanCai Open-Jev, APUS-OpenJev, Laya, Manchego, Eikos, JevK5, Winnow, and 20+ more), competing leaderboards, and live Spaces. This board tracks 40+ distinct families and ranks the ones that are actually measurable on a neutral, labeled test set. Systems whose axis differs (game-playing bots, routers, constrained decoders, entity-extraction encoders) are listed off-ranking with the reason.

ZTC (Zero-Token Confidence) by VIDRAFT / FINAL-Bench is an open-weight verifier that reads a model's own hidden state to judge its answer — no generated tokens, no separate API.

한국어 요약

답변 검증기(타입드 디시전·System-1 스코어러) 리더보드. LLM의 답이 맞는지 토큰 생성 없이 판정하는 모델들을 같은 2,018문항·같은 라벨로 재고, 점수·라벨·채점 코드를 공개합니다. 각 경쟁 모델은 제작자 자신의 코드로 측정합니다.

  • 부문 분리: 오픈 웨이트(로컬) vs 상용 API
  • 기준선 동봉: 답 길이·서식만 보는 기준선(0.6223)을 못 넘으면 내용을 못 읽는 것
  • 현재 1위 그룹(무승부): JEV 0.7350 · ZTC-27B 0.7289 · ZTC-397B 0.7272
  • 처음 보는 답(Haiku)에선 ZTC가 JEV를 이김 (0.7752 vs 0.7521)
  • JEV·open-jev·Bespoke Nimble·CLM-8B·decider·Laya 등 40여 계열 추적

Keywords: JEV alternative, open-jev leaderboard, answer verification benchmark, typed decisions, hallucination detection, LLM judge, calibrated confidence, ECE, AUC, zero-token verifier, System-1 decision model, Bespoke Nimble, CLM-8B, decider-2b, Laya, ZTC, factuality checker, 답변 검증기, 환각 탐지, 리더보드.