π Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).
S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.
π Top of the board π₯ Darwin ZTC v2 (FINAL-Bench) 89.58 π₯ OpenJev-27B 87.50 π₯ AutoJev-27B 87.07 4οΈβ£ Eikos 27B 85.43 5οΈβ£ Jev 1.13 85.05
π Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.
βοΈ Why a zero-token judge wins here πΉ It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass. πΉ Zero generated tokens, no decoding loop, so latency and cost stay low. πΉ Holds up out of distribution too: General Noul 96.00, General Choice 99.34.
It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.
π Darwin-27B-ZTC-v2 just took #1 on the System One Mosaic Benchmark (S1MB).
S1MB compares 102 models across 137 specialized benchmarks, in three task types: Noul (assess a condition), Choice (select an option), Score (rate on a scale). Ranking is by overall Borda score.
π Top of the board π₯ Darwin ZTC v2 (FINAL-Bench) 89.58 π₯ OpenJev-27B 87.50 π₯ AutoJev-27B 87.07 4οΈβ£ Eikos 27B 85.43 5οΈβ£ Jev 1.13 85.05
π Ranks 2 to 5 are all the JEV family (TypeSafe AI's System One model, from ex-OpenAI researchers). S1MB exists to compare these System One judges, so leading it is the headline.
βοΈ Why a zero-token judge wins here πΉ It does not generate. It reads the input and typed questions and returns a calibrated distribution in a single forward pass. πΉ Zero generated tokens, no decoding loop, so latency and cost stay low. πΉ Holds up out of distribution too: General Noul 96.00, General Choice 99.34.
It is also #1 on the typed-decisions leaderboard (0.743, zero-shot). Same message from both: a deterministic, calibrated judge at one forward pass per call.
π§ We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
βοΈ How it works πΉ It makes its call in a single forward pass. πΉ Zero generated tokens, and no decoding loop. πΉ That keeps latency and cost far below what a generative model needs.
π― What it judges πΉ It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). πΉ For each one it hands back a calibrated confidence, not just an answer.
π How well calibrated (measured) πΉ KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. πΉ 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. πΉ By type: noul 0.847, choice 0.723, score 0.675. πΉ None of the benchmark's train split went into it. It is pure zero-shot.
π Where it fits πΉ Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
π It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).