Leading the System One Mosaic Benchmark: What Darwin-27B-ZTC-v2's #1 Means

Community Article
Published October 9, 2026

A new category: models that decide instead of chat

A distinct class of models has emerged in 2026: "System One" models that do not generate free text at all. Instead of writing an answer token by token, they read a state and a set of typed questions and return structured outputs, a label and a calibrated probability, in a single forward pass. TypeSafe AI's JEV, from former OpenAI researchers, popularized the framing. The appeal is operational: for routing, grading, and gating inside agent systems, you do not want prose, you want a fast, typed decision with a confidence you can threshold.

The System One Mosaic Benchmark (S1MB, maintained by hotchpotch) was built to measure exactly these models. This article looks at how it scores, where Darwin-27B-ZTC-v2 lands, and why a zero-token judge does well on it.

How S1MB scores

S1MB aggregates 137 specialized benchmarks over three task families:

  • Noul: assess a condition (open-ended correctness of a statement).
  • Choice: select the correct option among candidates.
  • Score: rate on a scale (ordinal judgment).

These three are not the same problem. Noul is open-ended verification, Choice is categorical classification, and Score is closer to ordinal regression. A single model that must cover all three cannot overfit to one decision geometry.

The leaderboard ranks by a Borda score. Rather than averaging raw task numbers, which lets one easy benchmark dominate, Borda converts each benchmark into a rank and sums rank-based points across all 137. That rewards broad, consistent strength over a spike on a few favorable tasks. It is a deliberately conservative aggregation, which makes a top Borda position harder to game than a top raw average.

Results

As of this writing, Darwin-27B-ZTC-v2 is #1 overall.

# Model Borda Task Avg Noul Choice Score
1 Darwin ZTC v2 (FINAL-Bench) 89.58 66.46 67.66 71.51 60.21
2 OpenJev-27B 87.50 62.60 65.22 67.74 54.83
3 AutoJev-27B 87.07 60.80 64.96 68.11 49.34
4 Eikos 27B 85.43 59.86 64.14 67.61 47.83
5 Jev 1.13 85.05 59.59 64.63 67.22 46.92

Two things stand out. First, the margin at the top is driven by Task Avg (66.46 vs 62.60 for #2) and by the Score column (60.21 vs 54.83), the hardest of the three families. Ordinal scoring is where most judges fall off, and it is where the gap is widest. Second, ranks 2 through 5 are all JEV family models (OpenJev, AutoJev, and JEV itself). S1MB is the home benchmark for the System One category, so leading it against the models that defined the category is the substantive result, not a win on a niche test.

On generalization splits the model stays high: General Noul 96.00, General Choice 99.34, General Score 88.35, with coverage across all 137 benchmarks. The generalization numbers matter because a judge is only useful if it holds on inputs it was not tuned for.

Why a zero-token judge fits S1MB

Darwin-27B-ZTC is a zero-token classifier. It does not sample. It reads the input and the typed questions and emits a probability distribution directly, in one forward pass. Three properties follow, and all three line up with what S1MB measures.

  • Determinism. No sampling means no output variance. The same input returns the same distribution, so the reported numbers are exact values rather than sample estimates. A generative judge reporting a mean over samples carries a standard deviation that S1MB's typed scoring does not have to absorb here.
  • Single-pass cost. Latency does not scale with output length because there is no decoding loop. For a benchmark of 137 suites, and for real deployments doing millions of judgments, this is a structural cost difference, not a constant factor.
  • Calibrated confidence. Because the model emits a distribution, it can be scored on whether its confidence is honest, not just whether it is right. On the related typed-decisions leaderboard the model is also #1 (accuracy 0.743, zero-shot), with calibration KL 0.204 and Brier 0.097. Brier is a strictly proper scoring rule, so those numbers cannot be gamed by inflating or deflating confidence.

The match is not an accident. S1MB is itself a suite of typed decisions. A model that natively outputs typed distributions is answering in the benchmark's own language, while a generative model has to be coerced into that shape through prompting or logprob extraction, both of which add noise and cost.

Reading the result honestly

Public leaderboards move. New models are added, and a top position is a snapshot, not a permanent title. The numbers here reflect the board at the time of writing, and we state the rank with its date rather than as an absolute claim.

It is also worth being precise about scope. S1MB measures typed decision quality. It does not measure open-ended generation, long-form reasoning, or tool use. Darwin-27B-ZTC is a judge, not a chat model, and this result is about the judge category specifically. For generation we point to separate models and separate benchmarks.

Takeaway

The System One category is being valued as a distinct market, and the benchmark built to measure it now has a zero-token judge from FINAL-Bench at the top, ahead of the models that named the category. The lead is widest on the hardest task family (Score) and holds on generalization splits, which is the profile you want from a judge you intend to put in front of a router or a safety gate.

Links

Note: public leaderboard standings change as new models are added. All numbers reflect the board at the time of writing.

Community

Sign up or log in to comment