S1MB: comparing System One Decision Models across 100+ benchmarks

Community Article
Published September 30, 2026

Since TypeSafe AI introduced Jev, its System One Decision Model, other developers have been working toward similar capabilities. I built a small model of my own, bekko-system-one-v0, and wanted to see how it compared with other models under the same evaluation conditions.

There are many ways to design a more comprehensive evaluation. What I needed first, though, was a practical reference point for comparing models. I put together tasks drawn from existing public NLP datasets and called the result System One Mosaic Benchmark, or S1MB. That collection of pieces is where “Mosaic” comes from.

The S1MB leaderboard for comparing System One Decision Models

I plan to expand both the benchmark suite and model coverage. The leaderboard shows the latest results; the screenshot and tables in this article capture a snapshot collected on September 30, 2026.

The evaluation code, dataset, and results are also available:

What it evaluates

System One Decision Models use text and instructions to make yes/no judgments, select an option, or assign a numeric score. S1MB evaluates these as Noul, Choice, and Score decisions, respectively.

This release covers English-language tasks. The evaluation data includes tasks from existing NLP datasets, tasks available through Open-Jev and Laya, and synthetic datasets I prepared for this benchmark. It is not exclusively a collection of the original datasets' official test sets. Within S1MB, however, the distributed data is an evaluation-only test split.

The English evaluation suite in this snapshot has the following composition:

Decision type What it does Benchmarks
Noul Makes a yes/no judgment 59
Choice Selects an option 57
Score Assigns a numeric score using supplied criteria 21
Total 137

In this snapshot, the data is organized into 106 subsets, containing 14,009 cases and 26,269 judgments. A subset is a collection of data that can contain more than one decision type. Evaluating each subset separately for its Noul, Choice, and Score tasks gives 137 benchmarks. A case can contain more than one judgment, so those counts differ. The six generalization benchmarks described later are included in the 137 benchmarks.

Overall results

These results cover 23 models and were collected on September 30, 2026. Every model completed all 137 benchmarks. The table is sorted by Borda Score, which converts relative model rankings into points. Task Avg averages the adjusted scores for the three decision types. Neither is raw accuracy; I explain both below the table.

The models were evaluated at different times. As models and benchmarks are added, the comparison scope and scores can change; these tables retain the scope of the snapshot above.

# Model Borda Score Task Avg Noul Choice Score TP AP
1 Jev 1.13 86.51 59.59 64.63 67.22 46.92 — —
2 bekko-system-one-v0-400m 76.21 50.60 51.24 61.32 39.25 395M 343M
3 Open-Jev-9B (ZefanCai) 72.54 52.10 55.76 55.51 45.04 7.94B 6.93B
4 Open-Jev-27B-v1.1 (ZefanCai) 72.35 56.88 51.48 66.91 52.26 25.64B 24.37B
5 Tev1-4B-experimental 71.67 47.07 50.73 53.36 37.11 4.54B 4.54B
6 kev-9b 70.52 42.40 50.21 56.36 20.64 7.94B 6.92B
7 kev-4b 68.81 42.33 48.48 52.03 26.47 4.21B 3.57B
8 JevK5 67.49 41.14 48.93 52.84 21.65 4.21B 4.21B
9 OpenJev-4B (Alex) 63.12 36.46 42.77 47.49 19.11 4.54B 3.9B
10 bekko-system-one-v0-68m 62.23 40.46 42.91 51.62 26.85 68M 42M
11 Open-Jev-2B (ZefanCai) 50.56 34.64 38.83 39.82 25.27 1.88B 1.38B
12 bekko-system-one-v0-17m 46.57 27.57 31.43 35.57 15.70 17M 4M
13 OpenJev-2B (Alex) 44.38 25.21 29.72 37.80 8.11 2.21B 1.7B
14 decider-0.8b 42.65 22.57 28.97 32.36 6.39 752M 752M
15 Tev1-0.8B-experimental 38.59 20.43 25.97 27.15 8.16 853M 851M
16 kev-0.8b 35.20 18.68 23.61 27.36 5.07 753M 499M
17 von 32.40 16.21 20.15 23.99 4.48 395M 343M
18 OpenJev-0.8B (Alex) 31.40 16.02 18.90 23.79 5.36 853M 597M
19 laya-typed-decisions 31.22 15.00 20.06 18.93 5.99 421M 370M
20 laya 29.31 13.36 20.19 16.31 3.58 421M 370M
21 laya-multilingual 24.14 9.28 14.01 13.03 0.79 322M 125M
22 minojev 19.99 7.93 8.57 14.56 0.67 1.72B 1.72B
23 JevForge-0.8B 12.14 1.92 0.00 5.75 0.00 854M 598M

TP is the total parameter count. AP excludes embeddings used only for lookup. This is not the MoE meaning of active parameters, where a different set of experts may be used for each input.

Borda Score and Task Avg

For each benchmark, Borda Score assigns 100 points to first place and 0 to last place, with evenly spaced points between them. Ties use average ranks. The points are then averaged equally across all benchmarks in the selected evaluation scope. It measures how consistently a model ranks near the top of the comparison, so changing the model roster or benchmark set can change its value.

Task Avg and the Noul, Choice, and Score columns instead use scores adjusted against simple predictions that ignore the input. A model performing at or below a baseline such as always choosing the same option or returning the same number receives 0; the reference ceiling is 100. A score of 50 therefore means closing half the gap from the baseline to the reference ceiling, not getting half the answers right.

The adjustment differs by task. For Noul, balanced accuracy of 50% maps to 0, and 100% maps to 100; values below the baseline are clipped to 0. Choice compares against the best of uniform random and fixed-answer baselines. Score compares mean absolute error (MAE) against a constant prediction.

The adjusted values are clipped to 0–100 for each benchmark and averaged within each decision type. Task Avg is then the equal-weighted mean of Noul, Choice, and Score. The details are in the scoring definitions.

Borda gives equal weight to each benchmark. In this snapshot, with 59 Noul benchmarks and 21 Score benchmarks, those decision types therefore have different weights in the overall result. Task Avg weights the three types equally, so the rankings need not agree. For example, Open-Jev-27B-v1.1 (ZefanCai) ranks second by Task Avg at 56.88, but fourth by Borda Score. The individual task columns help show what an overall ranking leaves out.

Reading the overall results

Jev 1.13 ranks first with a Borda Score of 86.51. My bekko-system-one-v0-400m is second at 76.21.

There is relevant training history behind that result. My model uses training sets from datasets that also appear in this evaluation. It has not trained on the test sets, but familiarity with the tasks and domains makes high scores on their test sets more likely. The training manifest and S1MB evaluation manifest have 77 subset names in common. That is worth keeping in mind when reading the results.

Even with that familiarity, Jev 1.13 scores higher on both overall Borda Score and Task Avg. Seeing that gap, despite my model having trained on many of the same domains, was one of the surprises of running this comparison.

An overall ranking does not tell us whether a model can handle unfamiliar instructions, though. I wanted to look at that separately.

Evaluating generalization

Performance on an existing dataset can depend on how much experience a model has with that task. What I also want from a System One Decision Model is the ability to adapt when instructions, context, or decision criteria change. To probe that, I used GPT-6-Astra to create a synthetic evaluation set.

For each of Noul, Choice, and Score, I prepared two kinds of benchmarks: Diverse and Contextual.

Kind Content Noul Choice Score Total
Diverse Varied instructions, contexts, and decision criteria 100 100 100 300
Contextual Related groups of questions with changes in context and other conditions 100 100 100 300
Total 6 benchmarks 200 200 200 600

Each Contextual task contains 20 groups of five questions. For Choice and Score, the question and criteria stay fixed while the state changes. For Noul, both the state and the false/true criteria change. The intention is to see whether a model adapts its judgment to the supplied conditions instead of relying on isolated words.

Here is a summarized example from Diverse Choice:

Field Content
State The account is still in use. The user wants to stop weekly promotional emails but keep security alerts and invoices.
Instruction Select the communication the user wants to stop.
Example options All communications, security alerts, invoices, promotional email, account access, and others
Intended answer Promotional email

The same model generated the questions and intended answers and checked its own work. They have not undergone independent human validation. Nor do I guarantee that the evaluated models' training data contains no overlap. What this set measures is adaptation to the particular instructions, contexts, and criteria supplied here.

Generalization results

This table contains the same 23 models, evaluated only on those six benchmarks. Borda Score is recomputed over the six benchmarks, so its values differ from the overall table. The rows are again sorted by Borda Score.

# Model Borda Score Task Avg General Noul General Choice General Score
1 Jev 1.13 97.73 96.27 99.00 98.68 91.14
2 Open-Jev-9B (ZefanCai) 88.26 85.63 93.00 96.70 67.21
3 Open-Jev-27B-v1.1 (ZefanCai) 87.50 82.97 93.00 100.00 55.91
4 JevK5 85.98 85.01 92.00 93.39 69.64
5 kev-9b 83.71 84.58 87.00 96.05 70.69
6 Tev1-4B-experimental 83.33 84.07 88.00 94.05 70.15
7 kev-4b 73.48 77.50 83.00 93.44 56.06
8 OpenJev-4B (Alex) 72.73 76.75 82.00 94.12 54.13
9 Open-Jev-2B (ZefanCai) 60.61 65.57 71.00 85.56 40.15
10 OpenJev-2B (Alex) 58.71 64.08 70.00 88.85 33.39
11 decider-0.8b 50.76 55.92 57.00 82.90 27.86
12 bekko-system-one-v0-400m 48.11 54.48 52.00 80.30 31.14
13 Tev1-0.8B-experimental 45.45 54.61 59.00 76.40 28.43
14 kev-0.8b 43.18 51.22 53.00 75.15 25.51
15 OpenJev-0.8B (Alex) 42.05 52.33 57.00 77.09 22.89
16 von 31.44 43.19 41.00 65.89 22.67
17 bekko-system-one-v0-68m 21.21 32.48 31.00 57.52 8.90
18 laya-typed-decisions 20.83 31.82 22.00 58.80 14.67
19 minojev 18.18 25.08 11.00 63.37 0.87
20 laya 17.80 26.50 9.00 54.83 15.67
21 bekko-system-one-v0-17m 10.23 18.96 15.00 39.11 2.76
22 laya-multilingual 7.58 14.47 6.00 36.67 0.75
23 JevForge-0.8B 1.14 3.74 0.00 11.22 0.00

Jev 1.13 reaches a Task Avg of 96.27 and matches the intended answers on almost every Noul and Choice question. Checking the saved predictions, Noul matches on 199 of 200 questions (99.5%) and Choice on 198 of 200 (99.0%). The table's General Noul score of 99.00 and General Choice score of 98.68 are adjusted scores, so they differ from those raw accuracies.

That consistency as instructions and context change is what stands out to me about Jev 1.13.

Open-Jev-27B-v1.1 (ZefanCai) and Open-Jev-9B (ZefanCai) also perform well. Both score 93.00 on General Noul, and their General Choice scores are 100.00 and 96.70, respectively. Their General Score results are 55.91 and 67.21, however, so neither has reached the ceiling across all three decision types.

How I would read these scores

This is a synthetic set I prepared with GPT-6-Astra, and the questions are not particularly complex. As its author, I expect models with some generalization capability to approach the ceiling, especially on Noul and Choice. I would use it as a reference point for asking: can a model get close to the maximum score on this level of variation in instructions and context? Small gaps between the leading models are not a basis for finely ranking their generalization abilities.

Here, approaching the maximum means getting close to 100 on the adjusted task scores, not on Borda's relative ranking points. A high score does not mean a model can handle any unfamiliar task. These 600 questions also do not tell us how the models would separate on harder problems.

For Score, the evaluation measures agreement with criteria and intended scores generated by GPT-6-Astra. Jev 1.13's 91.14 is not an accuracy percentage: it compares error in the expected numeric score against a constant prediction. Exact agreement with an intended score is not necessarily the only valid real-world judgment, either.

Meanwhile, bekko-system-one-v0-400m, second in the overall Borda ranking, has a Task Avg of 54.48 on these six benchmarks. Scoring well on familiar task families and adapting to freely specified instructions and changing contexts are worth examining separately.

Closing thoughts

S1MB started from wanting to evaluate my own small model. Running the comparisons showed me how different the picture can look between an overall ranking and the generalization tests.

Both the collection of existing datasets and these relatively simple synthetic tasks have limited scope. Even so, comparing models on the same inputs and metrics, and seeing which decision types or tasks they struggle with, gives me a more concrete basis for deciding what to improve next.

Instructions for evaluating your own model and adding its results to the leaderboard are available. After validating and exporting the results, you can submit them through a pull request to the results dataset.

I hope S1MB gives other people building System One Decision Models a useful reference point for their own experiments.

References

Community

Sign up or log in to comment