S1MB: comparing System One Decision Models across 100+ benchmarks
There are many ways to design a more comprehensive evaluation. What I needed first, though, was a practical reference point for comparing models. I put together tasks drawn from existing public NLP datasets and called the result System One Mosaic Benchmark, or S1MB. That collection of pieces is where “Mosaic” comes from.
I plan to expand both the benchmark suite and model coverage. The leaderboard shows the latest results; the screenshot and tables in this article capture a snapshot collected on September 30, 2026.
The evaluation code, dataset, and results are also available:
What it evaluates
System One Decision Models use text and instructions to make yes/no judgments, select an option, or assign a numeric score. S1MB evaluates these as Noul, Choice, and Score decisions, respectively.
This release covers English-language tasks. The evaluation data includes tasks from existing NLP datasets, tasks available through Open-Jev and Laya, and synthetic datasets I prepared for this benchmark. It is not exclusively a collection of the original datasets' official test sets. Within S1MB, however, the distributed data is an evaluation-only test split.
The English evaluation suite in this snapshot has the following composition:
| Decision type | What it does | Benchmarks |
|---|---|---|
| Noul | Makes a yes/no judgment | 59 |
| Choice | Selects an option | 57 |
| Score | Assigns a numeric score using supplied criteria | 21 |
| Total | 137 |
In this snapshot, the data is organized into 106 subsets, containing 14,009 cases and 26,269 judgments. A subset is a collection of data that can contain more than one decision type. Evaluating each subset separately for its Noul, Choice, and Score tasks gives 137 benchmarks. A case can contain more than one judgment, so those counts differ. The six generalization benchmarks described later are included in the 137 benchmarks.
Overall results
These results cover 23 models and were collected on September 30, 2026. Every model completed all 137 benchmarks. The table is sorted by Borda Score, which converts relative model rankings into points. Task Avg averages the adjusted scores for the three decision types. Neither is raw accuracy; I explain both below the table.
The models were evaluated at different times. As models and benchmarks are added, the comparison scope and scores can change; these tables retain the scope of the snapshot above.
| # | Model | Borda Score | Task Avg | Noul | Choice | Score | TP | AP |
|---|---|---|---|---|---|---|---|---|
| 1 | Jev 1.13 | 86.51 | 59.59 | 64.63 | 67.22 | 46.92 | — | — |
| 2 | bekko-system-one-v0-400m | 76.21 | 50.60 | 51.24 | 61.32 | 39.25 | 395M | 343M |
| 3 | Open-Jev-9B (ZefanCai) | 72.54 | 52.10 | 55.76 | 55.51 | 45.04 | 7.94B | 6.93B |
| 4 | Open-Jev-27B-v1.1 (ZefanCai) | 72.35 | 56.88 | 51.48 | 66.91 | 52.26 | 25.64B | 24.37B |
| 5 | Tev1-4B-experimental | 71.67 | 47.07 | 50.73 | 53.36 | 37.11 | 4.54B | 4.54B |
| 6 | kev-9b | 70.52 | 42.40 | 50.21 | 56.36 | 20.64 | 7.94B | 6.92B |
| 7 | kev-4b | 68.81 | 42.33 | 48.48 | 52.03 | 26.47 | 4.21B | 3.57B |
| 8 | JevK5 | 67.49 | 41.14 | 48.93 | 52.84 | 21.65 | 4.21B | 4.21B |
| 9 | OpenJev-4B (Alex) | 63.12 | 36.46 | 42.77 | 47.49 | 19.11 | 4.54B | 3.9B |
| 10 | bekko-system-one-v0-68m | 62.23 | 40.46 | 42.91 | 51.62 | 26.85 | 68M | 42M |
| 11 | Open-Jev-2B (ZefanCai) | 50.56 | 34.64 | 38.83 | 39.82 | 25.27 | 1.88B | 1.38B |
| 12 | bekko-system-one-v0-17m | 46.57 | 27.57 | 31.43 | 35.57 | 15.70 | 17M | 4M |
| 13 | OpenJev-2B (Alex) | 44.38 | 25.21 | 29.72 | 37.80 | 8.11 | 2.21B | 1.7B |
| 14 | decider-0.8b | 42.65 | 22.57 | 28.97 | 32.36 | 6.39 | 752M | 752M |
| 15 | Tev1-0.8B-experimental | 38.59 | 20.43 | 25.97 | 27.15 | 8.16 | 853M | 851M |
| 16 | kev-0.8b | 35.20 | 18.68 | 23.61 | 27.36 | 5.07 | 753M | 499M |
| 17 | von | 32.40 | 16.21 | 20.15 | 23.99 | 4.48 | 395M | 343M |
| 18 | OpenJev-0.8B (Alex) | 31.40 | 16.02 | 18.90 | 23.79 | 5.36 | 853M | 597M |
| 19 | laya-typed-decisions | 31.22 | 15.00 | 20.06 | 18.93 | 5.99 | 421M | 370M |
| 20 | laya | 29.31 | 13.36 | 20.19 | 16.31 | 3.58 | 421M | 370M |
| 21 | laya-multilingual | 24.14 | 9.28 | 14.01 | 13.03 | 0.79 | 322M | 125M |
| 22 | minojev | 19.99 | 7.93 | 8.57 | 14.56 | 0.67 | 1.72B | 1.72B |
| 23 | JevForge-0.8B | 12.14 | 1.92 | 0.00 | 5.75 | 0.00 | 854M | 598M |
TP is the total parameter count. AP excludes embeddings used only for lookup. This is not the MoE meaning of active parameters, where a different set of experts may be used for each input.
Borda Score and Task Avg
For each benchmark, Borda Score assigns 100 points to first place and 0 to last place, with evenly spaced points between them. Ties use average ranks. The points are then averaged equally across all benchmarks in the selected evaluation scope. It measures how consistently a model ranks near the top of the comparison, so changing the model roster or benchmark set can change its value.
Task Avg and the Noul, Choice, and Score columns instead use scores adjusted against simple predictions that ignore the input. A model performing at or below a baseline such as always choosing the same option or returning the same number receives 0; the reference ceiling is 100. A score of 50 therefore means closing half the gap from the baseline to the reference ceiling, not getting half the answers right.
The adjustment differs by task. For Noul, balanced accuracy of 50% maps to 0, and 100% maps to 100; values below the baseline are clipped to 0. Choice compares against the best of uniform random and fixed-answer baselines. Score compares mean absolute error (MAE) against a constant prediction.
The adjusted values are clipped to 0–100 for each benchmark and averaged within each decision type. Task Avg is then the equal-weighted mean of Noul, Choice, and Score. The details are in the scoring definitions.
Borda gives equal weight to each benchmark. In this snapshot, with 59 Noul benchmarks and 21 Score benchmarks, those decision types therefore have different weights in the overall result. Task Avg weights the three types equally, so the rankings need not agree. For example, Open-Jev-27B-v1.1 (ZefanCai) ranks second by Task Avg at 56.88, but fourth by Borda Score. The individual task columns help show what an overall ranking leaves out.
Reading the overall results
Jev 1.13 ranks first with a Borda Score of 86.51. My bekko-system-one-v0-400m is second at 76.21.
There is relevant training history behind that result. My model uses training sets from datasets that also appear in this evaluation. It has not trained on the test sets, but familiarity with the tasks and domains makes high scores on their test sets more likely. The training manifest and S1MB evaluation manifest have 77 subset names in common. That is worth keeping in mind when reading the results.
Even with that familiarity, Jev 1.13 scores higher on both overall Borda Score and Task Avg. Seeing that gap, despite my model having trained on many of the same domains, was one of the surprises of running this comparison.
An overall ranking does not tell us whether a model can handle unfamiliar instructions, though. I wanted to look at that separately.
Evaluating generalization
Performance on an existing dataset can depend on how much experience a model has with that task. What I also want from a System One Decision Model is the ability to adapt when instructions, context, or decision criteria change. To probe that, I used GPT-6-Astra to create a synthetic evaluation set.
For each of Noul, Choice, and Score, I prepared two kinds of benchmarks: Diverse and Contextual.
| Kind | Content | Noul | Choice | Score | Total |
|---|---|---|---|---|---|
| Diverse | Varied instructions, contexts, and decision criteria | 100 | 100 | 100 | 300 |
| Contextual | Related groups of questions with changes in context and other conditions | 100 | 100 | 100 | 300 |
| Total | 6 benchmarks | 200 | 200 | 200 | 600 |
Each Contextual task contains 20 groups of five questions. For Choice and Score, the question and criteria stay fixed while the state changes. For Noul, both the state and the false/true criteria change. The intention is to see whether a model adapts its judgment to the supplied conditions instead of relying on isolated words.
Here is a summarized example from Diverse Choice:
| Field | Content |
|---|---|
| State | The account is still in use. The user wants to stop weekly promotional emails but keep security alerts and invoices. |
| Instruction | Select the communication the user wants to stop. |
| Example options | All communications, security alerts, invoices, promotional email, account access, and others |
| Intended answer | Promotional email |
The same model generated the questions and intended answers and checked its own work. They have not undergone independent human validation. Nor do I guarantee that the evaluated models' training data contains no overlap. What this set measures is adaptation to the particular instructions, contexts, and criteria supplied here.
Generalization results
This table contains the same 23 models, evaluated only on those six benchmarks. Borda Score is recomputed over the six benchmarks, so its values differ from the overall table. The rows are again sorted by Borda Score.
| # | Model | Borda Score | Task Avg | General Noul | General Choice | General Score |
|---|---|---|---|---|---|---|
| 1 | Jev 1.13 | 97.73 | 96.27 | 99.00 | 98.68 | 91.14 |
| 2 | Open-Jev-9B (ZefanCai) | 88.26 | 85.63 | 93.00 | 96.70 | 67.21 |
| 3 | Open-Jev-27B-v1.1 (ZefanCai) | 87.50 | 82.97 | 93.00 | 100.00 | 55.91 |
| 4 | JevK5 | 85.98 | 85.01 | 92.00 | 93.39 | 69.64 |
| 5 | kev-9b | 83.71 | 84.58 | 87.00 | 96.05 | 70.69 |
| 6 | Tev1-4B-experimental | 83.33 | 84.07 | 88.00 | 94.05 | 70.15 |
| 7 | kev-4b | 73.48 | 77.50 | 83.00 | 93.44 | 56.06 |
| 8 | OpenJev-4B (Alex) | 72.73 | 76.75 | 82.00 | 94.12 | 54.13 |
| 9 | Open-Jev-2B (ZefanCai) | 60.61 | 65.57 | 71.00 | 85.56 | 40.15 |
| 10 | OpenJev-2B (Alex) | 58.71 | 64.08 | 70.00 | 88.85 | 33.39 |
| 11 | decider-0.8b | 50.76 | 55.92 | 57.00 | 82.90 | 27.86 |
| 12 | bekko-system-one-v0-400m | 48.11 | 54.48 | 52.00 | 80.30 | 31.14 |
| 13 | Tev1-0.8B-experimental | 45.45 | 54.61 | 59.00 | 76.40 | 28.43 |
| 14 | kev-0.8b | 43.18 | 51.22 | 53.00 | 75.15 | 25.51 |
| 15 | OpenJev-0.8B (Alex) | 42.05 | 52.33 | 57.00 | 77.09 | 22.89 |
| 16 | von | 31.44 | 43.19 | 41.00 | 65.89 | 22.67 |
| 17 | bekko-system-one-v0-68m | 21.21 | 32.48 | 31.00 | 57.52 | 8.90 |
| 18 | laya-typed-decisions | 20.83 | 31.82 | 22.00 | 58.80 | 14.67 |
| 19 | minojev | 18.18 | 25.08 | 11.00 | 63.37 | 0.87 |
| 20 | laya | 17.80 | 26.50 | 9.00 | 54.83 | 15.67 |
| 21 | bekko-system-one-v0-17m | 10.23 | 18.96 | 15.00 | 39.11 | 2.76 |
| 22 | laya-multilingual | 7.58 | 14.47 | 6.00 | 36.67 | 0.75 |
| 23 | JevForge-0.8B | 1.14 | 3.74 | 0.00 | 11.22 | 0.00 |
Jev 1.13 reaches a Task Avg of 96.27 and matches the intended answers on almost every Noul and Choice question. Checking the saved predictions, Noul matches on 199 of 200 questions (99.5%) and Choice on 198 of 200 (99.0%). The table's General Noul score of 99.00 and General Choice score of 98.68 are adjusted scores, so they differ from those raw accuracies.
That consistency as instructions and context change is what stands out to me about Jev 1.13.
Open-Jev-27B-v1.1 (ZefanCai) and Open-Jev-9B (ZefanCai) also perform well. Both score 93.00 on General Noul, and their General Choice scores are 100.00 and 96.70, respectively. Their General Score results are 55.91 and 67.21, however, so neither has reached the ceiling across all three decision types.
How I would read these scores
This is a synthetic set I prepared with GPT-6-Astra, and the questions are not particularly complex. As its author, I expect models with some generalization capability to approach the ceiling, especially on Noul and Choice. I would use it as a reference point for asking: can a model get close to the maximum score on this level of variation in instructions and context? Small gaps between the leading models are not a basis for finely ranking their generalization abilities.
Here, approaching the maximum means getting close to 100 on the adjusted task scores, not on Borda's relative ranking points. A high score does not mean a model can handle any unfamiliar task. These 600 questions also do not tell us how the models would separate on harder problems.
For Score, the evaluation measures agreement with criteria and intended scores generated by GPT-6-Astra. Jev 1.13's 91.14 is not an accuracy percentage: it compares error in the expected numeric score against a constant prediction. Exact agreement with an intended score is not necessarily the only valid real-world judgment, either.
Meanwhile, bekko-system-one-v0-400m, second in the overall Borda ranking, has a Task Avg of 54.48 on these six benchmarks. Scoring well on familiar task families and adapting to freely specified instructions and changing contexts are worth examining separately.
Closing thoughts
S1MB started from wanting to evaluate my own small model. Running the comparisons showed me how different the picture can look between an overall ranking and the generalization tests.
Both the collection of existing datasets and these relatively simple synthetic tasks have limited scope. Even so, comparing models on the same inputs and metrics, and seeing which decision types or tasks they struggle with, gives me a more concrete basis for deciding what to improve next.
Instructions for evaluating your own model and adding its results to the leaderboard are available. After validating and exporting the results, you can submit them through a pull request to the results dataset.
I hope S1MB gives other people building System One Decision Models a useful reference point for their own experiments.


