Bekko System One: building ultra-small decision models that run in a browser
I wanted to see how much of that capability I could get from a much smaller model. The result is bekko-system-one-v0, shortened below to BS1. Jev remains well ahead on generalization. But on S1MB (System One Mosaic Benchmark), a benchmark I built, these small models produce results comparable to some much larger models. They are most useful on tasks from domains represented in their training data.
The models make three kinds of decisions: Noul for yes/no judgments, Choice for selecting an option, and Score for assigning a numeric rating.
Running in a browser
The models support English and come in three sizes: 17M, 68M, and 400M parameters. The 17M and 68M models are small enough to load into a browser and run on the CPU at practical speeds.
I deliberately chose demo examples that BS1 handles well. They make it look capable, but they should not be read as evidence of Jev-like performance across a broad range of tasks.
The 17M label counts all parameters. Excluding the embedding lookup table leaves about 4M, which I call active parameters (AP) here. For the browser export, I store only the token lookup table in INT8 and convert the model to ONNX. The model file itself is about 29 MB.
Benchmark results
Running in a browser is useful only if the model can do the job. I evaluated the models with S1MB, my benchmark for Noul, Choice, and Score decisions. It contains 137 benchmarks: 59 Noul, 57 Choice, and 21 Score. These draw mainly from the test sets of the datasets described later, generally sampling 100 cases per benchmark.
I describe the benchmark in the S1MB article. Here are the results restricted to models with at most 500M total parameters, with Jev 1.13 included for reference. BS1 does reasonably well within this small-model group.
| Model | Avg | Noul | Choice | Score | Total Params | Active Params |
|---|---|---|---|---|---|---|
| Jev 1.13 | 59.59 | 64.63 | 67.22 | 46.92 | — | — |
| bekko-system-one-v0-400m | 50.60 | 51.24 | 61.32 | 39.25 | 395M | 343M |
| bekko-system-one-v0-68m | 40.46 | 42.91 | 51.62 | 26.85 | 68M | 42M |
| bekko-system-one-v0-17m | 27.57 | 31.43 | 35.57 | 15.70 | 17M | 4M |
| von | 16.21 | 20.15 | 23.99 | 4.48 | 395M | 343M |
| laya-typed-decisions | 15.00 | 20.06 | 18.93 | 5.99 | 421M | 370M |
| laya | 13.36 | 20.19 | 16.31 | 3.58 | 421M | 370M |
| laya-multilingual | 9.28 | 14.01 | 13.03 | 0.79 | 322M | 125M |
Avg is the equal-weighted mean of the adjusted Noul, Choice, and Score scores. Higher is better. These are not raw accuracy percentages.
The generalization results tell a different story. On the synthetic benchmarks generated with GPT-6-Astra to test generalization, BS1 falls well behind Jev. My reading is that it has learned useful behavior within familiar task domains, but struggles with open-ended instructions and unfamiliar domains. That gap is why I am releasing it as version 0, with the -v0 suffix. It has not yet reached the generalization level I would want from a System One Decision Model.
Here are the same models on the six generalization benchmarks, sorted by Avg.
| Model | Avg | General Noul | General Choice | General Score |
|---|---|---|---|---|
| Jev 1.13 | 96.27 | 99.00 | 98.68 | 91.14 |
| bekko-system-one-v0-400m | 54.48 | 52.00 | 80.30 | 31.14 |
| von | 43.19 | 41.00 | 65.89 | 22.67 |
| bekko-system-one-v0-68m | 32.48 | 31.00 | 57.52 | 8.90 |
| laya-typed-decisions | 31.82 | 22.00 | 58.80 | 14.67 |
| laya | 26.50 | 9.00 | 54.83 | 15.67 |
| bekko-system-one-v0-17m | 18.96 | 15.00 | 39.11 | 2.76 |
| laya-multilingual | 14.47 | 6.00 | 36.67 | 0.75 |
Jev performs well across the domain-specific benchmarks as well as the generalization tests. On the generalization benchmarks it is close to the ceiling. For me, matching that level on this suite is a minimum target before claiming strong generalization for a System One Decision Model.
Model size and inference time
Size and speed are where BS1 is useful. On an RTX 5090, evaluating all 26,269 S1MB judgments took about 13.33 seconds for 17M, 34.04 seconds for 68M, and 132.97 seconds (about 2 minutes 13 seconds) for 400M. For 17M, the recorded intervals including model loading and per-benchmark data loading summed to about 21.38 seconds. The browser demo offers another way to try the small model on local hardware.
For Jev, accessed through its API, the sum of the per-benchmark evaluation times was about 58 minutes 37 seconds, including network waits. API response times can vary. These numbers measure evaluation work, not pure model computation. The 21.38-second figure for 17M also excludes warmup and other work outside the recorded intervals; it is not the time from starting the command to its completion.
Reusing the shared input
The main architectural choice in BS1 is to reuse the computation for the input shared by several candidate answers. Before describing that, it helps to look at two common ways to use an encoder for this kind of decision.
cross encoder
A conventional cross encoder pairs the shared input with each candidate and processes them together using a bidirectional encoder. The special tokens below are schematic.
[CLS] {state} {instruction} [SEP] {criteria A} [SEP]
[CLS] {state} {instruction} [SEP] {criteria B} [SEP]
[CLS] {state} {instruction} [SEP] {criteria C} [SEP]
A trained head produces a score from the leading CLS token's representation, or from a mean-pooled vector of token representations. This is also a common approach for retrieval rerankers.
The input and candidate can attend to each other, allowing the model to capture detailed relationships. But when the state is long, the same text is processed once for every candidate. The candidate affects the representation of the shared input, so identical input text does not imply that its intermediate representations can be reused.
In retrieval reranking, the query at the start is often short. There is less computation to save by sharing it, which makes a conventional cross encoder a reasonable choice in many cases.
listwise encoder
Another approach is to concatenate all candidates and process them in one pass:
{state} {instruction} [SEP]
{criteria A} [SEP] {criteria B} [SEP] {criteria C} [SEP]
Scores can be computed by pooling each candidate's representations, or by using special tokens around the candidates. The shared input is processed once, and candidates can attend to one another.
The tradeoff is input length. With full attention across all tokens, attention computation grows quadratically with sequence length. Concatenating many candidates can therefore become expensive. This does not make a listwise encoder universally slower than a cross encoder; the balance depends on input lengths and the number of candidates.
BS1's shared prefix
BS1 processes the state and instruction together as a shared prefix, then stores the keys and values (K/V) from each layer for reuse.
{state} {instruction} → shared K/V at each layer
├─ {criteria A} → score A
├─ {criteria B} → score B
└─ {criteria C} → score C
The difference from a conventional cross encoder is which tokens can attend to which. The shared prefix cannot attend to candidates. Each candidate can attend only to the shared prefix and its own tokens. Changing a candidate therefore does not change the prefix's intermediate representations, so its K/V can be computed once and reused.
This uses stored K/V, as an LLM's KV cache does, but it does not generate text one token at a time. After computing the prefix, BS1 processes candidates in parallel, mean-pools the candidate representations, and passes them to the scoring head.
Pooling only the candidate tokens does not mean the model sees only the candidate text. Their representations incorporate the state and instruction through attention to the prefix. A CLS token at the start of the shared prefix, on the other hand, cannot see the candidates and cannot by itself provide a candidate-specific judgment.
This avoids processing a long state and instruction from scratch for every candidate. It also changes the computation. In a conventional cross encoder, a candidate changes the shared input's representations, which can then feed back into the candidate. BS1 restricts that exchange so it can reuse the prefix computation.
A conventional cross encoder therefore has more freedom to model detailed interactions. How much that matters for accuracy depends on the task and training conditions.
The state and instruction are processed bidirectionally together. This is not a cache of the state alone: changing the instruction requires recomputing the prefix.
I did not include attention between candidates in this release. For Choice in particular, being able to compare candidates directly seems potentially useful. I tried it briefly, but did not obtain a training result that improved accuracy. That says something about the setup I tried, not that candidate interaction cannot work.
Choosing the base model
BS1 starts from Ettin-reranker. Its lineage is Ettin-encoder, an English encoder built on the ModernBERT architecture, followed by cross encoder training with a teacher reranker.
A model pretrained with masked language modeling, such as ModernBERT, has not yet been trained to score query-document relevance directly. Retrieval-trained dense embedding models and cross encoders have already learned relationships between queries and documents. That made Ettin-reranker, which has strong English reranking results for its size, a useful starting point for further training.
I compared Ettin-encoder and Ettin-reranker initialization under the same training conditions, using the full training pool after applying the per-dataset caps. At 68M, reranker initialization gave lower overall cross-entropy and higher accuracy for every decision type. At 17M, overall cross-entropy was nearly tied: the reranker was better on Noul and Choice accuracy, while the encoder was better on Score.
| Size | Initialization | Overall CE ↓ | Noul accuracy ↑ | Choice accuracy ↑ | Score accuracy ↑ |
|---|---|---|---|---|---|
| 17M | Ettin-encoder | 0.8694 | 68.71% | 60.14% | 56.57% |
| 17M | Ettin-reranker | 0.8713 | 69.88% | 60.77% | 53.88% |
| 68M | Ettin-encoder | 0.7625 | 72.87% | 69.78% | 58.16% |
| 68M | Ettin-reranker | 0.7180 | 77.01% | 71.13% | 61.34% |
Lower cross-entropy (CE) and higher accuracy are better. Results are aggregated per dataset and averaged over the two input orders. These metrics differ from the adjusted scores in the earlier tables. Each condition uses one seed.
Training data
I trained BS1 on hotchpotch/bekko-system-one-dataset-v0. It contains 153 training subsets, drawn from existing NLP and retrieval datasets and converted into Noul, Choice, and Score decisions. The training set contains about 6.59 million cases, with some synthetic datasets added as well.
The test sets also feed into S1MB. Since BS1 trains on the corresponding training domains, some of its S1MB scores benefit from that familiarity.
Training ran on an RTX 5090 and took about 2 hours for 17M, 5.5 hours for 68M, and 25 hours for 400M.
Training settings
I drew on the findings from my Bekko Embedding paper, which covers small multilingual dense retrieval models, and tried several parameter settings. These are the values that worked well within those experiments.
| Setting | 17M | 68M | 400M |
|---|---|---|---|
| Optimizer | AdamW | AdamW | AdamW |
| Batch size | 512 | 512 | 512 |
| Encoder learning rate | 1e-4 | 3e-5 | 1e-5 |
| Head learning rate | 5e-4 | 2e-4 | 2e-4 |
| Weight decay | 0.01 | 0.01 | 0.01 |
| Warmup ratio | 10% | 10% | 10% |
| Learning-rate schedule | Cosine | Cosine | Cosine |
| Query / candidate token limits | 4,096 / 2,048 tokens | 4,096 / 2,048 tokens | 4,096 / 2,048 tokens |
The training implementation is available if you want to train a similar model or continue the experiments with your own data.
Can a very small model generalize?
I think a System One Decision Model with fewer than 100M active parameters could eventually reach Jev-like scores on S1MB's generalization benchmarks. My view is that high-quality synthetic training data is the most important part of getting there.
TypeSafe AI CEO Diogo Almeida has described their data this way: “100% of our data is synthetic (but not the type of crap that is just spit out from an LLM obviously).” If we can build Noul, Choice, and Score training data covering a broad range of questions with appropriate targets, I think much smaller models could become useful across domains. Architecture attracts a lot of attention, but the harder work may be producing diverse, consistent data and finding how to train on it.
My bekko-embedding models are part of why I think this is worth pursuing. Training with large batches on a corpus of more than a billion examples, including about 160 million synthetic examples, produced retrieval models with about 8M and 25M active parameters that compete with models several to tens of times larger. That is a retrieval result, but it gives me a reason to keep investigating small models.
Stronger base models are another part of the picture. Making a very small decoder LLM with practical text-generation quality is difficult. If text generation is not required, there may be room for smaller, capable base models; ModernBERT is an example of that direction. Continued work on LLM architectures may also help small encoders trained on large, varied corpora, giving downstream training a better starting point. I still see synthetic data as the more important factor.
Closing thoughts
What I find compelling about Jev is the usefulness of a classification model that generalizes. TypeSafe AI set a high target: a System One Decision Model with the kind of broad capability I associate with frontier LLMs.
Classification covers more useful problems than it might first suggest. Since BERT, transformer encoders have been widely used for NLP tasks involving classification and scoring.
With bekko-system-one-v0, I explored how much useful accuracy on familiar tasks could fit into a model small enough to run in a browser. I am releasing the models, dataset, and training code as a starting point for trying small models on those tasks or training them further on your own data.
Handling a broad range of unfamiliar tasks remains open work. I expect better datasets and continued progress in small-model architectures to make more general System One Decision Models possible. I hope this release is a useful step for others who want to explore that direction.
References
bekko-system-one
- bekko-system-one collection
- bekko-system-one-v0-17m
- bekko-system-one-v0-68m
- bekko-system-one-v0-400m
- Browser demo
- Training and inference code
- Training dataset
S1MB
Base models and related work
- TypeSafe AI — System One / Jev
- TypeSafe AI — Models
- Diogo Almeida on synthetic data
- Laya: code / models
- ModernBERT-base
- Ettin: Seq vs Seq paper
- Ettin-reranker-17m-v1
- Ettin-reranker-68m-v1
- Ettin-reranker-400m-v1
- bekko-embedding-v1-a8m
- bekko-embedding-v1-a25m
- Bekko Embedding paper
- Sentence Transformers — Cross-Encoders


