State
Questions
Answers
TypeSafe Jev · same request
Raw request / response
Test bench — hand-written questions the models were never trained on
Every question has a known correct answer written by hand. Choice: the picked option must match. Yes/no: P(yes) must be on the right side of 50 %. Score: the expected level must be within the stated tolerance. "p(gold)" is the probability the model gave to the correct answer (1.0 = certain and right).
Summary
Questions
Run history (saved in this browser)
Model versions
Fit test — "would you describe X as …?"
The same question format the Almost Certain demo sends to Jev, but with your own description, context and candidates. Each candidate is one yes/no question; all of them run in a single batched request.
P(yes) per candidate
Raw request
Compare with TypeSafe Jev (via the public "Almost Certain" demo)
Each row asks almost-certain.vercel.app how well a guess fits a hidden challenge description (a yes/no question, e.g. "would you describe a hat as a disappointing souvenir?"). Their reply includes the exact request Jev received and Jev's probability; we replay the same request on our model and show both. Challenge ids are 0–49-ish; unknown ids return an error row.
| # | challenge | guess | description (from Jev's request) | context | Jev | ours | Δ | agree |
|---|
Last replayed request (what both models were asked)
What is Certus?
Certus is a Jev-like "System One" decision model: it reads a state (any text or JSON) and
typed questions, and returns calibrated probability distributions over answers you define. It never generates
text — every answer is one of your options, produced in a single forward pass, so outputs are always well-typed
and the cost is fixed. The wire format is the same as TypeSafe's POST /v1/systemone.
| primitive | you give | you get |
|---|---|---|
choice | N named options, each with an optional description | choice, a probability per option, confidence |
score | an ordered rubric of levels ("calm → annoyed → angry") | score = expected level (fractional), a probability per level |
noul | a yes/no question | noul = P(yes) |
How it works
Base model Qwen2.5-1.5B-Instruct (Apache-2.0) + a LoRA trunk fine-tuned on 61 public
datasets with a proper-scoring loss (NLL + Brier) and temperature calibration, + eight domain adapters
(reasoning, math, language, support, sentiment, safety, spam, workflow) kept separate on top of the merged trunk.
auto runs a router on the whole request — every question with its type and options, then the
state — applies the adapters whose heads fire (at most two, stacked) and otherwise answers with the trunk.
The answer's model field shows which adapter answered. Adding an adapter later is one folder and one
router head; nothing else is retrained.
Results (hand-written held-out sets, never trained on)
| set | trunk v0007 | auto (router + adapters) | Jev |
|---|---|---|---|
| 196-question test bench, 20 categories | 0.842 | 0.949 | – |
| typed decisions (Laya benchmark, JSON workflow states) | 0.526 | 0.796 | 0.727 |
| 10 hard yes/no logic traps | 4/10 | 9/10 | 10/10 |
| 10 mixed hard questions (arithmetic, units, fallacies, dates, sarcasm) | 4/10 | 6/10 | 6/10 |
| 15-task unseen public suite, mean accuracy / ECE | 0.736 / 0.074 | – | – |
Speed and this Space
On a single consumer GPU (RTX 2080 Ti) a request answers in 40–70 ms, about 110–200 ms when the router applies an adapter. This Space runs on ZeroGPU: the model is materialised on a shared GPU per request (about 0.5 s round trip, longer after a pause); GPU seconds count against the visitor's Hugging Face quota, so log in on huggingface.co for more; one request at a time, up to 8 questions, 255 options and 6 000 characters of state per request. The weights are public: ait-hf/certus-jev-like-v0007.
Adapters for your own use case
Every domain here is a small LoRA adapter (8–140 MB) trained on top of the trunk plus one router head. A new domain needs a few hundred labelled examples in the API's own shape (state, question, options, answer) and minutes of training on one consumer GPU — the spam adapter came from 600 emails — and it joins the router without retraining anything else. Details in the model card.
Known limits
A 1.5B single-pass reader: exact computation (prices × quantities, unit conversions) and multi-step logic traps are hit-and-miss even with the math and reasoning adapters. States are truncated to 700 tokens. English first; a multilingual adapter is planned for the next version.
Calling the API
The body is identical to TypeSafe Jev's POST /v1/systemone. Every question becomes one forward pass; all questions in a request are batched together.
Endpoints: POST /v1/systemone · GET /v1/models · GET /v1/models/{id} · GET /health. Use "model": "auto" for the router, "v0007" for the trunk alone or one adapter such as "v0007-math".