card: the decider family table, contents, changelog
Browse files
README.md
CHANGED
|
@@ -8,12 +8,31 @@ tags: [decision-model, calibrated, structured-output, multi-task, system-one, on
|
|
| 8 |
|
| 9 |
# decider-0.8b: typed decisions with calibrated probabilities in one forward pass
|
| 10 |
|
| 11 |
-
The
|
| 12 |
a **state** (a string or any JSON value) and a set of **typed questions** and returns a probability distribution for every question
|
| 13 |
from a single forward pass: Choice (2-255 options, optionally described), Score (2-10 described levels), Noul (probability of yes).
|
| 14 |
No decoding, no parsing, no output outside the options you defined. Same code, same wire format (`POST /v1/systemone`, TypeSafe
|
| 15 |
Jev's format), same training recipe as the 2B: one epoch of `scripts/train.sh full` from `Qwen/Qwen3.5-0.8B-Base` over the full
|
| 16 |
-
mixture (1.47M examples, 455M tokens, 4.5 h on one GH200).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
```python
|
| 19 |
# pip install git+https://github.com/Mapika/decider
|
|
@@ -27,6 +46,9 @@ d.system_one(
|
|
| 27 |
"frustration": {"type": "score", "instructions": "How frustrated is the customer?", "criteria": ["calm", "frustrated", "very frustrated"]}})
|
| 28 |
```
|
| 29 |
|
|
|
|
|
|
|
|
|
|
| 30 |
## How it compares with the 2B
|
| 31 |
|
| 32 |
Same 93 public tasks, same protocol, one temperature fitted on in-task data (it came out at 1.03 for both: the recipe calibrates
|
|
@@ -57,3 +79,12 @@ Those of decider-2b, more so: a small model without reasoning; English only; rul
|
|
| 57 |
otherwise skip") are not followed reliably, so state the decision as a plain question with described options; knowledge-heavy
|
| 58 |
multiple choice is close to the base model; calibration is measured on public datasets and teacher-labelled probes, not on your
|
| 59 |
traffic. The teacher-written training data comes from Qwen3.5-27B and carries its biases.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|
| 9 |
# decider-0.8b: typed decisions with calibrated probabilities in one forward pass
|
| 10 |
|
| 11 |
+
The smallest decider: a language model that does not generate text. It reads
|
| 12 |
a **state** (a string or any JSON value) and a set of **typed questions** and returns a probability distribution for every question
|
| 13 |
from a single forward pass: Choice (2-255 options, optionally described), Score (2-10 described levels), Noul (probability of yes).
|
| 14 |
No decoding, no parsing, no output outside the options you defined. Same code, same wire format (`POST /v1/systemone`, TypeSafe
|
| 15 |
Jev's format), same training recipe as the 2B: one epoch of `scripts/train.sh full` from `Qwen/Qwen3.5-0.8B-Base` over the full
|
| 16 |
+
mixture (1.47M examples, 455M tokens, 4.5 h on one GH200).
|
| 17 |
+
|
| 18 |
+
**Contents:** [The decider family](#the-decider-family) 路 [Usage](#usage) 路 [How it compares with the 2B](#how-it-compares-with-the-2b) 路 [Limitations](#limitations) 路 [Changelog](#changelog)
|
| 19 |
+
|
| 20 |
+
## The decider family
|
| 21 |
+
|
| 22 |
+
All five repositories share one interface (`decider.infer.Decider`, `POST /v1/systemone` in TypeSafe's format) and one
|
| 23 |
+
readout: the letter logits at an answer slot, softmaxed over the options. Pick by size and input.
|
| 24 |
+
|
| 25 |
+
| model | base | weights | use it for | numbers |
|
| 26 |
+
|---|---|---|---|---|
|
| 27 |
+
| [decider-2b](https://huggingface.co/Mapika/decider-2b) v10 | Qwen3.5-2B-Base | 3.5 GB bf16 | the default: routing, classification, judgments, browser agents; 4 ms per request with CUDA graphs on one GPU | regression set 0.805 in-task / 0.755 held-out; live browser 93%; Bespoke suite 0.704 |
|
| 28 |
+
| [decider-35b-a3b](https://huggingface.co/Mapika/decider-35b-a3b) v1 | Qwen3.5-35B-A3B-Base (3B active) | 65 GB bf16 | when accuracy is worth 3 to 4 times the cost per decision: knowledge and multi-step questions, long policies | 0.855 / 0.810, above the 2B on 93 of 95 tasks; JevBench hard 0.676; Bespoke 0.774; no RL stage |
|
| 29 |
+
| [decider-35b-a3b-nvfp4](https://huggingface.co/Mapika/decider-35b-a3b-nvfp4) | the 35B in NVFP4 | 19.6 GB | the 35B on Blackwell through vLLM or TensorRT-LLM | 1.0 to 1.5 points under bf16 on the measured fixtures |
|
| 30 |
+
| [decider-0.8b](https://huggingface.co/Mapika/decider-0.8b) | Qwen3.5-0.8B-Base | 1.4 GB bf16 | the smallest: routing, yes/no and short-state lookups within 1 to 4 points of the 2B, 1.5x faster | 0.776 / 0.707 on the single-run protocol (2B: 0.809 / 0.739) |
|
| 31 |
+
| [decider-2b-vision](https://huggingface.co/Mapika/decider-2b-vision) | Qwen3.5-2B vision-language, v5 text weights | 4.1 GB bf16 | decisions from an image plus a question; game frames | Visual7W 0.89; Breakout 41 from pixels |
|
| 32 |
+
|
| 33 |
+
Code, data registry, training scripts, the changelog and the per-version history: https://github.com/Mapika/decider.
|
| 34 |
+
|
| 35 |
+
## Usage
|
| 36 |
|
| 37 |
```python
|
| 38 |
# pip install git+https://github.com/Mapika/decider
|
|
|
|
| 46 |
"frustration": {"type": "score", "instructions": "How frustrated is the customer?", "criteria": ["calm", "frustrated", "very frustrated"]}})
|
| 47 |
```
|
| 48 |
|
| 49 |
+
`decide_batch`, `schema` (the cached question set), `decider.serve` and the TypeSafe request shape work as in the decider-2b
|
| 50 |
+
card; `decider/` in this repository is the inference subset of the GitHub package.
|
| 51 |
+
|
| 52 |
## How it compares with the 2B
|
| 53 |
|
| 54 |
Same 93 public tasks, same protocol, one temperature fitted on in-task data (it came out at 1.03 for both: the recipe calibrates
|
|
|
|
| 79 |
otherwise skip") are not followed reliably, so state the decision as a plain question with described options; knowledge-heavy
|
| 80 |
multiple choice is close to the base model; calibration is measured on public datasets and teacher-labelled probes, not on your
|
| 81 |
traffic. The teacher-written training data comes from Qwen3.5-27B and carries its biases.
|
| 82 |
+
|
| 83 |
+
## Changelog
|
| 84 |
+
|
| 85 |
+
| version | what changed |
|
| 86 |
+
|---|---|
|
| 87 |
+
| **v1** (these weights) | one epoch of `scripts/train.sh full` from Qwen3.5-0.8B-Base, the single-run form of the supervised recipe that produced decider-2b v8 to v9 |
|
| 88 |
+
|
| 89 |
+
Every decider release is listed in [docs/CHANGELOG.md](https://github.com/Mapika/decider/blob/main/docs/CHANGELOG.md) of the
|
| 90 |
+
GitHub repository. Code and the training recipe: https://github.com/Mapika/decider.
|