File size: 10,871 Bytes
e033ec6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
---
license: apache-2.0
base_model: Qwen/Qwen3.5-0.8B-Base
language: [en]
pipeline_tag: text-classification
tags: [decision-model, calibrated, structured-output, system-one, one-pass, jevcontrol]
---

# jev-control-core: a small, fast decision model for agent decision sites

A single-forward-pass typed-decision model (open-spark-Jev's "System One" class, the same family as
[spark-s1](https://huggingface.co/abhishek085/spark-s1-4b-v6)), purpose-built for a narrower target than
spark-s1's JevBench-style general decision benchmark: the decision *sites* inside a real agent
harness — guardrail/injection gates, tool routing, context ranking, answer sufficiency, triage,
moderation, claim verification, next-action selection, entity matching, escalate-to-human. These are
the ten patterns [JevControl](https://github.com/abhishek085/JevControl)'s own in-app Guide lists as
"where a decision model tends to fit" — short state, 2-5 options, embedded in a larger agent loop, not
an adversarial benchmark question.

Base model: [Qwen/Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base) (24 layers, hybrid
gated-delta-net linear attention + full attention, hidden size 1,024). Qwen ships this size only as a
vision-language checkpoint (`Qwen3_5ForConditionalGeneration`); this repository is the extracted
text-only decoder (`Qwen3_5ForCausalLM`, 0.75B), full-parameter fine-tuned -- 0.8B is small enough and
the target distribution narrow enough that a full tune is cheap and reaches higher accuracy than a
rank-16 LoRA adapter would on this task.

**Two sizes, one family:** `jev-control-core` (this repository, 0.8B dense) and
[`jev-control-es`](https://huggingface.co/abhishek085/jev-control-es) (149M, ModernBERT encoder,
Laya-style option-marker readout) -- pick by latency budget.

## How it works

Same readout as spark-s1: the state and a typed question (Choice/Score/Noul, each with an explicit
option list) are rendered into one prompt; the answer-slot letter logits are read from a single forward
pass, restricted to the valid option letters, and softmaxed. No decoding, no parsing, no output outside
the options you defined. Served through the same `/v1/systemone` gateway as spark-s1
(`open_spark_jev.serve.gateway`), so it's a drop-in smaller sibling, not a different serving stack.

## Training

Fully fine-tuned (all 0.75B parameters, no LoRA) for 3 epochs on 20,000 rows from
`os_datagen.control`'s ten decision-site families (2,000 rows/family) -- code-generated, code-verified
gold, zero overlap with any benchmark (`scripts/tools/benchmark_overlap.py`, 0/26,000 rows flagged).
lr 3e-5, batch 16 x grad-accum 2, cosine schedule, `lambda_brier` 0.5 (same KL + Brier-regularised loss
as spark-s1's `sft.py`). Every choice-type family's option list is rendered in a per-row-random order,
and every family draws from substantially widened prompt/topic pools rather than a handful of fixed
templates (see History below for why both of these matter more than they sound).

## Evaluation

**Own held-out splits** (`os_datagen.control`'s `test_locked`/`challenge`, same families as training,
different generated instances, scored both at the option order the row was written in and averaged over
3 random re-permutations of that order):

| split | accuracy | accuracy (mean over 3 option-order permutations) | ECE (calibrated) |
|---|---|---|---|
| test_locked | 1.000 | 1.000 | 0.000 |
| challenge | 1.000 | 1.000 | 0.000 |

**Latency**, isolated single-decision calls, in-process HF backend, one NVIDIA GB10: **18.5 ms** per
decision (median). Faster than decider-4b (32-35 ms) and well under spark-s1-4b-v6 (75 ms).

**Real harness (JevControl, `support_desk` demo, 203 tasks, 4 decision sites/task, against a Gemma-4-E4B
baseline that decides by prompting):**

| arm | accuracy | p50 latency | verdict |
|---|---|---|---|
| Gemma 4 (prompted) | 0.892 | 2,933 ms | baseline |
| **jev-control-core** | **0.837** | **4,563 ms*** | close (Δ -0.054) |

Per decision site: injection 0.911, route 0.988, sufficiency 0.915, relevance 0.770.

*\*p50 wall-clock for the harness run as a whole, not model latency alone -- this run's decider shared the
box with other GPU work; the isolated 18.5ms/decision figure above is the honest per-call latency number.*

## Why the real-harness number moved from 0.438 to 0.837 (three fixes, kept here on purpose)

The first trained version of this model scored **0.438** on the real harness despite 0.956 in-distribution
accuracy -- a large, genuine transfer gap. Three separate, diagnosed issues accounted for essentially all of
it, and all three are worth knowing if you retrain this family on your own decision sites:

1. **Templated, fixed-vocabulary synthetic state.** The `route`/`sufficient`/`relevance` families originally
   rendered abstract, symbolic state (e.g. `"Retrieved so far: the price: known."`) instead of realistic
   customer-message and article prose. Rewriting them to use varied greetings/sign-offs and real-looking KB
   article text over a 12-topic corpus took the real-harness number to 0.635.
2. **Fixed option order.** Four choice-type families (`tool_routing`, `moderation_class`, `next_action`,
   `escalate_human`) always rendered their options in the same order every training row. The model could
   solve every training example by learning "the answer is at position N" without ever reading the option
   text -- invisible on our own eval (which never varies the order) but fatal the moment a real caller
   enumerates its own options in its own order. The `route` site's real-harness confusion matrix showed the
   unmistakable signature: a clean positional shift (`kb`->`orders` 94/203, `orders`->`human` 48/203) rather
   than random noise. Randomizing option order per training row (`_shuffled()` in
   `datagen-pipeline/src/os_datagen/control/families.py`) fixed `route` specifically from 0.152 to 0.978 and
   took the overall real-harness number to 0.813.
3. **Narrow templates within each family.** Even with (1) and (2) fixed, training loss collapsed to 0.0000
   within the first ~100 of ~560 steps -- the 5,000-row dataset was diverse enough to fix the two bugs above
   but still narrow enough (e.g. one family's negative case was drawn from just 8 topics) to overfit near-
   instantly rather than learn a robust rule. Widening every family's template/topic/signal pools and scaling
   to 20,000 rows (2,000/family) pushed the loss curve out past the first epoch and took this model's
   real-harness number to 0.837. (The same fix made the smaller `jev-control-es` sibling's `injection` site
   *worse*, not better -- see that model's card for why more data alone isn't a universal fix.)

## Usage

**As an API (recommended -- restricted-option decoding, calibration and abstain-threshold logic all live
server-side).** Run the OpenAI-compatible gateway this repository ships with:

```bash
python -m open_spark_jev.serve.gateway --backend hf --default-model jev-control-core --port 8620
```

```bash
curl -s http://localhost:8620/v1/systemone -H "Content-Type: application/json" -d '{
  "model": "jev-control-core",
  "state": "Customer: my order #A1006 never arrived.",
  "question": {"type": "choice", "prompt": "route to:",
               "options": ["kb", "orders", "human"]}
}'
```

**Direct, CUDA or CPU (`transformers`):**

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("abhishek085/jev-control-core")
model = AutoModelForCausalLM.from_pretrained("abhishek085/jev-control-core", torch_dtype="bfloat16", device_map="cuda")
```

Restricted-option decoding (reading only the answer-slot letter logits, softmaxed over the valid options)
is what actually makes this a calibrated decision model rather than a free-generation chat model -- that
logic lives in `open_spark_jev.serve.gateway` / `open_spark_jev/eval/osdg.py`, not in `transformers` alone,
so prefer the gateway path above unless you're reimplementing that readout yourself.

**Apple Silicon (verified path -- GGUF + llama.cpp Metal):** this repo includes an f16 GGUF build
(`jev-control-core-f16.gguf`, see Files below), verified to give byte-identical top-logit ordering and
confidence to the safetensors checkpoint on a real decision prompt. This is the Metal path to use today.

**MLX:** not currently supported. `mlx-lm` does not ship a model class for Qwen3.5's hybrid
gated-delta-net + full-attention architecture as of this writing, and this checkpoint was built and
tested on Linux/CUDA hardware with no Apple Silicon available to verify an MLX conversion -- so rather
than claim untested support, use the GGUF/llama.cpp Metal path above on Apple Silicon.

## JevControl

This model is purpose-built for the decision sites [JevControl](https://github.com/abhishek085/JevControl)'s
own in-app Guide documents -- JevControl is the tool used to produce the real-harness numbers on this card
(`support_desk` demo, 203 tasks) and is the recommended way to measure this model's savings against your
own agent harness before adopting it.

## Limitations

* Not evaluated on JevBench: this model is intentionally scoped to JevControl-shaped decision sites, not
  general benchmark decisions -- use spark-s1 for that.
* No reinforcement-learning stage.
* The remaining 0.079 real-harness gap to the Gemma-4 baseline is concentrated in `sufficiency` (0.815) and
  `relevance` (0.801) -- both require judging whether a short article's prose actually contains an answer,
  the hardest reading-comprehension step of the four sites, and the most likely to still carry some
  synthetic-corpus vocabulary bias even after the fixes above.
* Per-question-type temperatures were fitted on the synthetic calibration split only; refit on your own
  labelled traffic before trusting confidence-gated escalation (`python -m open_spark_jev.eval.osdg` or
  JevControl's own `python -m decider.calibrate`-equivalent path).

## Files

`config.json`, `tokenizer.json`/`tokenizer_config.json`, `model.safetensors` (bf16), `calibration.json`
(per-type temperatures), `chat_template.jinja`, `generation_config.json`. A GGUF build
(`jev-control-core-f16.gguf`, f16, no quantisation) is included for llama.cpp / Metal serving; verified to
give byte-identical top-logit ordering and confidence to the safetensors checkpoint on a real decision
prompt. Convert with `--no-mtp` if rebuilding from the HF checkpoint (this size of Qwen3.5 was extracted
from a vision-language release and no longer carries the speculative-decoding head llama.cpp's converter
otherwise expects).

## Reproduction

Data generation, training config and the extraction/GGUF scripts are in
[abhishek085/open-spark-jev](https://github.com/abhishek085/open-spark-jev):
`datagen-pipeline/src/os_datagen/control/`, `configs/train/sft_control_jev.yaml`,
`scripts/tools/extract_qwen35_text.py`.