jev-control-core: a small, fast decision model for agent decision sites

A single-forward-pass typed-decision model (open-spark-Jev's "System One" class, the same family as spark-s1), purpose-built for a narrower target than spark-s1's JevBench-style general decision benchmark: the decision sites inside a real agent harness — guardrail/injection gates, tool routing, context ranking, answer sufficiency, triage, moderation, claim verification, next-action selection, entity matching, escalate-to-human. These are the ten patterns JevControl's own in-app Guide lists as "where a decision model tends to fit" — short state, 2-5 options, embedded in a larger agent loop, not an adversarial benchmark question.

Base model: Qwen/Qwen3.5-0.8B-Base (24 layers, hybrid gated-delta-net linear attention + full attention, hidden size 1,024). Qwen ships this size only as a vision-language checkpoint (Qwen3_5ForConditionalGeneration); this repository is the extracted text-only decoder (Qwen3_5ForCausalLM, 0.75B), full-parameter fine-tuned -- 0.8B is small enough and the target distribution narrow enough that a full tune is cheap and reaches higher accuracy than a rank-16 LoRA adapter would on this task.

Two sizes, one family: jev-control-core (this repository, 0.8B dense) and jev-control-es (149M, ModernBERT encoder, Laya-style option-marker readout) -- pick by latency budget.

How it works

Same readout as spark-s1: the state and a typed question (Choice/Score/Noul, each with an explicit option list) are rendered into one prompt; the answer-slot letter logits are read from a single forward pass, restricted to the valid option letters, and softmaxed. No decoding, no parsing, no output outside the options you defined. Served through the same /v1/systemone gateway as spark-s1 (open_spark_jev.serve.gateway), so it's a drop-in smaller sibling, not a different serving stack.

Training

Fully fine-tuned (all 0.75B parameters, no LoRA) for 3 epochs on 20,000 rows from os_datagen.control's ten decision-site families (2,000 rows/family) -- code-generated, code-verified gold, zero overlap with any benchmark (scripts/tools/benchmark_overlap.py, 0/26,000 rows flagged). lr 3e-5, batch 16 x grad-accum 2, cosine schedule, lambda_brier 0.5 (same KL + Brier-regularised loss as spark-s1's sft.py). Every choice-type family's option list is rendered in a per-row-random order, and every family draws from substantially widened prompt/topic pools rather than a handful of fixed templates (see History below for why both of these matter more than they sound).

Evaluation

Own held-out splits (os_datagen.control's test_locked/challenge, same families as training, different generated instances, scored both at the option order the row was written in and averaged over 3 random re-permutations of that order):

split accuracy accuracy (mean over 3 option-order permutations) ECE (calibrated)
test_locked 1.000 1.000 0.000
challenge 1.000 1.000 0.000

Latency, isolated single-decision calls, in-process HF backend, one NVIDIA GB10: 18.5 ms per decision (median). Faster than decider-4b (32-35 ms) and well under spark-s1-4b-v6 (75 ms).

Real harness (JevControl, support_desk demo, 203 tasks, 4 decision sites/task, against a Gemma-4-E4B baseline that decides by prompting):

arm accuracy p50 latency verdict
Gemma 4 (prompted) 0.892 2,933 ms baseline
jev-control-core 0.837 4,563 ms* close (Δ -0.054)

Per decision site: injection 0.911, route 0.988, sufficiency 0.915, relevance 0.770.

*p50 wall-clock for the harness run as a whole, not model latency alone -- this run's decider shared the box with other GPU work; the isolated 18.5ms/decision figure above is the honest per-call latency number.

Why the real-harness number moved from 0.438 to 0.837 (three fixes, kept here on purpose)

The first trained version of this model scored 0.438 on the real harness despite 0.956 in-distribution accuracy -- a large, genuine transfer gap. Three separate, diagnosed issues accounted for essentially all of it, and all three are worth knowing if you retrain this family on your own decision sites:

  1. Templated, fixed-vocabulary synthetic state. The route/sufficient/relevance families originally rendered abstract, symbolic state (e.g. "Retrieved so far: the price: known.") instead of realistic customer-message and article prose. Rewriting them to use varied greetings/sign-offs and real-looking KB article text over a 12-topic corpus took the real-harness number to 0.635.
  2. Fixed option order. Four choice-type families (tool_routing, moderation_class, next_action, escalate_human) always rendered their options in the same order every training row. The model could solve every training example by learning "the answer is at position N" without ever reading the option text -- invisible on our own eval (which never varies the order) but fatal the moment a real caller enumerates its own options in its own order. The route site's real-harness confusion matrix showed the unmistakable signature: a clean positional shift (kb->orders 94/203, orders->human 48/203) rather than random noise. Randomizing option order per training row (_shuffled() in datagen-pipeline/src/os_datagen/control/families.py) fixed route specifically from 0.152 to 0.978 and took the overall real-harness number to 0.813.
  3. Narrow templates within each family. Even with (1) and (2) fixed, training loss collapsed to 0.0000 within the first ~100 of ~560 steps -- the 5,000-row dataset was diverse enough to fix the two bugs above but still narrow enough (e.g. one family's negative case was drawn from just 8 topics) to overfit near- instantly rather than learn a robust rule. Widening every family's template/topic/signal pools and scaling to 20,000 rows (2,000/family) pushed the loss curve out past the first epoch and took this model's real-harness number to 0.837. (The same fix made the smaller jev-control-es sibling's injection site worse, not better -- see that model's card for why more data alone isn't a universal fix.)

Usage

As an API (recommended -- restricted-option decoding, calibration and abstain-threshold logic all live server-side). Run the OpenAI-compatible gateway this repository ships with:

python -m open_spark_jev.serve.gateway --backend hf --default-model jev-control-core --port 8620
curl -s http://localhost:8620/v1/systemone -H "Content-Type: application/json" -d '{
  "model": "jev-control-core",
  "state": "Customer: my order #A1006 never arrived.",
  "question": {"type": "choice", "prompt": "route to:",
               "options": ["kb", "orders", "human"]}
}'

Direct, CUDA or CPU (transformers):

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("abhishek085/jev-control-core")
model = AutoModelForCausalLM.from_pretrained("abhishek085/jev-control-core", torch_dtype="bfloat16", device_map="cuda")

Restricted-option decoding (reading only the answer-slot letter logits, softmaxed over the valid options) is what actually makes this a calibrated decision model rather than a free-generation chat model -- that logic lives in open_spark_jev.serve.gateway / open_spark_jev/eval/osdg.py, not in transformers alone, so prefer the gateway path above unless you're reimplementing that readout yourself.

Apple Silicon (verified path -- GGUF + llama.cpp Metal): this repo includes an f16 GGUF build (jev-control-core-f16.gguf, see Files below), verified to give byte-identical top-logit ordering and confidence to the safetensors checkpoint on a real decision prompt. This is the Metal path to use today.

MLX: not currently supported. mlx-lm does not ship a model class for Qwen3.5's hybrid gated-delta-net + full-attention architecture as of this writing, and this checkpoint was built and tested on Linux/CUDA hardware with no Apple Silicon available to verify an MLX conversion -- so rather than claim untested support, use the GGUF/llama.cpp Metal path above on Apple Silicon.

JevControl

This model is purpose-built for the decision sites JevControl's own in-app Guide documents -- JevControl is the tool used to produce the real-harness numbers on this card (support_desk demo, 203 tasks) and is the recommended way to measure this model's savings against your own agent harness before adopting it.

Limitations

  • Not evaluated on JevBench: this model is intentionally scoped to JevControl-shaped decision sites, not general benchmark decisions -- use spark-s1 for that.
  • No reinforcement-learning stage.
  • The remaining 0.079 real-harness gap to the Gemma-4 baseline is concentrated in sufficiency (0.815) and relevance (0.801) -- both require judging whether a short article's prose actually contains an answer, the hardest reading-comprehension step of the four sites, and the most likely to still carry some synthetic-corpus vocabulary bias even after the fixes above.
  • Per-question-type temperatures were fitted on the synthetic calibration split only; refit on your own labelled traffic before trusting confidence-gated escalation (python -m open_spark_jev.eval.osdg or JevControl's own python -m decider.calibrate-equivalent path).

Files

config.json, tokenizer.json/tokenizer_config.json, model.safetensors (bf16), calibration.json (per-type temperatures), chat_template.jinja, generation_config.json. A GGUF build (jev-control-core-f16.gguf, f16, no quantisation) is included for llama.cpp / Metal serving; verified to give byte-identical top-logit ordering and confidence to the safetensors checkpoint on a real decision prompt. Convert with --no-mtp if rebuilding from the HF checkpoint (this size of Qwen3.5 was extracted from a vision-language release and no longer carries the speculative-decoding head llama.cpp's converter otherwise expects).

Reproduction

Data generation, training config and the extraction/GGUF scripts are in abhishek085/open-spark-jev: datagen-pipeline/src/os_datagen/control/, configs/train/sft_control_jev.yaml, scripts/tools/extract_qwen35_text.py.

Downloads last month
149
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abhishek085/jev-control-core

Quantized
(46)
this model