abhishek085 commited on
Commit
e033ec6
·
verified ·
1 Parent(s): b2bfd3a

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +178 -0
README.md ADDED
@@ -0,0 +1,178 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-0.8B-Base
4
+ language: [en]
5
+ pipeline_tag: text-classification
6
+ tags: [decision-model, calibrated, structured-output, system-one, one-pass, jevcontrol]
7
+ ---
8
+
9
+ # jev-control-core: a small, fast decision model for agent decision sites
10
+
11
+ A single-forward-pass typed-decision model (open-spark-Jev's "System One" class, the same family as
12
+ [spark-s1](https://huggingface.co/abhishek085/spark-s1-4b-v6)), purpose-built for a narrower target than
13
+ spark-s1's JevBench-style general decision benchmark: the decision *sites* inside a real agent
14
+ harness — guardrail/injection gates, tool routing, context ranking, answer sufficiency, triage,
15
+ moderation, claim verification, next-action selection, entity matching, escalate-to-human. These are
16
+ the ten patterns [JevControl](https://github.com/abhishek085/JevControl)'s own in-app Guide lists as
17
+ "where a decision model tends to fit" — short state, 2-5 options, embedded in a larger agent loop, not
18
+ an adversarial benchmark question.
19
+
20
+ Base model: [Qwen/Qwen3.5-0.8B-Base](https://huggingface.co/Qwen/Qwen3.5-0.8B-Base) (24 layers, hybrid
21
+ gated-delta-net linear attention + full attention, hidden size 1,024). Qwen ships this size only as a
22
+ vision-language checkpoint (`Qwen3_5ForConditionalGeneration`); this repository is the extracted
23
+ text-only decoder (`Qwen3_5ForCausalLM`, 0.75B), full-parameter fine-tuned -- 0.8B is small enough and
24
+ the target distribution narrow enough that a full tune is cheap and reaches higher accuracy than a
25
+ rank-16 LoRA adapter would on this task.
26
+
27
+ **Two sizes, one family:** `jev-control-core` (this repository, 0.8B dense) and
28
+ [`jev-control-es`](https://huggingface.co/abhishek085/jev-control-es) (149M, ModernBERT encoder,
29
+ Laya-style option-marker readout) -- pick by latency budget.
30
+
31
+ ## How it works
32
+
33
+ Same readout as spark-s1: the state and a typed question (Choice/Score/Noul, each with an explicit
34
+ option list) are rendered into one prompt; the answer-slot letter logits are read from a single forward
35
+ pass, restricted to the valid option letters, and softmaxed. No decoding, no parsing, no output outside
36
+ the options you defined. Served through the same `/v1/systemone` gateway as spark-s1
37
+ (`open_spark_jev.serve.gateway`), so it's a drop-in smaller sibling, not a different serving stack.
38
+
39
+ ## Training
40
+
41
+ Fully fine-tuned (all 0.75B parameters, no LoRA) for 3 epochs on 20,000 rows from
42
+ `os_datagen.control`'s ten decision-site families (2,000 rows/family) -- code-generated, code-verified
43
+ gold, zero overlap with any benchmark (`scripts/tools/benchmark_overlap.py`, 0/26,000 rows flagged).
44
+ lr 3e-5, batch 16 x grad-accum 2, cosine schedule, `lambda_brier` 0.5 (same KL + Brier-regularised loss
45
+ as spark-s1's `sft.py`). Every choice-type family's option list is rendered in a per-row-random order,
46
+ and every family draws from substantially widened prompt/topic pools rather than a handful of fixed
47
+ templates (see History below for why both of these matter more than they sound).
48
+
49
+ ## Evaluation
50
+
51
+ **Own held-out splits** (`os_datagen.control`'s `test_locked`/`challenge`, same families as training,
52
+ different generated instances, scored both at the option order the row was written in and averaged over
53
+ 3 random re-permutations of that order):
54
+
55
+ | split | accuracy | accuracy (mean over 3 option-order permutations) | ECE (calibrated) |
56
+ |---|---|---|---|
57
+ | test_locked | 1.000 | 1.000 | 0.000 |
58
+ | challenge | 1.000 | 1.000 | 0.000 |
59
+
60
+ **Latency**, isolated single-decision calls, in-process HF backend, one NVIDIA GB10: **18.5 ms** per
61
+ decision (median). Faster than decider-4b (32-35 ms) and well under spark-s1-4b-v6 (75 ms).
62
+
63
+ **Real harness (JevControl, `support_desk` demo, 203 tasks, 4 decision sites/task, against a Gemma-4-E4B
64
+ baseline that decides by prompting):**
65
+
66
+ | arm | accuracy | p50 latency | verdict |
67
+ |---|---|---|---|
68
+ | Gemma 4 (prompted) | 0.892 | 2,933 ms | baseline |
69
+ | **jev-control-core** | **0.837** | **4,563 ms*** | close (Δ -0.054) |
70
+
71
+ Per decision site: injection 0.911, route 0.988, sufficiency 0.915, relevance 0.770.
72
+
73
+ *\*p50 wall-clock for the harness run as a whole, not model latency alone -- this run's decider shared the
74
+ box with other GPU work; the isolated 18.5ms/decision figure above is the honest per-call latency number.*
75
+
76
+ ## Why the real-harness number moved from 0.438 to 0.837 (three fixes, kept here on purpose)
77
+
78
+ The first trained version of this model scored **0.438** on the real harness despite 0.956 in-distribution
79
+ accuracy -- a large, genuine transfer gap. Three separate, diagnosed issues accounted for essentially all of
80
+ it, and all three are worth knowing if you retrain this family on your own decision sites:
81
+
82
+ 1. **Templated, fixed-vocabulary synthetic state.** The `route`/`sufficient`/`relevance` families originally
83
+ rendered abstract, symbolic state (e.g. `"Retrieved so far: the price: known."`) instead of realistic
84
+ customer-message and article prose. Rewriting them to use varied greetings/sign-offs and real-looking KB
85
+ article text over a 12-topic corpus took the real-harness number to 0.635.
86
+ 2. **Fixed option order.** Four choice-type families (`tool_routing`, `moderation_class`, `next_action`,
87
+ `escalate_human`) always rendered their options in the same order every training row. The model could
88
+ solve every training example by learning "the answer is at position N" without ever reading the option
89
+ text -- invisible on our own eval (which never varies the order) but fatal the moment a real caller
90
+ enumerates its own options in its own order. The `route` site's real-harness confusion matrix showed the
91
+ unmistakable signature: a clean positional shift (`kb`->`orders` 94/203, `orders`->`human` 48/203) rather
92
+ than random noise. Randomizing option order per training row (`_shuffled()` in
93
+ `datagen-pipeline/src/os_datagen/control/families.py`) fixed `route` specifically from 0.152 to 0.978 and
94
+ took the overall real-harness number to 0.813.
95
+ 3. **Narrow templates within each family.** Even with (1) and (2) fixed, training loss collapsed to 0.0000
96
+ within the first ~100 of ~560 steps -- the 5,000-row dataset was diverse enough to fix the two bugs above
97
+ but still narrow enough (e.g. one family's negative case was drawn from just 8 topics) to overfit near-
98
+ instantly rather than learn a robust rule. Widening every family's template/topic/signal pools and scaling
99
+ to 20,000 rows (2,000/family) pushed the loss curve out past the first epoch and took this model's
100
+ real-harness number to 0.837. (The same fix made the smaller `jev-control-es` sibling's `injection` site
101
+ *worse*, not better -- see that model's card for why more data alone isn't a universal fix.)
102
+
103
+ ## Usage
104
+
105
+ **As an API (recommended -- restricted-option decoding, calibration and abstain-threshold logic all live
106
+ server-side).** Run the OpenAI-compatible gateway this repository ships with:
107
+
108
+ ```bash
109
+ python -m open_spark_jev.serve.gateway --backend hf --default-model jev-control-core --port 8620
110
+ ```
111
+
112
+ ```bash
113
+ curl -s http://localhost:8620/v1/systemone -H "Content-Type: application/json" -d '{
114
+ "model": "jev-control-core",
115
+ "state": "Customer: my order #A1006 never arrived.",
116
+ "question": {"type": "choice", "prompt": "route to:",
117
+ "options": ["kb", "orders", "human"]}
118
+ }'
119
+ ```
120
+
121
+ **Direct, CUDA or CPU (`transformers`):**
122
+
123
+ ```python
124
+ from transformers import AutoModelForCausalLM, AutoTokenizer
125
+ tok = AutoTokenizer.from_pretrained("abhishek085/jev-control-core")
126
+ model = AutoModelForCausalLM.from_pretrained("abhishek085/jev-control-core", torch_dtype="bfloat16", device_map="cuda")
127
+ ```
128
+
129
+ Restricted-option decoding (reading only the answer-slot letter logits, softmaxed over the valid options)
130
+ is what actually makes this a calibrated decision model rather than a free-generation chat model -- that
131
+ logic lives in `open_spark_jev.serve.gateway` / `open_spark_jev/eval/osdg.py`, not in `transformers` alone,
132
+ so prefer the gateway path above unless you're reimplementing that readout yourself.
133
+
134
+ **Apple Silicon (verified path -- GGUF + llama.cpp Metal):** this repo includes an f16 GGUF build
135
+ (`jev-control-core-f16.gguf`, see Files below), verified to give byte-identical top-logit ordering and
136
+ confidence to the safetensors checkpoint on a real decision prompt. This is the Metal path to use today.
137
+
138
+ **MLX:** not currently supported. `mlx-lm` does not ship a model class for Qwen3.5's hybrid
139
+ gated-delta-net + full-attention architecture as of this writing, and this checkpoint was built and
140
+ tested on Linux/CUDA hardware with no Apple Silicon available to verify an MLX conversion -- so rather
141
+ than claim untested support, use the GGUF/llama.cpp Metal path above on Apple Silicon.
142
+
143
+ ## JevControl
144
+
145
+ This model is purpose-built for the decision sites [JevControl](https://github.com/abhishek085/JevControl)'s
146
+ own in-app Guide documents -- JevControl is the tool used to produce the real-harness numbers on this card
147
+ (`support_desk` demo, 203 tasks) and is the recommended way to measure this model's savings against your
148
+ own agent harness before adopting it.
149
+
150
+ ## Limitations
151
+
152
+ * Not evaluated on JevBench: this model is intentionally scoped to JevControl-shaped decision sites, not
153
+ general benchmark decisions -- use spark-s1 for that.
154
+ * No reinforcement-learning stage.
155
+ * The remaining 0.079 real-harness gap to the Gemma-4 baseline is concentrated in `sufficiency` (0.815) and
156
+ `relevance` (0.801) -- both require judging whether a short article's prose actually contains an answer,
157
+ the hardest reading-comprehension step of the four sites, and the most likely to still carry some
158
+ synthetic-corpus vocabulary bias even after the fixes above.
159
+ * Per-question-type temperatures were fitted on the synthetic calibration split only; refit on your own
160
+ labelled traffic before trusting confidence-gated escalation (`python -m open_spark_jev.eval.osdg` or
161
+ JevControl's own `python -m decider.calibrate`-equivalent path).
162
+
163
+ ## Files
164
+
165
+ `config.json`, `tokenizer.json`/`tokenizer_config.json`, `model.safetensors` (bf16), `calibration.json`
166
+ (per-type temperatures), `chat_template.jinja`, `generation_config.json`. A GGUF build
167
+ (`jev-control-core-f16.gguf`, f16, no quantisation) is included for llama.cpp / Metal serving; verified to
168
+ give byte-identical top-logit ordering and confidence to the safetensors checkpoint on a real decision
169
+ prompt. Convert with `--no-mtp` if rebuilding from the HF checkpoint (this size of Qwen3.5 was extracted
170
+ from a vision-language release and no longer carries the speculative-decoding head llama.cpp's converter
171
+ otherwise expects).
172
+
173
+ ## Reproduction
174
+
175
+ Data generation, training config and the extraction/GGUF scripts are in
176
+ [abhishek085/open-spark-jev](https://github.com/abhishek085/open-spark-jev):
177
+ `datagen-pipeline/src/os_datagen/control/`, `configs/train/sft_control_jev.yaml`,
178
+ `scripts/tools/extract_qwen35_text.py`.