Instructions to use abhishek085/jev-control-core with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use abhishek085/jev-control-core with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf abhishek085/jev-control-core:F16 # Run inference directly in the terminal: llama cli -hf abhishek085/jev-control-core:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf abhishek085/jev-control-core:F16 # Run inference directly in the terminal: llama cli -hf abhishek085/jev-control-core:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf abhishek085/jev-control-core:F16 # Run inference directly in the terminal: ./llama-cli -hf abhishek085/jev-control-core:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf abhishek085/jev-control-core:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf abhishek085/jev-control-core:F16
Use Docker
docker model run hf.co/abhishek085/jev-control-core:F16
- LM Studio
- Jan
- Ollama
How to use abhishek085/jev-control-core with Ollama:
ollama run hf.co/abhishek085/jev-control-core:F16
- Unsloth Desktop
- Pi
How to use abhishek085/jev-control-core with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf abhishek085/jev-control-core:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "abhishek085/jev-control-core:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use abhishek085/jev-control-core with Docker Model Runner:
docker model run hf.co/abhishek085/jev-control-core:F16
- Lemonade
How to use abhishek085/jev-control-core with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull abhishek085/jev-control-core:F16
Run and chat with the model
lemonade run user.jev-control-core-F16
List all available models
lemonade list
- Hermes Agent
How to use abhishek085/jev-control-core with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf abhishek085/jev-control-core:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default abhishek085/jev-control-core:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use abhishek085/jev-control-core with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf abhishek085/jev-control-core:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "abhishek085/jev-control-core:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf abhishek085/jev-control-core:F16# Run inference directly in the terminal:
llama cli -hf abhishek085/jev-control-core:F16Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf abhishek085/jev-control-core:F16# Run inference directly in the terminal:
./llama-cli -hf abhishek085/jev-control-core:F16Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf abhishek085/jev-control-core:F16# Run inference directly in the terminal:
./build/bin/llama-cli -hf abhishek085/jev-control-core:F16Use Docker
docker model run hf.co/abhishek085/jev-control-core:F16jev-control-core: a small, fast decision model for agent decision sites
A single-forward-pass typed-decision model (open-spark-Jev's "System One" class, the same family as spark-s1), purpose-built for a narrower target than spark-s1's JevBench-style general decision benchmark: the decision sites inside a real agent harness — guardrail/injection gates, tool routing, context ranking, answer sufficiency, triage, moderation, claim verification, next-action selection, entity matching, escalate-to-human. These are the ten patterns JevControl's own in-app Guide lists as "where a decision model tends to fit" — short state, 2-5 options, embedded in a larger agent loop, not an adversarial benchmark question.
Base model: Qwen/Qwen3.5-0.8B-Base (24 layers, hybrid
gated-delta-net linear attention + full attention, hidden size 1,024). Qwen ships this size only as a
vision-language checkpoint (Qwen3_5ForConditionalGeneration); this repository is the extracted
text-only decoder (Qwen3_5ForCausalLM, 0.75B), full-parameter fine-tuned -- 0.8B is small enough and
the target distribution narrow enough that a full tune is cheap and reaches higher accuracy than a
rank-16 LoRA adapter would on this task.
Two sizes, one family: jev-control-core (this repository, 0.8B dense) and
jev-control-es (149M, ModernBERT encoder,
Laya-style option-marker readout) -- pick by latency budget.
How it works
Same readout as spark-s1: the state and a typed question (Choice/Score/Noul, each with an explicit
option list) are rendered into one prompt; the answer-slot letter logits are read from a single forward
pass, restricted to the valid option letters, and softmaxed. No decoding, no parsing, no output outside
the options you defined. Served through the same /v1/systemone gateway as spark-s1
(open_spark_jev.serve.gateway), so it's a drop-in smaller sibling, not a different serving stack.
Training
Fully fine-tuned (all 0.75B parameters, no LoRA) for 3 epochs on 20,000 rows from
os_datagen.control's ten decision-site families (2,000 rows/family) -- code-generated, code-verified
gold, zero overlap with any benchmark (scripts/tools/benchmark_overlap.py, 0/26,000 rows flagged).
lr 3e-5, batch 16 x grad-accum 2, cosine schedule, lambda_brier 0.5 (same KL + Brier-regularised loss
as spark-s1's sft.py). Every choice-type family's option list is rendered in a per-row-random order,
and every family draws from substantially widened prompt/topic pools rather than a handful of fixed
templates (see History below for why both of these matter more than they sound).
Evaluation
Own held-out splits (os_datagen.control's test_locked/challenge, same families as training,
different generated instances, scored both at the option order the row was written in and averaged over
3 random re-permutations of that order):
| split | accuracy | accuracy (mean over 3 option-order permutations) | ECE (calibrated) |
|---|---|---|---|
| test_locked | 1.000 | 1.000 | 0.000 |
| challenge | 1.000 | 1.000 | 0.000 |
Latency, isolated single-decision calls, in-process HF backend, one NVIDIA GB10: 18.5 ms per decision (median). Faster than decider-4b (32-35 ms) and well under spark-s1-4b-v6 (75 ms).
Real harness (JevControl, support_desk demo, 203 tasks, 4 decision sites/task, against a Gemma-4-E4B
baseline that decides by prompting):
| arm | accuracy | p50 latency | verdict |
|---|---|---|---|
| Gemma 4 (prompted) | 0.892 | 2,933 ms | baseline |
| jev-control-core | 0.837 | 4,563 ms* | close (Δ -0.054) |
Per decision site: injection 0.911, route 0.988, sufficiency 0.915, relevance 0.770.
*p50 wall-clock for the harness run as a whole, not model latency alone -- this run's decider shared the box with other GPU work; the isolated 18.5ms/decision figure above is the honest per-call latency number.
Why the real-harness number moved from 0.438 to 0.837 (three fixes, kept here on purpose)
The first trained version of this model scored 0.438 on the real harness despite 0.956 in-distribution accuracy -- a large, genuine transfer gap. Three separate, diagnosed issues accounted for essentially all of it, and all three are worth knowing if you retrain this family on your own decision sites:
- Templated, fixed-vocabulary synthetic state. The
route/sufficient/relevancefamilies originally rendered abstract, symbolic state (e.g."Retrieved so far: the price: known.") instead of realistic customer-message and article prose. Rewriting them to use varied greetings/sign-offs and real-looking KB article text over a 12-topic corpus took the real-harness number to 0.635. - Fixed option order. Four choice-type families (
tool_routing,moderation_class,next_action,escalate_human) always rendered their options in the same order every training row. The model could solve every training example by learning "the answer is at position N" without ever reading the option text -- invisible on our own eval (which never varies the order) but fatal the moment a real caller enumerates its own options in its own order. Theroutesite's real-harness confusion matrix showed the unmistakable signature: a clean positional shift (kb->orders94/203,orders->human48/203) rather than random noise. Randomizing option order per training row (_shuffled()indatagen-pipeline/src/os_datagen/control/families.py) fixedroutespecifically from 0.152 to 0.978 and took the overall real-harness number to 0.813. - Narrow templates within each family. Even with (1) and (2) fixed, training loss collapsed to 0.0000
within the first ~100 of ~560 steps -- the 5,000-row dataset was diverse enough to fix the two bugs above
but still narrow enough (e.g. one family's negative case was drawn from just 8 topics) to overfit near-
instantly rather than learn a robust rule. Widening every family's template/topic/signal pools and scaling
to 20,000 rows (2,000/family) pushed the loss curve out past the first epoch and took this model's
real-harness number to 0.837. (The same fix made the smaller
jev-control-essibling'sinjectionsite worse, not better -- see that model's card for why more data alone isn't a universal fix.)
Usage
As an API (recommended -- restricted-option decoding, calibration and abstain-threshold logic all live server-side). Run the OpenAI-compatible gateway this repository ships with:
python -m open_spark_jev.serve.gateway --backend hf --default-model jev-control-core --port 8620
curl -s http://localhost:8620/v1/systemone -H "Content-Type: application/json" -d '{
"model": "jev-control-core",
"state": "Customer: my order #A1006 never arrived.",
"question": {"type": "choice", "prompt": "route to:",
"options": ["kb", "orders", "human"]}
}'
Direct, CUDA or CPU (transformers):
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("abhishek085/jev-control-core")
model = AutoModelForCausalLM.from_pretrained("abhishek085/jev-control-core", torch_dtype="bfloat16", device_map="cuda")
Restricted-option decoding (reading only the answer-slot letter logits, softmaxed over the valid options)
is what actually makes this a calibrated decision model rather than a free-generation chat model -- that
logic lives in open_spark_jev.serve.gateway / open_spark_jev/eval/osdg.py, not in transformers alone,
so prefer the gateway path above unless you're reimplementing that readout yourself.
Apple Silicon (verified path -- GGUF + llama.cpp Metal): this repo includes an f16 GGUF build
(jev-control-core-f16.gguf, see Files below), verified to give byte-identical top-logit ordering and
confidence to the safetensors checkpoint on a real decision prompt. This is the Metal path to use today.
MLX: not currently supported. mlx-lm does not ship a model class for Qwen3.5's hybrid
gated-delta-net + full-attention architecture as of this writing, and this checkpoint was built and
tested on Linux/CUDA hardware with no Apple Silicon available to verify an MLX conversion -- so rather
than claim untested support, use the GGUF/llama.cpp Metal path above on Apple Silicon.
JevControl
This model is purpose-built for the decision sites JevControl's
own in-app Guide documents -- JevControl is the tool used to produce the real-harness numbers on this card
(support_desk demo, 203 tasks) and is the recommended way to measure this model's savings against your
own agent harness before adopting it.
Limitations
- Not evaluated on JevBench: this model is intentionally scoped to JevControl-shaped decision sites, not general benchmark decisions -- use spark-s1 for that.
- No reinforcement-learning stage.
- The remaining 0.079 real-harness gap to the Gemma-4 baseline is concentrated in
sufficiency(0.815) andrelevance(0.801) -- both require judging whether a short article's prose actually contains an answer, the hardest reading-comprehension step of the four sites, and the most likely to still carry some synthetic-corpus vocabulary bias even after the fixes above. - Per-question-type temperatures were fitted on the synthetic calibration split only; refit on your own
labelled traffic before trusting confidence-gated escalation (
python -m open_spark_jev.eval.osdgor JevControl's ownpython -m decider.calibrate-equivalent path).
Files
config.json, tokenizer.json/tokenizer_config.json, model.safetensors (bf16), calibration.json
(per-type temperatures), chat_template.jinja, generation_config.json. A GGUF build
(jev-control-core-f16.gguf, f16, no quantisation) is included for llama.cpp / Metal serving; verified to
give byte-identical top-logit ordering and confidence to the safetensors checkpoint on a real decision
prompt. Convert with --no-mtp if rebuilding from the HF checkpoint (this size of Qwen3.5 was extracted
from a vision-language release and no longer carries the speculative-decoding head llama.cpp's converter
otherwise expects).
Reproduction
Data generation, training config and the extraction/GGUF scripts are in
abhishek085/open-spark-jev:
datagen-pipeline/src/os_datagen/control/, configs/train/sft_control_jev.yaml,
scripts/tools/extract_qwen35_text.py.
- Downloads last month
- 205
Model tree for abhishek085/jev-control-core
Base model
Qwen/Qwen3.5-0.8B-Base
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf abhishek085/jev-control-core:F16# Run inference directly in the terminal: llama cli -hf abhishek085/jev-control-core:F16