SmolLM3-3B (APO checkpoint), readout fine-tuned (LoRA, supervised)

A System One decision model: it reads a state, answers typed questions (choice, score, noul) and returns calibrated probability distributions your code can branch on β€” it never writes text. This repo is a rank-16 LoRA (30,228,480 parameters) on HuggingFaceTB/SmolLM3-3B-checkpoints, merged into the weights at load, trained on its own decision readout.

At a glance β€” accuracy 0.705 Β· ECE 0.058 Β· held-out 0.741 Β· TVD to human labels 0.337 Β· sure loss 0.292
same order, SmolLM3-3B APO checkpoint, untuned (Tier 0): 0.539 / 0.117 / 0.579 / 0.460 / 0.226
same order, Jev 1.13.0: 0.733 / 0.113 / 0.835 / 0.432 / 0.081

jev-bench Β· leaderboard Β· findings Β· code

Use it

from jevify import load_jevified

model = load_jevified("Praveenrajus/jevify-smollm3-3b-apo-readout")
model.ask({"text": "The battery lasted two days on a single charge."},
          {"q": {"type": "noul", "instructions": "Is the review positive?"}})

jevify-serve --model Praveenrajus/jevify-smollm3-3b-apo-readout serves it as a drop-in for the TypeSafe SDK (TYPESAFE_BASE_URL=http://localhost:8000). The backbone is pulled from its own repo at load, pinned to commit cfb32d505f5025ec9be4e704f70cfbf5bdf8da94.

Results

Every number is on the jev-bench test splits (22,773 records) or the study's other test suites, scored the same way for every model; the rows under this model are references from the same study.

Decisions and calibration

model acc ECE Brier held-out acc TVD to human labels
this model 0.705 0.058 0.363 0.741 0.337
SmolLM3-3B APO checkpoint, untuned (Tier 0) 0.539 0.117 0.526 0.579 0.460
same recipe + coherence 0.708 0.055 0.360 0.749 0.318
same recipe from the SFT checkpoint 0.705 0.057 0.365 0.738 0.350
Jev 1.13.0 (TypeSafe API) 0.733 0.113 0.349 0.835 0.432

Coherence and invariance β€” sure loss: mean dΒ² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers.

model sure loss share incoherent order flip tag TVD K=2β†’max acc drop
this model 0.292 0.964 β€” β€” β€”
SmolLM3-3B APO checkpoint, untuned (Tier 0) 0.226 0.990 0.435 0.077 0.566
same recipe + coherence 0.041 0.640 β€” β€” β€”
same recipe from the SFT checkpoint 0.315 0.966 β€” β€” β€”
Jev 1.13.0 (TypeSafe API) 0.081 0.725 0.046 β€” 0.246

Out of distribution β€” stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks.

model stated rule 'none' when gone hijack phishing AUROC tool risk
this model β€” β€” β€” β€” β€”
SmolLM3-3B APO checkpoint, untuned (Tier 0) 0.588 0.552 0.464 0.794 0.800
same recipe + coherence β€” β€” β€” β€” β€”
same recipe from the SFT checkpoint β€” β€” β€” β€” β€”
Jev 1.13.0 (TypeSafe API) 0.924 0.744 0.205 0.688 0.933

Reproduction check. Loading this folder with load_jevified and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Ξ”p| 0.004, max 0.030 (the adapter is merged into bf16 weights at load).

How it was trained

The model is trained on its own decision readout β€” the distribution over the allowed answers read at the answer position, one forward pass, no decoding β€” with the primitive's proper scoring rule. Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources (5,885 families, at most 400 records per source); lr 3e-05, 2 epochs, best epoch by validation loss (epoch 1), seed 0. A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits. The six held-out sources (clinc150, arc_challenge, yelp5, measuring_hate_speech, fever_evidence, strategyqa_grounded) never appeared in training.

Files

  • jevify_config.json β€” the recipe, the backbone and the training settings load_jevified reads
  • lora/ β€” the adapter, merged into the backbone at load
  • results/test_metrics.json β€” every jev-bench config; recipe.json β€” the fitted recipe
  • results/coherence.json, probes.json, tags.json β€” the coherence, probe and tag tests
  • results/train.json β€” the training log; summary.json β€” this model's row of the study table
  • results/verification.json β€” the reproduction check reported under Results

Related models

Limitations

  • One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions.
  • The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it.
  • English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Praveenrajus/jevify-smollm3-3b-apo-readout

Finetuned
(5)
this model

Dataset used to train Praveenrajus/jevify-smollm3-3b-apo-readout