Llama-3.1-Tülu-3-8B-SFT, readout fine-tuned (LoRA, + coherence)

Built with Llama. Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

A System One decision model: it reads a state, answers typed questions (choice, score, noul) and returns calibrated probability distributions your code can branch on — it never writes text. This repo is a rank-16 LoRA (41,943,040 parameters) on allenai/Llama-3.1-Tulu-3-8B-SFT, merged into the weights at load, trained on its own decision readout with a coherence penalty.

At a glance — accuracy 0.722 · ECE 0.064 · held-out 0.758 · TVD to human labels 0.318 · sure loss 0.030
same order, Tülu-3-8B-SFT, untuned (Tier 0): 0.603 / 0.099 / 0.644 / 0.448 / 0.200
same order, Jev 1.13.0: 0.733 / 0.113 / 0.835 / 0.432 / 0.081

jev-bench · leaderboard · findings · code

Use it

from jevify import load_jevified

model = load_jevified("Praveenrajus/Llama-3.1-Tulu-3-8B-SFT-jevify-readout-coh")
model.ask({"text": "The battery lasted two days on a single charge."},
          {"q": {"type": "noul", "instructions": "Is the review positive?"}})

jevify-serve --model Praveenrajus/Llama-3.1-Tulu-3-8B-SFT-jevify-readout-coh serves it as a drop-in for the TypeSafe SDK (TYPESAFE_BASE_URL=http://localhost:8000). The backbone is pulled from its own repo at load, pinned to commit f2a0b46b0cfda21003c6141b1ff837b7e165524d.

Results

Every number is on the jev-bench test splits (22,773 records) or the study's other test suites, scored the same way for every model; the rows under this model are references from the same study.

Decisions and calibration

model acc ECE Brier held-out acc TVD to human labels
this model 0.722 0.064 0.347 0.758 0.318
Tülu-3-8B-SFT, untuned (Tier 0) 0.603 0.099 0.470 0.644 0.448
same recipe, supervised only 0.708 0.070 0.362 0.739 0.321
same recipe from the DPO checkpoint 0.708 0.064 0.359 0.744 0.314
Jev 1.13.0 (TypeSafe API) 0.733 0.113 0.349 0.835 0.432

Coherence and invariance — sure loss: mean d² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers.

model sure loss share incoherent order flip tag TVD K=2→max acc drop
this model 0.030 0.516 0.090 0.033 0.226
Tülu-3-8B-SFT, untuned (Tier 0) 0.200 0.978 0.311 0.078 0.420
same recipe, supervised only 0.317 0.955 0.098 0.035 0.237
same recipe from the DPO checkpoint 0.033 0.464 0.096 0.035 0.241
Jev 1.13.0 (TypeSafe API) 0.081 0.725 0.046 — 0.246

Out of distribution — stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks.

model stated rule 'none' when gone hijack phishing AUROC tool risk
this model 0.575 0.532 0.044 0.499 0.733
Tülu-3-8B-SFT, untuned (Tier 0) 0.558 0.212 0.165 0.771 0.700
same recipe, supervised only 0.577 0.470 0.092 0.476 0.817
same recipe from the DPO checkpoint 0.575 0.514 0.016 0.485 0.700
Jev 1.13.0 (TypeSafe API) 0.924 0.744 0.205 0.688 0.933

Reproduction check. Loading this folder with load_jevified and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Δp| 0.004, max 0.023 (the adapter is merged into bf16 weights at load).

How it was trained

The model is trained on its own decision readout — the distribution over the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring rule, plus a coherence penalty (weight 1.0): every training question comes with automatically derived siblings (the options as yes/no questions, the negation, the threshold questions of a scale), and the de Finetti sure loss of the family's answers is penalised, so the model's answers to related questions stay mutually consistent. Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources (5,885 families, at most 400 records per source); lr 3e-05, 2 epochs, best epoch by validation loss (epoch 1), seed 0. A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits. The six held-out sources (clinc150, arc_challenge, yelp5, measuring_hate_speech, fever_evidence, strategyqa_grounded) never appeared in training.

Files

  • jevify_config.json — the recipe, the backbone and the training settings load_jevified reads
  • lora/ — the adapter, merged into the backbone at load
  • results/test_metrics.json — every jev-bench config; recipe.json — the fitted recipe
  • results/coherence.json, probes.json, tags.json — the coherence, probe and tag tests
  • results/train.json — the training log; summary.json — this model's row of the study table
  • results/verification.json — the reproduction check reported under Results

Related models

Limitations

  • One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions.
  • The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it.
  • English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Praveenrajus/Llama-3.1-Tulu-3-8B-SFT-jevify-readout-coh

Finetuned
(25)
this model

Dataset used to train Praveenrajus/Llama-3.1-Tulu-3-8B-SFT-jevify-readout-coh