Llama-3.1-Tülu-3-8B-DPO, readout fine-tuned (LoRA, supervised)

Built with Llama. Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

A System One decision model: it reads a state, answers typed questions (choice, score, noul) and returns calibrated probability distributions your code can branch on — it never writes text. This repo is a rank-16 LoRA (41,943,040 parameters) on allenai/Llama-3.1-Tulu-3-8B-DPO, merged into the weights at load, trained on its own decision readout.

At a glance — accuracy 0.704 · ECE 0.072 · held-out 0.739 · TVD to human labels 0.328 · sure loss 0.319
same order, Tülu-3-8B-DPO, untuned (Tier 0): 0.607 / 0.110 / 0.637 / 0.451 / 0.271
same order, Jev 1.13.0: 0.733 / 0.113 / 0.835 / 0.432 / 0.081

jev-bench · leaderboard · findings · code

Use it

from jevify import load_jevified

model = load_jevified("Praveenrajus/Llama-3.1-Tulu-3-8B-DPO-jevify-readout")
model.ask({"text": "The battery lasted two days on a single charge."},
          {"q": {"type": "noul", "instructions": "Is the review positive?"}})

jevify-serve --model Praveenrajus/Llama-3.1-Tulu-3-8B-DPO-jevify-readout serves it as a drop-in for the TypeSafe SDK (TYPESAFE_BASE_URL=http://localhost:8000). The backbone is pulled from its own repo at load, pinned to commit a7beb67e33ffd01cc87ac3b46cadc1000985b8db.

Results

Every number is on the jev-bench test splits (22,773 records) or the study's other test suites, scored the same way for every model; the rows under this model are references from the same study.

Decisions and calibration

model acc ECE Brier held-out acc TVD to human labels
this model 0.704 0.072 0.367 0.739 0.328
Tülu-3-8B-DPO, untuned (Tier 0) 0.607 0.110 0.473 0.637 0.451
same recipe + coherence 0.708 0.064 0.359 0.744 0.314
same recipe from the SFT checkpoint 0.708 0.070 0.362 0.739 0.321
Jev 1.13.0 (TypeSafe API) 0.733 0.113 0.349 0.835 0.432

Coherence and invariance — sure loss: mean d² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers.

model sure loss share incoherent order flip tag TVD K=2→max acc drop
this model 0.319 0.948 0.104 0.036 0.248
Tülu-3-8B-DPO, untuned (Tier 0) 0.271 0.966 0.308 0.065 0.398
same recipe + coherence 0.033 0.464 0.096 0.035 0.241
same recipe from the SFT checkpoint 0.317 0.955 0.098 0.035 0.237
Jev 1.13.0 (TypeSafe API) 0.081 0.725 0.046 — 0.246

Out of distribution — stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks.

model stated rule 'none' when gone hijack phishing AUROC tool risk
this model 0.565 0.436 0.082 0.516 0.717
Tülu-3-8B-DPO, untuned (Tier 0) 0.560 0.120 0.188 0.709 0.733
same recipe + coherence 0.575 0.514 0.016 0.485 0.700
same recipe from the SFT checkpoint 0.577 0.470 0.092 0.476 0.817
Jev 1.13.0 (TypeSafe API) 0.924 0.744 0.205 0.688 0.933

Reproduction check. Loading this folder with load_jevified and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Δp| 0.004, max 0.019 (the adapter is merged into bf16 weights at load).

How it was trained

The model is trained on its own decision readout — the distribution over the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring rule. Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources (5,885 families, at most 400 records per source); lr 3e-05, 2 epochs, best epoch by validation loss (epoch 0), seed 0. A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits. The six held-out sources (clinc150, arc_challenge, yelp5, measuring_hate_speech, fever_evidence, strategyqa_grounded) never appeared in training.

Files

  • jevify_config.json — the recipe, the backbone and the training settings load_jevified reads
  • lora/ — the adapter, merged into the backbone at load
  • results/test_metrics.json — every jev-bench config; recipe.json — the fitted recipe
  • results/coherence.json, probes.json, tags.json — the coherence, probe and tag tests
  • results/train.json — the training log; summary.json — this model's row of the study table
  • results/verification.json — the reproduction check reported under Results

Related models

Limitations

  • One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions.
  • The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it.
  • English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Praveenrajus/Llama-3.1-Tulu-3-8B-DPO-jevify-readout

Finetuned
(9)
this model

Dataset used to train Praveenrajus/Llama-3.1-Tulu-3-8B-DPO-jevify-readout