open-jev-base (base-none-v2)

A pointwise cross-encoder: one forward pass scores one (question, option, text) triple and returns the probability that this option is the right answer for this text. A K-way decision is K passes, renormalised over the K offered options; a yes/no (noul) decision is one pass. Nothing about K or the option wording lives in the head, so both are free at inference — a new question with new options needs no retraining, only a new prompt string.

Trained by open-jev on the pooled soft labels of 40 typed-decision tasks (29 labelled by the Jev teacher, 11 by a local Qwen). This is the --fold none artefact: every task in the mixture is held in, so it must not be scored on any of them.

Use it

from openjev import base
from openjev.spec import load_task

tok, model = base.load("sshalimov04/open-jev-base")
task = load_task("tasks/kinopoisk.yaml")          # any task YAML: your question, your options
logp = base.predict_logits(tok, model, task, ["Фильм затянут, но актёры вытягивают."])
probs = logp.exp()                                # [N, K], renormalised over the K options

The string the model actually sees (openjev/pairs.py:render) is a sentence pair — segment A is the decision and is never truncated, segment B is the text and takes the rest of the 512-token window:

A: choice | Q: <question> | option: <this option's line> | options: <all option lines, cut at 96 tokens>
B: <the text>

Calibrate it before you trust the probabilities. Every number below uses the gold500 variant: a temperature + per-class bias fitted on that task's own calib rows (200 on the unseen sets, 500 on the leave-one-source-out tasks). Raw, the renormalised sigmoids are usable at small K and unusable at K = 77 (ECE 0.307 on banking77, repaired to 0.055 by the fit).

What it does on questions it was never trained on

Three preregistered sets, chosen and frozen before any of this model existed, over three text sources not in the mixture. n = 500 eval rows each, 200 gold calibration rows per task. Teacher-normalised = (metric − chance) / (teacher − chance), sign-flipped for MAE.

task metric n K chance teacher this model advantage over chance (95 % CI) teacher-normalised
yahoo-topics acc 500 10 0.088 0.718 0.528 +0.440 (+0.386..+0.492) 0.70
sst5 mae 500 5 1.152 0.493 0.775 +0.377 (+0.324..+0.435) 0.57
ru-inappropriate auroc 500 2 0.500 0.891 0.619 +0.119 (+0.072..+0.168) 0.30

It beats chance on all three with the CI excluding 0, and recovers 70 % / 57 % / 30 % of the teacher's own margin over chance. That is the whole claim. It is not a general model and it does not answer any question: the rule it was written against passed at ≥ 0.5 on 2 of 3 sets, the weakest result is the Russian safety question, and under a no-gold calibration (teacher500) sst5 falls to 0.48 and the rule would be met on 1 of 3. A deployment on a new question has to supply those 200 gold rows.

What it loses to (the same recipe, held-out text sources)

A card that only lists wins is the thing this repo exists not to be. These are the fold models (base-F1-v2, base-F2-v2) — same architecture, same recipe, same mixture minus the held-out source — scored zero-shot against the per-task distilled student for that task:

task fold metric n chance teacher per-task student this recipe, zero-shot Δ vs student (95 % CI)
kinopoisk F1 acc 1500 0.333 0.652 0.657 0.519 -0.138 (-0.164..-0.112)
swde-field F1 acc 2000 0.052 0.803 0.869 0.317 -0.552 (-0.574..-0.529)
toxic F1 auroc 3000 0.500 0.817 0.856 0.829 -0.027 (-0.055..-0.000)
georeview F2 mae 2000 1.201 0.619 0.607 0.767 +0.161 (+0.143..+0.179)
banking77 F2 acc 2000 0.015 0.764 0.750 0.372 -0.378 (-0.403..-0.355)
m2w-element F2 acc 900 0.049 0.649 0.059 0.136 +0.077 (+0.051..+0.103)

It loses to every per-task student except one, every CI excluding 0. The exception is m2w-element, where the per-task student is a fixed 16-way head over row-specific candidates and collapses to 0.059 — a pair scorer reads the candidate's own line, which is the argument for this architecture and not evidence that the model is useful there (it is still 51 pts below the teacher). Converging the training made transfer worse on 5 of these 6 tasks than an earlier 45-minute budget did. Note that results/base/base-F2-v2-zs-m2w-element-gold500.json carries nll: inf and ece: 0.855: the vector fit drove one option's calibrated probability to 0 on a row that carries it as gold. Accuracy is unaffected; anything reading nll off that file is reading a degenerate fit.

What is not measured

The mixture contains 27 small tasks beyond the original ones, and whether they bought anything is unknown. The ablation that was built to answer it is void by its own preregistered STOP rule: 2 of the 3 control-arm seeds never converged, and the prereg forbids comparing a converged arm with a non-converged one. Its numbers happen to favour the full mixture, which is exactly the direction an under-trained control would fake, so they are not cited here. This model's mixture advantage is unmeasured.

Also not measured: any recalibration under drift, any language beyond the en/ru mix of the training tasks, and any behaviour at K far above the 77 seen in training.

Training

  • Init: jhu-clsp/mmBERT-small (~140M), AutoModelForSequenceClassification(num_labels=1), BCE against the teacher's probability of that option.
  • Teachers: typesafe/jev-1.13-20260917 (29 run dirs) and a local Qwen/Qwen3.8-27B-FP8 (11 run dirs, the original tasks). No row was re-labelled for this model.
  • Mixture: 234,316 pairs/epoch (cap 12,000 per task per epoch) over 80,872 teacher rows.
  • Run: batch 32, lr 5e-05, max_len 512, 15,000 steps (3 epochs), early-stopped on held-in calib BCE (min_delta 0.001, patience 3), best 0.3569, converged: true. 100 min, 8.91 GB peak on one GB10.

Licence and provenance

MIT. Base model jhu-clsp/mmBERT-small. Code, task YAMLs and every table above: open-jev — the lab notebook is docs/experiments.md §B2 and the preregistration that fixed these thresholds before the runs existed is docs/prereg/base-v2.md.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sshalimov04/open-jev-base

Finetuned
(53)
this model