Instructions to use sshalimov04/open-jev-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sshalimov04/open-jev-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="sshalimov04/open-jev-base")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("sshalimov04/open-jev-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
open-jev-base (base-none-v2)
A pointwise cross-encoder: one forward pass scores one (question, option, text) triple and
returns the probability that this option is the right answer for this text. A K-way decision is K
passes, renormalised over the K offered options; a yes/no (noul) decision is one pass. Nothing
about K or the option wording lives in the head, so both are free at inference — a new question
with new options needs no retraining, only a new prompt string.
Trained by open-jev on the pooled soft labels of
40 typed-decision tasks (29 labelled by the Jev teacher, 11 by a
local Qwen). This is the --fold none artefact: every task in the mixture is held in, so it
must not be scored on any of them.
Use it
from openjev import base
from openjev.spec import load_task
tok, model = base.load("sshalimov04/open-jev-base")
task = load_task("tasks/kinopoisk.yaml") # any task YAML: your question, your options
logp = base.predict_logits(tok, model, task, ["Фильм затянут, но актёры вытягивают."])
probs = logp.exp() # [N, K], renormalised over the K options
The string the model actually sees (openjev/pairs.py:render) is a sentence pair — segment A is the
decision and is never truncated, segment B is the text and takes the rest of the 512-token window:
A: choice | Q: <question> | option: <this option's line> | options: <all option lines, cut at 96 tokens>
B: <the text>
Calibrate it before you trust the probabilities. Every number below uses the gold500 variant:
a temperature + per-class bias fitted on that task's own calib rows (200 on the unseen sets, 500 on
the leave-one-source-out tasks). Raw, the renormalised sigmoids are usable at small K and unusable
at K = 77 (ECE 0.307 on banking77, repaired to 0.055 by the fit).
What it does on questions it was never trained on
Three preregistered sets, chosen and frozen before any of this model existed, over three text sources not in the mixture. n = 500 eval rows each, 200 gold calibration rows per task. Teacher-normalised = (metric − chance) / (teacher − chance), sign-flipped for MAE.
| task | metric | n | K | chance | teacher | this model | advantage over chance (95 % CI) | teacher-normalised |
|---|---|---|---|---|---|---|---|---|
| yahoo-topics | acc | 500 | 10 | 0.088 | 0.718 | 0.528 | +0.440 (+0.386..+0.492) | 0.70 |
| sst5 | mae | 500 | 5 | 1.152 | 0.493 | 0.775 | +0.377 (+0.324..+0.435) | 0.57 |
| ru-inappropriate | auroc | 500 | 2 | 0.500 | 0.891 | 0.619 | +0.119 (+0.072..+0.168) | 0.30 |
It beats chance on all three with the CI excluding 0, and recovers 70 % / 57 % / 30 % of the
teacher's own margin over chance. That is the whole claim. It is not a general model and it does
not answer any question: the rule it was written against passed at ≥ 0.5 on 2 of 3 sets, the
weakest result is the Russian safety question, and under a no-gold calibration (teacher500) sst5
falls to 0.48 and the rule would be met on 1 of 3. A deployment on a new question has to supply
those 200 gold rows.
What it loses to (the same recipe, held-out text sources)
A card that only lists wins is the thing this repo exists not to be. These are the fold models
(base-F1-v2, base-F2-v2) — same architecture, same recipe, same mixture minus the held-out
source — scored zero-shot against the per-task distilled student for that task:
| task | fold | metric | n | chance | teacher | per-task student | this recipe, zero-shot | Δ vs student (95 % CI) |
|---|---|---|---|---|---|---|---|---|
| kinopoisk | F1 | acc | 1500 | 0.333 | 0.652 | 0.657 | 0.519 | -0.138 (-0.164..-0.112) |
| swde-field | F1 | acc | 2000 | 0.052 | 0.803 | 0.869 | 0.317 | -0.552 (-0.574..-0.529) |
| toxic | F1 | auroc | 3000 | 0.500 | 0.817 | 0.856 | 0.829 | -0.027 (-0.055..-0.000) |
| georeview | F2 | mae | 2000 | 1.201 | 0.619 | 0.607 | 0.767 | +0.161 (+0.143..+0.179) |
| banking77 | F2 | acc | 2000 | 0.015 | 0.764 | 0.750 | 0.372 | -0.378 (-0.403..-0.355) |
| m2w-element | F2 | acc | 900 | 0.049 | 0.649 | 0.059 | 0.136 | +0.077 (+0.051..+0.103) |
It loses to every per-task student except one, every CI excluding 0. The exception is
m2w-element, where the per-task student is a fixed 16-way head over row-specific candidates and
collapses to 0.059 — a pair scorer reads the candidate's own line, which is the argument for this
architecture and not evidence that the model is useful there (it is still 51 pts below the teacher).
Converging the training made transfer worse on 5 of these 6 tasks than an earlier 45-minute
budget did. Note that results/base/base-F2-v2-zs-m2w-element-gold500.json carries nll: inf and
ece: 0.855: the vector fit drove one option's calibrated probability to 0 on a row that carries it
as gold. Accuracy is unaffected; anything reading nll off that file is reading a degenerate fit.
What is not measured
The mixture contains 27 small tasks beyond the original ones, and whether they bought anything is unknown. The ablation that was built to answer it is void by its own preregistered STOP rule: 2 of the 3 control-arm seeds never converged, and the prereg forbids comparing a converged arm with a non-converged one. Its numbers happen to favour the full mixture, which is exactly the direction an under-trained control would fake, so they are not cited here. This model's mixture advantage is unmeasured.
Also not measured: any recalibration under drift, any language beyond the en/ru mix of the training tasks, and any behaviour at K far above the 77 seen in training.
Training
- Init:
jhu-clsp/mmBERT-small(~140M),AutoModelForSequenceClassification(num_labels=1), BCE against the teacher's probability of that option. - Teachers:
typesafe/jev-1.13-20260917(29 run dirs) and a localQwen/Qwen3.8-27B-FP8(11 run dirs, the original tasks). No row was re-labelled for this model. - Mixture: 234,316 pairs/epoch (cap 12,000 per task per epoch) over 80,872 teacher rows.
- Run: batch 32, lr 5e-05, max_len 512, 15,000 steps (3 epochs), early-stopped on held-in calib BCE (min_delta 0.001, patience 3), best 0.3569, converged: true. 100 min, 8.91 GB peak on one GB10.
Licence and provenance
MIT. Base model jhu-clsp/mmBERT-small.
Code, task YAMLs and every table above: open-jev —
the lab notebook is docs/experiments.md §B2 and the preregistration that fixed these thresholds
before the runs existed is docs/prereg/base-v2.md.
Model tree for sshalimov04/open-jev-base
Base model
jhu-clsp/mmBERT-small