File size: 6,909 Bytes
bf071a6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
---
library_name: transformers
pipeline_tag: text-classification
base_model: jhu-clsp/mmBERT-small
license: mit
language: [en, ru]
tags: [open-jev, distillation, cross-encoder, calibrated]
---

# open-jev-base (base-none-v2)

A **pointwise cross-encoder**: one forward pass scores one `(question, option, text)` triple and
returns the probability that this option is the right answer for this text. A K-way decision is K
passes, renormalised over the K offered options; a yes/no (`noul`) decision is one pass. Nothing
about K or the option wording lives in the head, so **both are free at inference** — a new question
with new options needs no retraining, only a new prompt string.

Trained by [open-jev](https://github.com/sshalimov04/open-jev) on the pooled soft labels of
40 typed-decision tasks (29 labelled by the Jev teacher, 11 by a
local Qwen). This is the `--fold none` artefact: **every** task in the mixture is held in, so it
must not be scored on any of them.

## Use it

```python
from openjev import base
from openjev.spec import load_task

tok, model = base.load("sshalimov04/open-jev-base")
task = load_task("tasks/kinopoisk.yaml")          # any task YAML: your question, your options
logp = base.predict_logits(tok, model, task, ["Фильм затянут, но актёры вытягивают."])
probs = logp.exp()                                # [N, K], renormalised over the K options
```

The string the model actually sees (`openjev/pairs.py:render`) is a sentence pair — segment A is the
decision and is never truncated, segment B is the text and takes the rest of the 512-token window:

```
A: choice | Q: <question> | option: <this option's line> | options: <all option lines, cut at 96 tokens>
B: <the text>
```

**Calibrate it before you trust the probabilities.** Every number below uses the `gold500` variant:
a temperature + per-class bias fitted on that task's own calib rows (200 on the unseen sets, 500 on
the leave-one-source-out tasks). Raw, the renormalised sigmoids are usable at small K and unusable
at K = 77 (ECE 0.307 on banking77, repaired to 0.055 by the fit).

## What it does on questions it was never trained on

Three preregistered sets, chosen and frozen before any of this model existed, over three text
sources not in the mixture. n = 500 eval rows each, **200 gold calibration rows per task**.
Teacher-normalised = (metric − chance) / (teacher − chance), sign-flipped for MAE.

| task | metric | n | K | chance | teacher | this model | advantage over chance (95 % CI) | teacher-normalised |
|---|---|---|---|---|---|---|---|---|
| yahoo-topics | acc | 500 | 10 | 0.088 | 0.718 | **0.528** | +0.440 (+0.386..+0.492) | 0.70 |
| sst5 | mae | 500 | 5 | 1.152 | 0.493 | **0.775** | +0.377 (+0.324..+0.435) | 0.57 |
| ru-inappropriate | auroc | 500 | 2 | 0.500 | 0.891 | **0.619** | +0.119 (+0.072..+0.168) | 0.30 |

It beats chance on all three with the CI excluding 0, and recovers 70 % / 57 % / 30 % of the
teacher's own margin over chance. That is the whole claim. It is **not** a general model and it does
**not** answer any question: the rule it was written against passed at ≥ 0.5 on 2 of 3 sets, the
weakest result is the Russian safety question, and under a no-gold calibration (`teacher500`) sst5
falls to 0.48 and the rule would be met on 1 of 3. A deployment on a new question has to supply
those 200 gold rows.

## What it loses to (the same recipe, held-out text sources)

A card that only lists wins is the thing this repo exists not to be. These are the fold models
(`base-F1-v2`, `base-F2-v2`) — same architecture, same recipe, same mixture minus the held-out
source — scored zero-shot against the per-task distilled student for that task:

| task | fold | metric | n | chance | teacher | per-task student | this recipe, zero-shot | Δ vs student (95 % CI) |
|---|---|---|---|---|---|---|---|---|
| kinopoisk | F1 | acc | 1500 | 0.333 | 0.652 | 0.657 | 0.519 | -0.138 (-0.164..-0.112) |
| swde-field | F1 | acc | 2000 | 0.052 | 0.803 | 0.869 | 0.317 | -0.552 (-0.574..-0.529) |
| toxic | F1 | auroc | 3000 | 0.500 | 0.817 | 0.856 | 0.829 | -0.027 (-0.055..-0.000) |
| georeview | F2 | mae | 2000 | 1.201 | 0.619 | 0.607 | 0.767 | +0.161 (+0.143..+0.179) |
| banking77 | F2 | acc | 2000 | 0.015 | 0.764 | 0.750 | 0.372 | -0.378 (-0.403..-0.355) |
| m2w-element | F2 | acc | 900 | 0.049 | 0.649 | 0.059 | 0.136 | +0.077 (+0.051..+0.103) |

**It loses to every per-task student except one**, every CI excluding 0. The exception is
`m2w-element`, where the per-task student is a fixed 16-way head over row-specific candidates and
collapses to 0.059 — a pair scorer reads the candidate's own line, which is the argument for this
architecture and not evidence that the model is useful there (it is still 51 pts below the teacher).
Converging the training made transfer **worse** on 5 of these 6 tasks than an earlier 45-minute
budget did. Note that `results/base/base-F2-v2-zs-m2w-element-gold500.json` carries `nll: inf` and
`ece: 0.855`: the vector fit drove one option's calibrated probability to 0 on a row that carries it
as gold. Accuracy is unaffected; anything reading `nll` off that file is reading a degenerate fit.

## What is not measured

The mixture contains 27 small tasks beyond the original ones, and **whether they bought anything is
unknown**. The ablation that was built to answer it is **void by its own preregistered STOP rule**:
2 of the 3 control-arm seeds never converged, and the prereg forbids comparing a converged arm with
a non-converged one. Its numbers happen to favour the full mixture, which is exactly the direction
an under-trained control would fake, so they are not cited here. **This model's mixture advantage is
unmeasured.**

Also not measured: any recalibration under drift, any language beyond the en/ru mix of the training
tasks, and any behaviour at K far above the 77 seen in training.

## Training

- **Init**: `jhu-clsp/mmBERT-small` (~140M), `AutoModelForSequenceClassification(num_labels=1)`,
  BCE against the teacher's probability of that option.
- **Teachers**: `typesafe/jev-1.13-20260917` (29 run dirs) and a local `Qwen/Qwen3.8-27B-FP8`
  (11 run dirs, the original tasks). No row was re-labelled for this model.
- **Mixture**: 234,316 pairs/epoch (cap 12,000 per task per epoch) over 80,872 teacher rows.
- **Run**: batch 32, lr 5e-05, max_len 512, 15,000 steps (3 epochs),
  early-stopped on held-in calib BCE (min_delta 0.001, patience 3), best
  0.3569, converged: true. 100 min, 8.91 GB peak on one GB10.

## Licence and provenance

MIT. Base model [`jhu-clsp/mmBERT-small`](https://huggingface.co/jhu-clsp/mmBERT-small).
Code, task YAMLs and every table above: [open-jev](https://github.com/sshalimov04/open-jev) —
the lab notebook is `docs/experiments.md` §B2 and the preregistration that fixed these thresholds
before the runs existed is `docs/prereg/base-v2.md`.