Text Classification
Transformers
Safetensors
English
Russian
open-jev
distillation
cross-encoder
calibrated
Instructions to use sshalimov04/open-jev-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sshalimov04/open-jev-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="sshalimov04/open-jev-base")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("sshalimov04/open-jev-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from sshalimov04/open-jev-base: direct link, hf CLI and curl.
- Browser
- Download file 6.91 kB
-
https://huggingface.co/sshalimov04/open-jev-base/resolve/main/README.md
- Command line
-
hf download hf://sshalimov04/open-jev-base/README.md
-
curl -L -o README.md https://huggingface.co/sshalimov04/open-jev-base/resolve/main/README.md
6.91 kB
| library_name: transformers | |
| pipeline_tag: text-classification | |
| base_model: jhu-clsp/mmBERT-small | |
| license: mit | |
| language: [en, ru] | |
| tags: [open-jev, distillation, cross-encoder, calibrated] | |
| # open-jev-base (base-none-v2) | |
| A **pointwise cross-encoder**: one forward pass scores one `(question, option, text)` triple and | |
| returns the probability that this option is the right answer for this text. A K-way decision is K | |
| passes, renormalised over the K offered options; a yes/no (`noul`) decision is one pass. Nothing | |
| about K or the option wording lives in the head, so **both are free at inference** — a new question | |
| with new options needs no retraining, only a new prompt string. | |
| Trained by [open-jev](https://github.com/sshalimov04/open-jev) on the pooled soft labels of | |
| 40 typed-decision tasks (29 labelled by the Jev teacher, 11 by a | |
| local Qwen). This is the `--fold none` artefact: **every** task in the mixture is held in, so it | |
| must not be scored on any of them. | |
| ## Use it | |
| ```python | |
| from openjev import base | |
| from openjev.spec import load_task | |
| tok, model = base.load("sshalimov04/open-jev-base") | |
| task = load_task("tasks/kinopoisk.yaml") # any task YAML: your question, your options | |
| logp = base.predict_logits(tok, model, task, ["Фильм затянут, но актёры вытягивают."]) | |
| probs = logp.exp() # [N, K], renormalised over the K options | |
| ``` | |
| The string the model actually sees (`openjev/pairs.py:render`) is a sentence pair — segment A is the | |
| decision and is never truncated, segment B is the text and takes the rest of the 512-token window: | |
| ``` | |
| A: choice | Q: <question> | option: <this option's line> | options: <all option lines, cut at 96 tokens> | |
| B: <the text> | |
| ``` | |
| **Calibrate it before you trust the probabilities.** Every number below uses the `gold500` variant: | |
| a temperature + per-class bias fitted on that task's own calib rows (200 on the unseen sets, 500 on | |
| the leave-one-source-out tasks). Raw, the renormalised sigmoids are usable at small K and unusable | |
| at K = 77 (ECE 0.307 on banking77, repaired to 0.055 by the fit). | |
| ## What it does on questions it was never trained on | |
| Three preregistered sets, chosen and frozen before any of this model existed, over three text | |
| sources not in the mixture. n = 500 eval rows each, **200 gold calibration rows per task**. | |
| Teacher-normalised = (metric − chance) / (teacher − chance), sign-flipped for MAE. | |
| | task | metric | n | K | chance | teacher | this model | advantage over chance (95 % CI) | teacher-normalised | | |
| |---|---|---|---|---|---|---|---|---| | |
| | yahoo-topics | acc | 500 | 10 | 0.088 | 0.718 | **0.528** | +0.440 (+0.386..+0.492) | 0.70 | | |
| | sst5 | mae | 500 | 5 | 1.152 | 0.493 | **0.775** | +0.377 (+0.324..+0.435) | 0.57 | | |
| | ru-inappropriate | auroc | 500 | 2 | 0.500 | 0.891 | **0.619** | +0.119 (+0.072..+0.168) | 0.30 | | |
| It beats chance on all three with the CI excluding 0, and recovers 70 % / 57 % / 30 % of the | |
| teacher's own margin over chance. That is the whole claim. It is **not** a general model and it does | |
| **not** answer any question: the rule it was written against passed at ≥ 0.5 on 2 of 3 sets, the | |
| weakest result is the Russian safety question, and under a no-gold calibration (`teacher500`) sst5 | |
| falls to 0.48 and the rule would be met on 1 of 3. A deployment on a new question has to supply | |
| those 200 gold rows. | |
| ## What it loses to (the same recipe, held-out text sources) | |
| A card that only lists wins is the thing this repo exists not to be. These are the fold models | |
| (`base-F1-v2`, `base-F2-v2`) — same architecture, same recipe, same mixture minus the held-out | |
| source — scored zero-shot against the per-task distilled student for that task: | |
| | task | fold | metric | n | chance | teacher | per-task student | this recipe, zero-shot | Δ vs student (95 % CI) | | |
| |---|---|---|---|---|---|---|---|---| | |
| | kinopoisk | F1 | acc | 1500 | 0.333 | 0.652 | 0.657 | 0.519 | -0.138 (-0.164..-0.112) | | |
| | swde-field | F1 | acc | 2000 | 0.052 | 0.803 | 0.869 | 0.317 | -0.552 (-0.574..-0.529) | | |
| | toxic | F1 | auroc | 3000 | 0.500 | 0.817 | 0.856 | 0.829 | -0.027 (-0.055..-0.000) | | |
| | georeview | F2 | mae | 2000 | 1.201 | 0.619 | 0.607 | 0.767 | +0.161 (+0.143..+0.179) | | |
| | banking77 | F2 | acc | 2000 | 0.015 | 0.764 | 0.750 | 0.372 | -0.378 (-0.403..-0.355) | | |
| | m2w-element | F2 | acc | 900 | 0.049 | 0.649 | 0.059 | 0.136 | +0.077 (+0.051..+0.103) | | |
| **It loses to every per-task student except one**, every CI excluding 0. The exception is | |
| `m2w-element`, where the per-task student is a fixed 16-way head over row-specific candidates and | |
| collapses to 0.059 — a pair scorer reads the candidate's own line, which is the argument for this | |
| architecture and not evidence that the model is useful there (it is still 51 pts below the teacher). | |
| Converging the training made transfer **worse** on 5 of these 6 tasks than an earlier 45-minute | |
| budget did. Note that `results/base/base-F2-v2-zs-m2w-element-gold500.json` carries `nll: inf` and | |
| `ece: 0.855`: the vector fit drove one option's calibrated probability to 0 on a row that carries it | |
| as gold. Accuracy is unaffected; anything reading `nll` off that file is reading a degenerate fit. | |
| ## What is not measured | |
| The mixture contains 27 small tasks beyond the original ones, and **whether they bought anything is | |
| unknown**. The ablation that was built to answer it is **void by its own preregistered STOP rule**: | |
| 2 of the 3 control-arm seeds never converged, and the prereg forbids comparing a converged arm with | |
| a non-converged one. Its numbers happen to favour the full mixture, which is exactly the direction | |
| an under-trained control would fake, so they are not cited here. **This model's mixture advantage is | |
| unmeasured.** | |
| Also not measured: any recalibration under drift, any language beyond the en/ru mix of the training | |
| tasks, and any behaviour at K far above the 77 seen in training. | |
| ## Training | |
| - **Init**: `jhu-clsp/mmBERT-small` (~140M), `AutoModelForSequenceClassification(num_labels=1)`, | |
| BCE against the teacher's probability of that option. | |
| - **Teachers**: `typesafe/jev-1.13-20260917` (29 run dirs) and a local `Qwen/Qwen3.8-27B-FP8` | |
| (11 run dirs, the original tasks). No row was re-labelled for this model. | |
| - **Mixture**: 234,316 pairs/epoch (cap 12,000 per task per epoch) over 80,872 teacher rows. | |
| - **Run**: batch 32, lr 5e-05, max_len 512, 15,000 steps (3 epochs), | |
| early-stopped on held-in calib BCE (min_delta 0.001, patience 3), best | |
| 0.3569, converged: true. 100 min, 8.91 GB peak on one GB10. | |
| ## Licence and provenance | |
| MIT. Base model [`jhu-clsp/mmBERT-small`](https://huggingface.co/jhu-clsp/mmBERT-small). | |
| Code, task YAMLs and every table above: [open-jev](https://github.com/sshalimov04/open-jev) — | |
| the lab notebook is `docs/experiments.md` §B2 and the preregistration that fixed these thresholds | |
| before the runs existed is `docs/prereg/base-v2.md`. | |