modern-bert-jev
Pick the best option from a list β with percentages you can trust.
Give it some background, a question, and a list of options. It reads every option against the background and tells you how likely each one is.
Background: "How long will it take for my ID to verify?"
Question: Which banking support intent does this express?
Options: Card arrival / Lost or stolen card / Unable to verify identity / β¦ (77 total)
β Unable to verify identity ββββββββββββββββββββ 81%
Why verify identity βββ 11%
Verify my identity ββ 5%
Two options or 151 β same model either way. Nothing is wired to a fixed list, so you can hand it categories it has never seen and it will still rank them.
The percentages are calibrated: when it says 70%, it is right about 70% of the time. That is the unusual part β most classifiers are badly overconfident.
π Try the live demo Β· Code Β· Colab
What it is good at
Sorting text into a known list of categories. Tested on 9,599 examples it had never seen:
| Task | Options | Gets it right | Random guessing |
|---|---|---|---|
| Bank support messages | 77 | 86% | 1% |
| Voice assistant commands | 60 | 84% | 2% |
| Customer service intents | 151 | 83% | 1% |
| Sentence logic | 3 | 77% | 33% |
| Legal contract clauses | 100 | 76% | 1% |
| Emotion in a comment | 28 | 62% | 4% |
| Disputed sentence logic | 3 | 53% | 33% |
| School science questions | 4 | 35% | 25% |
| University exam questions | 4 | 28% | 25% |
Overall 64% correct where guessing gets 15%. Calibrated error is 0.053, meaning its stated confidence is off by about 5 points on average.
Against the benchmark's own reference system
jev-1.13.0, scored on the identical 9,599 test rows. This model wins 4 of 9 sources on
accuracy and 6 of 9 on NLL. The whole accuracy deficit is the two exam-question tasks;
drop those and it leads on accuracy as well:
| seven non-knowledge sources | this model | jev-1.13.0 |
|---|---|---|
| accuracy | 0.723 | 0.710 |
| macro ECE | 0.068 | 0.139 |
| macro NLL | 0.910 | 2.747 |
| macro Brier | 0.263 | 0.421 |
Full per-source table in BENCHMARK.md.
What it is bad at β read this first
Anything needing world knowledge. The bottom two rows are the honest warning: 28% on university exam questions against 25% for guessing is noise, not a model. It is small and has not memorised facts. Use a large language model for trivia and exams.
Also: English only. Trained on one random seed, so there is no error bar. It reads roughly the first 400 words of your background and ignores the rest. And it was stopped too early β the training curve was still improving when the budget ran out.
Using it
This is not a transformers model and will not load with AutoModel. It is a small
adapter β 6.5 MB β that sits on top of frozen ModernBERT-base. You need the repo's code.
git clone https://github.com/ali-rehman-ML/modern-bert-jev
cd modern-bert-jev && pip install torch transformers
from predict import Predictor
predictor = Predictor("runs/base-full")
print(predictor({
"question": "Which emotion does the comment primarily express?",
"context": "i stay as quiet as i can until im caught",
"choices": {"anger": "anger", "annoyance": "annoyance", "fear": "fear", "joy": "joy"},
}))
Prefer ONNX? onnx/model.onnx is the same model with the LoRA folded into the weights, so
it runs under onnxruntime with no PyTorch at all. It matches PyTorch to 1e-4 on identical
tokens. There is deliberately no int8 build: dynamic quantization, per-tensor and per-channel
alike, moved scores by 5β8 and flipped MNLI's argmax.
If you build your own loop: divide the scores by 1.2023 (in calibration.json) before
the softmax. Without it the percentages are overconfident. It cannot change which option wins,
only how sure the model claims to be.
How it works
Instead of one big output layer with a slot per category, it scores one option at a time. Each option is glued to your background and question, read by the model, and turned into a single number. A softmax over those numbers gives the percentages.
background + question + option 1 β model β 4.2 β
background + question + option 2 β model β 1.8 ββ softmax β 81% / 11% / 5%
background + question + option 3 β model β 0.3 β
The model never sees the competing options, which is exactly why their number does not matter. The cost is that it runs once per option β 151 options means 151 passes.
Technical details
| Trainable | 1,622,785 params (1.08%) β 1,622,016 LoRA + 769 head |
| Backbone | answerdotai/ModernBERT-base @ 8949b909ec900327062f0ebf497f51aef5e6f0c8, frozen |
| LoRA | rank 16, alpha 32, on attn.Wqkv and attn.Wo of all 22 layers |
| Head | mask-weighted mean pool of last_hidden_state, then Linear(768 β 1) |
| Input | Context: β¦ as segment A, Question: β¦\nCandidate: β¦ as segment B, 512 tokens |
| Truncation | context only; question and candidate are never cut |
Training. Full jev-bench choice dataset, no choice-count limit on any split: 49,364 train / 1,000 validation / 1,000 calibration / 9,599 test. 6,171 steps, 2 epochs, 114 minutes on one H100, 26.1 GiB peak.
Training samples at most 32 candidates per example β every candidate carrying target mass, plus random negatives, target renormalized over what survived. Mean candidates per example drops from 68.0 to 25.9. Validation, calibration and test always score the full published set.
AdamW (wd 0.01), lr 2e-4 with 5% linear warmup then cosine, grad clip 1.0, bf16 autocast, gradient checkpointing. Checkpoints selected on validation NLL, not accuracy.
Loss is soft-target cross-entropy, so human label distributions (GoEmotions, ChaosNLI) train against their real distribution rather than an argmax.
Full results.
| uncalibrated | calibrated (T = 1.2023) | |
|---|---|---|
| Accuracy | 0.6379 | 0.6379 |
| Target NLL | 1.0687 | 1.0351 |
| 10-bin ECE | 0.0837 | 0.0525 |
| Brier | 0.3715 | 0.3646 |
test_predictions.json ships raw per-candidate scores for all 9,599 rows, so every number
here is recomputable without a GPU.
A data bug worth knowing about. jev-bench publishes null option descriptions for LEDGAR
and GoEmotions in 100% of rows on every split. Every option then rendered as the literal string
Candidate: None, all options of an example shared one input, and the model returned identical
scores by construction β exactly ln K loss and 1/K confidence, untrainable at any budget.
The repo falls back to the option's ID and raises if a record's options still all read alike.
Legal clause classification went from 0.009 to 0.761 on that fix alone.
License and credit
Apache-2.0. Built on ModernBERT
(Apache-2.0). Trained on jev-bench
revision cbcb6703ef02b698bee32af39d3270fe1f356187; no dataset content is redistributed here.
Source licenses are recorded in data_audit.json and include CC-BY-SA-4.0 for
ARC-Challenge and ChaosNLI, CC-BY-4.0 for Banking77, MASSIVE and LEDGAR, CC-BY-3.0 for
CLINC150, Apache-2.0 for GoEmotions, MIT for MMLU, and research-use terms for MultiNLI. Check
those before commercial use.
- Downloads last month
- 38
Model tree for ali-rehman-ML/modern-bert-jev
Base model
answerdotai/ModernBERT-baseDataset used to train ali-rehman-ML/modern-bert-jev
Space using ali-rehman-ML/modern-bert-jev 1
Evaluation results
- Accuracy on jev-benchself-reported0.638
- Calibrated NLL on jev-benchself-reported1.035
- Calibrated ECE (10-bin) on jev-benchself-reported0.052