modern-bert-jev

Pick the best option from a list β€” with percentages you can trust.

Give it some background, a question, and a list of options. It reads every option against the background and tells you how likely each one is.

Background:  "How long will it take for my ID to verify?"
Question:    Which banking support intent does this express?
Options:     Card arrival / Lost or stolen card / Unable to verify identity / … (77 total)

          β†’  Unable to verify identity        β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ  81%
             Why verify identity              β–ˆβ–ˆβ–ˆ                   11%
             Verify my identity               β–ˆβ–ˆ                     5%

Two options or 151 β€” same model either way. Nothing is wired to a fixed list, so you can hand it categories it has never seen and it will still rank them.

The percentages are calibrated: when it says 70%, it is right about 70% of the time. That is the unusual part β€” most classifiers are badly overconfident.

πŸ‘‰ Try the live demo Β· Code Β· Colab

What it is good at

Sorting text into a known list of categories. Tested on 9,599 examples it had never seen:

Task Options Gets it right Random guessing
Bank support messages 77 86% 1%
Voice assistant commands 60 84% 2%
Customer service intents 151 83% 1%
Sentence logic 3 77% 33%
Legal contract clauses 100 76% 1%
Emotion in a comment 28 62% 4%
Disputed sentence logic 3 53% 33%
School science questions 4 35% 25%
University exam questions 4 28% 25%

Overall 64% correct where guessing gets 15%. Calibrated error is 0.053, meaning its stated confidence is off by about 5 points on average.

Against the benchmark's own reference system

jev-1.13.0, scored on the identical 9,599 test rows. This model wins 4 of 9 sources on accuracy and 6 of 9 on NLL. The whole accuracy deficit is the two exam-question tasks; drop those and it leads on accuracy as well:

seven non-knowledge sources this model jev-1.13.0
accuracy 0.723 0.710
macro ECE 0.068 0.139
macro NLL 0.910 2.747
macro Brier 0.263 0.421

Full per-source table in BENCHMARK.md.

What it is bad at β€” read this first

Anything needing world knowledge. The bottom two rows are the honest warning: 28% on university exam questions against 25% for guessing is noise, not a model. It is small and has not memorised facts. Use a large language model for trivia and exams.

Also: English only. Trained on one random seed, so there is no error bar. It reads roughly the first 400 words of your background and ignores the rest. And it was stopped too early β€” the training curve was still improving when the budget ran out.

Using it

This is not a transformers model and will not load with AutoModel. It is a small adapter β€” 6.5 MB β€” that sits on top of frozen ModernBERT-base. You need the repo's code.

git clone https://github.com/ali-rehman-ML/modern-bert-jev
cd modern-bert-jev && pip install torch transformers
from predict import Predictor

predictor = Predictor("runs/base-full")
print(predictor({
    "question": "Which emotion does the comment primarily express?",
    "context": "i stay as quiet as i can until im caught",
    "choices": {"anger": "anger", "annoyance": "annoyance", "fear": "fear", "joy": "joy"},
}))

Prefer ONNX? onnx/model.onnx is the same model with the LoRA folded into the weights, so it runs under onnxruntime with no PyTorch at all. It matches PyTorch to 1e-4 on identical tokens. There is deliberately no int8 build: dynamic quantization, per-tensor and per-channel alike, moved scores by 5–8 and flipped MNLI's argmax.

If you build your own loop: divide the scores by 1.2023 (in calibration.json) before the softmax. Without it the percentages are overconfident. It cannot change which option wins, only how sure the model claims to be.

How it works

Instead of one big output layer with a slot per category, it scores one option at a time. Each option is glued to your background and question, read by the model, and turned into a single number. A softmax over those numbers gives the percentages.

background + question + option 1  β†’  model  β†’  4.2  ┐
background + question + option 2  β†’  model  β†’  1.8  β”œβ†’  softmax  β†’  81% / 11% / 5%
background + question + option 3  β†’  model  β†’  0.3  β”˜

The model never sees the competing options, which is exactly why their number does not matter. The cost is that it runs once per option β€” 151 options means 151 passes.

Technical details
Trainable 1,622,785 params (1.08%) β€” 1,622,016 LoRA + 769 head
Backbone answerdotai/ModernBERT-base @ 8949b909ec900327062f0ebf497f51aef5e6f0c8, frozen
LoRA rank 16, alpha 32, on attn.Wqkv and attn.Wo of all 22 layers
Head mask-weighted mean pool of last_hidden_state, then Linear(768 β†’ 1)
Input Context: … as segment A, Question: …\nCandidate: … as segment B, 512 tokens
Truncation context only; question and candidate are never cut

Training. Full jev-bench choice dataset, no choice-count limit on any split: 49,364 train / 1,000 validation / 1,000 calibration / 9,599 test. 6,171 steps, 2 epochs, 114 minutes on one H100, 26.1 GiB peak.

Training samples at most 32 candidates per example β€” every candidate carrying target mass, plus random negatives, target renormalized over what survived. Mean candidates per example drops from 68.0 to 25.9. Validation, calibration and test always score the full published set.

AdamW (wd 0.01), lr 2e-4 with 5% linear warmup then cosine, grad clip 1.0, bf16 autocast, gradient checkpointing. Checkpoints selected on validation NLL, not accuracy.

Loss is soft-target cross-entropy, so human label distributions (GoEmotions, ChaosNLI) train against their real distribution rather than an argmax.

Full results.

uncalibrated calibrated (T = 1.2023)
Accuracy 0.6379 0.6379
Target NLL 1.0687 1.0351
10-bin ECE 0.0837 0.0525
Brier 0.3715 0.3646

test_predictions.json ships raw per-candidate scores for all 9,599 rows, so every number here is recomputable without a GPU.

A data bug worth knowing about. jev-bench publishes null option descriptions for LEDGAR and GoEmotions in 100% of rows on every split. Every option then rendered as the literal string Candidate: None, all options of an example shared one input, and the model returned identical scores by construction β€” exactly ln K loss and 1/K confidence, untrainable at any budget. The repo falls back to the option's ID and raises if a record's options still all read alike. Legal clause classification went from 0.009 to 0.761 on that fix alone.

License and credit

Apache-2.0. Built on ModernBERT (Apache-2.0). Trained on jev-bench revision cbcb6703ef02b698bee32af39d3270fe1f356187; no dataset content is redistributed here.

Source licenses are recorded in data_audit.json and include CC-BY-SA-4.0 for ARC-Challenge and ChaosNLI, CC-BY-4.0 for Banking77, MASSIVE and LEDGAR, CC-BY-3.0 for CLINC150, Apache-2.0 for GoEmotions, MIT for MMLU, and research-use terms for MultiNLI. Check those before commercial use.

Downloads last month
38
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ali-rehman-ML/modern-bert-jev

Adapter
(43)
this model

Dataset used to train ali-rehman-ML/modern-bert-jev

Space using ali-rehman-ML/modern-bert-jev 1

Evaluation results