--- license: apache-2.0 base_model: answerdotai/ModernBERT-base base_model_relation: adapter language: - en datasets: - Praveenrajus/jev-bench inference: false tags: - multiple-choice - calibration - uncertainty-estimation - confidence-estimation - lora - adapter - modernbert - encoder - intent-classification - emotion-classification - legal - natural-language-inference metrics: - accuracy model-index: - name: modern-bert-jev results: - task: type: multiple-choice name: Choice scoring (jev-bench, 9 sources) dataset: name: jev-bench type: Praveenrajus/jev-bench metrics: - type: accuracy value: 0.6379 name: Accuracy - type: loss value: 1.0351 name: Calibrated NLL - type: ece value: 0.0525 name: Calibrated ECE (10-bin) --- # modern-bert-jev **Pick the best option from a list — with percentages you can trust.** Give it some background, a question, and a list of options. It reads every option against the background and tells you how likely each one is. ``` Background: "How long will it take for my ID to verify?" Question: Which banking support intent does this express? Options: Card arrival / Lost or stolen card / Unable to verify identity / … (77 total) → Unable to verify identity ████████████████████ 81% Why verify identity ███ 11% Verify my identity ██ 5% ``` Two options or 151 — same model either way. Nothing is wired to a fixed list, so you can hand it categories it has never seen and it will still rank them. The percentages are **calibrated**: when it says 70%, it is right about 70% of the time. That is the unusual part — most classifiers are badly overconfident. 👉 **[Try the live demo](https://huggingface.co/spaces/ali-rehman-ML/modern-bert-jev-demo)** · **[Code](https://github.com/ali-rehman-ML/modern-bert-jev)** · **[Colab](https://colab.research.google.com/github/ali-rehman-ML/modern-bert-jev/blob/main/demo.ipynb)** ## What it is good at Sorting text into a known list of categories. Tested on 9,599 examples it had never seen: | Task | Options | Gets it right | Random guessing | |---|---:|---:|---:| | Bank support messages | 77 | **86%** | 1% | | Voice assistant commands | 60 | **84%** | 2% | | Customer service intents | 151 | **83%** | 1% | | Sentence logic | 3 | **77%** | 33% | | Legal contract clauses | 100 | **76%** | 1% | | Emotion in a comment | 28 | **62%** | 4% | | Disputed sentence logic | 3 | 53% | 33% | | School science questions | 4 | 35% | 25% | | University exam questions | 4 | 28% | 25% | Overall **64% correct** where guessing gets 15%. Calibrated error is 0.053, meaning its stated confidence is off by about 5 points on average. ### Against the benchmark's own reference system `jev-1.13.0`, scored on the identical 9,599 test rows. This model wins 4 of 9 sources on accuracy and **6 of 9 on NLL**. The whole accuracy deficit is the two exam-question tasks; drop those and it leads on accuracy as well: | seven non-knowledge sources | this model | jev-1.13.0 | |---|---:|---:| | accuracy | **0.723** | 0.710 | | macro ECE | **0.068** | 0.139 | | macro NLL | **0.910** | 2.747 | | macro Brier | **0.263** | 0.421 | Full per-source table in [BENCHMARK.md](BENCHMARK.md). ## What it is bad at — read this first **Anything needing world knowledge.** The bottom two rows are the honest warning: 28% on university exam questions against 25% for guessing is noise, not a model. It is small and has not memorised facts. Use a large language model for trivia and exams. Also: **English only**. Trained on **one random seed**, so there is no error bar. It reads roughly the first 400 words of your background and ignores the rest. And it was **stopped too early** — the training curve was still improving when the budget ran out. ## Using it This is **not** a `transformers` model and will not load with `AutoModel`. It is a small adapter — 6.5 MB — that sits on top of frozen `ModernBERT-base`. You need the repo's code. ```bash git clone https://github.com/ali-rehman-ML/modern-bert-jev cd modern-bert-jev && pip install torch transformers ``` ```python from predict import Predictor predictor = Predictor("runs/base-full") print(predictor({ "question": "Which emotion does the comment primarily express?", "context": "i stay as quiet as i can until im caught", "choices": {"anger": "anger", "annoyance": "annoyance", "fear": "fear", "joy": "joy"}, })) ``` **Prefer ONNX?** `onnx/model.onnx` is the same model with the LoRA folded into the weights, so it runs under `onnxruntime` with no PyTorch at all. It matches PyTorch to 1e-4 on identical tokens. There is deliberately no int8 build: dynamic quantization, per-tensor and per-channel alike, moved scores by 5–8 and flipped MNLI's argmax. **If you build your own loop:** divide the scores by **1.2023** (in `calibration.json`) before the softmax. Without it the percentages are overconfident. It cannot change which option wins, only how sure the model claims to be. ## How it works Instead of one big output layer with a slot per category, it scores **one option at a time**. Each option is glued to your background and question, read by the model, and turned into a single number. A softmax over those numbers gives the percentages. ``` background + question + option 1 → model → 4.2 ┐ background + question + option 2 → model → 1.8 ├→ softmax → 81% / 11% / 5% background + question + option 3 → model → 0.3 ┘ ``` The model never sees the competing options, which is exactly why their number does not matter. The cost is that it runs once per option — 151 options means 151 passes.
Technical details | | | |---|---| | Trainable | 1,622,785 params (1.08%) — 1,622,016 LoRA + 769 head | | Backbone | `answerdotai/ModernBERT-base` @ `8949b909ec900327062f0ebf497f51aef5e6f0c8`, frozen | | LoRA | rank 16, alpha 32, on `attn.Wqkv` and `attn.Wo` of all 22 layers | | Head | mask-weighted mean pool of `last_hidden_state`, then `Linear(768 → 1)` | | Input | `Context: …` as segment A, `Question: …\nCandidate: …` as segment B, 512 tokens | | Truncation | context only; question and candidate are never cut | **Training.** Full jev-bench choice dataset, no choice-count limit on any split: 49,364 train / 1,000 validation / 1,000 calibration / 9,599 test. 6,171 steps, 2 epochs, 114 minutes on one H100, 26.1 GiB peak. Training samples at most **32 candidates per example** — every candidate carrying target mass, plus random negatives, target renormalized over what survived. Mean candidates per example drops from 68.0 to 25.9. Validation, calibration and test always score the full published set. AdamW (wd 0.01), lr 2e-4 with 5% linear warmup then cosine, grad clip 1.0, bf16 autocast, gradient checkpointing. Checkpoints selected on **validation NLL, not accuracy**. Loss is soft-target cross-entropy, so human label distributions (GoEmotions, ChaosNLI) train against their real distribution rather than an argmax. **Full results.** | | uncalibrated | calibrated (T = 1.2023) | |---|---:|---:| | Accuracy | 0.6379 | 0.6379 | | Target NLL | 1.0687 | **1.0351** | | 10-bin ECE | 0.0837 | **0.0525** | | Brier | 0.3715 | 0.3646 | `test_predictions.json` ships raw per-candidate scores for all 9,599 rows, so every number here is recomputable without a GPU. **A data bug worth knowing about.** jev-bench publishes `null` option descriptions for LEDGAR and GoEmotions in 100% of rows on every split. Every option then rendered as the literal string `Candidate: None`, all options of an example shared one input, and the model returned identical scores by construction — exactly `ln K` loss and `1/K` confidence, untrainable at any budget. The repo falls back to the option's ID and raises if a record's options still all read alike. Legal clause classification went from 0.009 to 0.761 on that fix alone.
## License and credit Apache-2.0. Built on [ModernBERT](https://huggingface.co/answerdotai/ModernBERT-base) (Apache-2.0). Trained on [jev-bench](https://huggingface.co/datasets/Praveenrajus/jev-bench) revision `cbcb6703ef02b698bee32af39d3270fe1f356187`; no dataset content is redistributed here. Source licenses are recorded in `data_audit.json` and include **CC-BY-SA-4.0** for ARC-Challenge and ChaosNLI, CC-BY-4.0 for Banking77, MASSIVE and LEDGAR, CC-BY-3.0 for CLINC150, Apache-2.0 for GoEmotions, MIT for MMLU, and research-use terms for MultiNLI. Check those before commercial use.