Instructions to use mobarmg/jev-schema-scorer-deberta-v3-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mobarmg/jev-schema-scorer-deberta-v3-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mobarmg/jev-schema-scorer-deberta-v3-large")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("mobarmg/jev-schema-scorer-deberta-v3-large") model = AutoModelForSequenceClassification.from_pretrained("mobarmg/jev-schema-scorer-deberta-v3-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from mobarmg/jev-schema-scorer-deberta-v3-large: direct link, hf CLI and curl.
- Browser
- Download file 7.97 kB
-
https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large/resolve/4e52f50d5590662dc1ef11f273b3523123814a7d/README.md
- Command line
-
hf download hf://mobarmg/jev-schema-scorer-deberta-v3-large@4e52f50d5590662dc1ef11f273b3523123814a7d/README.md
-
curl -L -o README.md https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large/resolve/4e52f50d5590662dc1ef11f273b3523123814a7d/README.md
language: en
license: mit
library_name: transformers
base_model: microsoft/deberta-v3-large
pipeline_tag: text-classification
inference: false
tags:
- deberta-v3
- schema-conditioned
- candidate-scoring
- zero-shot-classification
- structured-output
Schema-conditioned candidate scorer (DeBERTa-v3-large)
One DeBERTa-v3-large encoder with a single scalar head scores (state, question + candidate) pairs.
Deterministic code groups the scalar logits per question and decodes them into three answer primitives:
| primitive | input schema | answer |
|---|---|---|
choice |
criteria: {option_id: description} |
argmax option id + probabilities |
noul |
optional criteria: {"true": ..., "false": ...} |
p(proposition is true) |
score |
criteria: [level 0 description, level 1, ...] |
expected level index + per-level probabilities |
The question text, criteria and option ids are read at inference time, never baked into the weights, so the same checkpoint answers new questions over new label sets without retraining.
Try it in the Space: mobarmg/jev-schema-scorer.
How it works
For every candidate of a question the model sees a sentence pair:
sequence_a = the state (free text, or a JSON object serialised)
sequence_b = {"candidate": {"id": "<option id>", "description": "<option description>"},
"type": "choice", "instructions": "...", "criteria": {...}}
The candidate sits right after [SEP], so the only tokens that differ between a question's candidates
are where the encoder attends most easily. Each pair yields one logit; a softmax over the question's
candidates gives the answer distribution. Training minimises cross-entropy between that grouped softmax
and a target distribution (one-hot for choice / score, [1-p, p] for noul).
Usage
The repo ships schema_scorer.py with the request compiler, decoder and a small adapter.
from huggingface_hub import hf_hub_download
import importlib.util
repo = "mobarmg/jev-schema-scorer-deberta-v3-large"
spec = importlib.util.spec_from_file_location("schema_scorer", hf_hub_download(repo, "schema_scorer.py"))
schema_scorer = importlib.util.module_from_spec(spec); spec.loader.exec_module(schema_scorer)
scorer = schema_scorer.LocalSystemOne(repo)
scorer.system_one(
"Nine days of silence on a signed quote is a joke. Our launch event is on the 28th and we still "
"do not have the licence keys your sales team promised. Somebody pick up a phone.",
{
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or plan questions"}},
"frustration": {"type": "score", "instructions": "How frustrated the customer appears",
"criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]},
"is_urgent": {"type": "noul", "instructions": "The message conveys urgency"},
},
)
# {'answers': {'department': {'type': 'choice', 'choice': 'sales', 'probabilities': {...}},
# 'frustration': {'type': 'score', 'score': 2.0, 'probabilities': [...], 'legend': {...}},
# 'is_urgent': {'type': 'noul', 'noul': 0.99}}}
Without the helper, it is a plain DebertaV2ForSequenceClassification with num_labels=1: tokenize
(state, serialised question + candidate) pairs, take logits[:, 0], and softmax over each question's
candidates.
Limits: state + schema + candidate must fit in 512 tokens; every question needs at least two candidates.
Training data
Fine-tuned from microsoft/deberta-v3-large on 7,650 questions over 3,000 short English texts spanning
30 domains (support triage, moderation, product reviews, email routing, resume screening, banking,
insurance claims, telehealth messages, IT helpdesk, dating safety, and more). Each domain defines one
choice, one score and one noul question. The texts and labels are synthetic, written to a per-domain
spec. To make the model read the schema instead of memorising label ids, each training example draws a
random instruction wording and criteria wording, shuffles the choice options, and replaces the option
ids with opaque ids (opt_a, k2, bravo, ...) half of the time.
Held-out eval split: 1,350 questions (15% of records per domain, stratified over label combinations).
Evaluation
Held-out eval split (1,350 questions, 45 per domain), measured with bench_dataset.py:
| primitive | metric | DeBERTa-v3-large scorer | chance |
|---|---|---|---|
choice |
accuracy | 0.889 | 0.255 |
noul |
accuracy | 0.940 | 0.500 |
noul |
Brier score (lower is better) | 0.052 | 0.250 |
score |
mean absolute error in levels (lower is better) | 0.183 | 0.700 |
Per-domain results
| domain | choice acc (chance) | noul acc / Brier | score MAE (uniform) |
|---|---|---|---|
| auto_service | 0.733 (0.200) | 0.933 / 0.067 | 0.105 (0.667) |
| banking | 0.867 (0.200) | 1.000 / 0.002 | 0.283 (0.600) |
| bug_report | 0.867 (0.250) | 0.933 / 0.067 | 0.097 (0.667) |
| dating_safety | 0.933 (0.250) | 0.933 / 0.067 | 0.176 (0.533) |
| ecommerce_order | 1.000 (0.250) | 1.000 / 0.000 | 0.133 (0.600) |
| education | 1.000 (0.250) | 1.000 / 0.000 | 0.064 (0.600) |
| email_routing | 0.800 (0.250) | 0.933 / 0.061 | 0.189 (0.733) |
| fitness_nutrition | 0.867 (0.250) | 1.000 / 0.000 | 0.024 (0.733) |
| gov_services | 0.800 (0.200) | 1.000 / 0.000 | 0.074 (0.667) |
| health_symptom | 1.000 (0.200) | 1.000 / 0.000 | 0.213 (1.100) |
| hr_workplace | 0.800 (0.200) | 1.000 / 0.000 | 0.041 (0.600) |
| insurance_claim | 1.000 (0.250) | 0.933 / 0.067 | 0.266 (0.733) |
| it_helpdesk | 0.800 (0.250) | 1.000 / 0.000 | 0.430 (0.533) |
| job_posting | 0.933 (0.250) | 0.933 / 0.066 | 0.116 (0.667) |
| legal_clause | 0.933 (0.250) | 0.800 / 0.205 | 0.421 (0.533) |
| mental_health | 0.667 (0.200) | 1.000 / 0.000 | 0.072 (0.667) |
| moderation | 1.000 (0.333) | 0.800 / 0.170 | 0.137 (0.667) |
| news | 1.000 (0.250) | 0.867 / 0.105 | 0.101 (0.667) |
| pharmacy | 0.933 (0.250) | 0.933 / 0.067 | 0.429 (0.600) |
| product_review | 0.867 (0.333) | 0.933 / 0.028 | 0.071 (1.133) |
| real_estate | 0.867 (0.250) | 1.000 / 0.000 | 0.000 (0.533) |
| restaurant_review | 0.933 (0.250) | 0.933 / 0.067 | 0.175 (1.267) |
| resume_screening | 0.933 (0.333) | 0.933 / 0.059 | 0.074 (0.667) |
| scientific_abstract | 1.000 (0.250) | 1.000 / 0.003 | 0.217 (0.733) |
| security_alert | 0.867 (0.200) | 0.800 / 0.134 | 0.814 (0.900) |
| smart_home | 0.933 (0.250) | 0.933 / 0.025 | 0.000 (0.733) |
| social_post | 0.800 (0.333) | 0.867 / 0.114 | 0.310 (0.667) |
| support_triage | 0.933 (0.333) | 0.867 / 0.134 | 0.009 (0.600) |
| survey | 0.800 (0.333) | 0.933 / 0.067 | 0.133 (0.600) |
| travel | 0.800 (0.250) | 1.000 / 0.000 | 0.304 (0.600) |
Chance is the uniform-guess baseline: 1/k accuracy for choice, 0.5 for noul, and the MAE of
predicting the middle level for score.
Intended use and caveats
- Intended for experiments with schema-conditioned classification: routing, triage, moderation-style labelling, ordinal scoring, and yes/no propositions over short English texts.
- Trained on synthetic data. Expect degraded accuracy far from the training domains, on long inputs, on non-English text, and on criteria that require world knowledge or reasoning across the text.
- Probabilities are often very peaked (the grouped softmax is trained on one-hot targets); treat them as rankings rather than calibrated confidences.
- Not a safety classifier. Do not use its
moderation,health, orsecurityoutputs to make consequential decisions without human review.