Text Classification
Transformers
Safetensors
English
deberta-v2
deberta-v3
schema-conditioned
candidate-scoring
zero-shot-classification
structured-output
text-embeddings-inference
Instructions to use mobarmg/jev-schema-scorer-deberta-v3-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mobarmg/jev-schema-scorer-deberta-v3-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mobarmg/jev-schema-scorer-deberta-v3-large")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("mobarmg/jev-schema-scorer-deberta-v3-large") model = AutoModelForSequenceClassification.from_pretrained("mobarmg/jev-schema-scorer-deberta-v3-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from mobarmg/jev-schema-scorer-deberta-v3-large: direct link, hf CLI and curl.
- Browser
- Download file 7.97 kB
-
https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large/resolve/4e52f50d5590662dc1ef11f273b3523123814a7d/README.md
- Command line
-
hf download hf://mobarmg/jev-schema-scorer-deberta-v3-large@4e52f50d5590662dc1ef11f273b3523123814a7d/README.md
-
curl -L -o README.md https://huggingface.co/mobarmg/jev-schema-scorer-deberta-v3-large/resolve/4e52f50d5590662dc1ef11f273b3523123814a7d/README.md
7.97 kB
| language: en | |
| license: mit | |
| library_name: transformers | |
| base_model: microsoft/deberta-v3-large | |
| pipeline_tag: text-classification | |
| inference: false | |
| tags: | |
| - deberta-v3 | |
| - schema-conditioned | |
| - candidate-scoring | |
| - zero-shot-classification | |
| - structured-output | |
| # Schema-conditioned candidate scorer (DeBERTa-v3-large) | |
| One DeBERTa-v3-large encoder with a **single scalar head** scores `(state, question + candidate)` pairs. | |
| Deterministic code groups the scalar logits per question and decodes them into three answer primitives: | |
| | primitive | input schema | answer | | |
| |---|---|---| | |
| | `choice` | `criteria: {option_id: description}` | argmax option id + probabilities | | |
| | `noul` | optional `criteria: {"true": ..., "false": ...}` | p(proposition is true) | | |
| | `score` | `criteria: [level 0 description, level 1, ...]` | expected level index + per-level probabilities | | |
| The question text, criteria and option ids are **read at inference time**, never baked into the weights, | |
| so the same checkpoint answers new questions over new label sets without retraining. | |
| Try it in the Space: **[mobarmg/jev-schema-scorer](https://huggingface.co/spaces/mobarmg/jev-schema-scorer)**. | |
| ## How it works | |
| For every candidate of a question the model sees a sentence pair: | |
| ``` | |
| sequence_a = the state (free text, or a JSON object serialised) | |
| sequence_b = {"candidate": {"id": "<option id>", "description": "<option description>"}, | |
| "type": "choice", "instructions": "...", "criteria": {...}} | |
| ``` | |
| The candidate sits right after `[SEP]`, so the only tokens that differ between a question's candidates | |
| are where the encoder attends most easily. Each pair yields one logit; a softmax over the question's | |
| candidates gives the answer distribution. Training minimises cross-entropy between that grouped softmax | |
| and a target distribution (one-hot for `choice` / `score`, `[1-p, p]` for `noul`). | |
| ## Usage | |
| The repo ships `schema_scorer.py` with the request compiler, decoder and a small adapter. | |
| ```python | |
| from huggingface_hub import hf_hub_download | |
| import importlib.util | |
| repo = "mobarmg/jev-schema-scorer-deberta-v3-large" | |
| spec = importlib.util.spec_from_file_location("schema_scorer", hf_hub_download(repo, "schema_scorer.py")) | |
| schema_scorer = importlib.util.module_from_spec(spec); spec.loader.exec_module(schema_scorer) | |
| scorer = schema_scorer.LocalSystemOne(repo) | |
| scorer.system_one( | |
| "Nine days of silence on a signed quote is a joke. Our launch event is on the 28th and we still " | |
| "do not have the licence keys your sales team promised. Somebody pick up a phone.", | |
| { | |
| "department": {"type": "choice", "instructions": "Which team should handle this?", | |
| "criteria": {"billing": "Payment or subscription issues", | |
| "technical": "Bugs or integration problems", | |
| "sales": "Pricing or plan questions"}}, | |
| "frustration": {"type": "score", "instructions": "How frustrated the customer appears", | |
| "criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]}, | |
| "is_urgent": {"type": "noul", "instructions": "The message conveys urgency"}, | |
| }, | |
| ) | |
| # {'answers': {'department': {'type': 'choice', 'choice': 'sales', 'probabilities': {...}}, | |
| # 'frustration': {'type': 'score', 'score': 2.0, 'probabilities': [...], 'legend': {...}}, | |
| # 'is_urgent': {'type': 'noul', 'noul': 0.99}}} | |
| ``` | |
| Without the helper, it is a plain `DebertaV2ForSequenceClassification` with `num_labels=1`: tokenize | |
| `(state, serialised question + candidate)` pairs, take `logits[:, 0]`, and softmax over each question's | |
| candidates. | |
| Limits: state + schema + candidate must fit in 512 tokens; every question needs at least two candidates. | |
| ## Training data | |
| Fine-tuned from `microsoft/deberta-v3-large` on 7,650 questions over 3,000 short English texts spanning | |
| 30 domains (support triage, moderation, product reviews, email routing, resume screening, banking, | |
| insurance claims, telehealth messages, IT helpdesk, dating safety, and more). Each domain defines one | |
| `choice`, one `score` and one `noul` question. The texts and labels are synthetic, written to a per-domain | |
| spec. To make the model read the schema instead of memorising label ids, each training example draws a | |
| random instruction wording and criteria wording, shuffles the `choice` options, and replaces the option | |
| ids with opaque ids (`opt_a`, `k2`, `bravo`, ...) half of the time. | |
| Held-out eval split: 1,350 questions (15% of records per domain, stratified over label combinations). | |
| ## Evaluation | |
| Held-out eval split (1,350 questions, 45 per domain), measured with `bench_dataset.py`: | |
| | primitive | metric | DeBERTa-v3-large scorer | chance | | |
| |---|---|---|---| | |
| | `choice` | accuracy | **0.889** | 0.255 | | |
| | `noul` | accuracy | **0.940** | 0.500 | | |
| | `noul` | Brier score (lower is better) | **0.052** | 0.250 | | |
| | `score` | mean absolute error in levels (lower is better) | **0.183** | 0.700 | | |
| <details> | |
| <summary>Per-domain results</summary> | |
| | domain | choice acc (chance) | noul acc / Brier | score MAE (uniform) | | |
| |---|---|---|---| | |
| | auto_service | 0.733 (0.200) | 0.933 / 0.067 | 0.105 (0.667) | | |
| | banking | 0.867 (0.200) | 1.000 / 0.002 | 0.283 (0.600) | | |
| | bug_report | 0.867 (0.250) | 0.933 / 0.067 | 0.097 (0.667) | | |
| | dating_safety | 0.933 (0.250) | 0.933 / 0.067 | 0.176 (0.533) | | |
| | ecommerce_order | 1.000 (0.250) | 1.000 / 0.000 | 0.133 (0.600) | | |
| | education | 1.000 (0.250) | 1.000 / 0.000 | 0.064 (0.600) | | |
| | email_routing | 0.800 (0.250) | 0.933 / 0.061 | 0.189 (0.733) | | |
| | fitness_nutrition | 0.867 (0.250) | 1.000 / 0.000 | 0.024 (0.733) | | |
| | gov_services | 0.800 (0.200) | 1.000 / 0.000 | 0.074 (0.667) | | |
| | health_symptom | 1.000 (0.200) | 1.000 / 0.000 | 0.213 (1.100) | | |
| | hr_workplace | 0.800 (0.200) | 1.000 / 0.000 | 0.041 (0.600) | | |
| | insurance_claim | 1.000 (0.250) | 0.933 / 0.067 | 0.266 (0.733) | | |
| | it_helpdesk | 0.800 (0.250) | 1.000 / 0.000 | 0.430 (0.533) | | |
| | job_posting | 0.933 (0.250) | 0.933 / 0.066 | 0.116 (0.667) | | |
| | legal_clause | 0.933 (0.250) | 0.800 / 0.205 | 0.421 (0.533) | | |
| | mental_health | 0.667 (0.200) | 1.000 / 0.000 | 0.072 (0.667) | | |
| | moderation | 1.000 (0.333) | 0.800 / 0.170 | 0.137 (0.667) | | |
| | news | 1.000 (0.250) | 0.867 / 0.105 | 0.101 (0.667) | | |
| | pharmacy | 0.933 (0.250) | 0.933 / 0.067 | 0.429 (0.600) | | |
| | product_review | 0.867 (0.333) | 0.933 / 0.028 | 0.071 (1.133) | | |
| | real_estate | 0.867 (0.250) | 1.000 / 0.000 | 0.000 (0.533) | | |
| | restaurant_review | 0.933 (0.250) | 0.933 / 0.067 | 0.175 (1.267) | | |
| | resume_screening | 0.933 (0.333) | 0.933 / 0.059 | 0.074 (0.667) | | |
| | scientific_abstract | 1.000 (0.250) | 1.000 / 0.003 | 0.217 (0.733) | | |
| | security_alert | 0.867 (0.200) | 0.800 / 0.134 | 0.814 (0.900) | | |
| | smart_home | 0.933 (0.250) | 0.933 / 0.025 | 0.000 (0.733) | | |
| | social_post | 0.800 (0.333) | 0.867 / 0.114 | 0.310 (0.667) | | |
| | support_triage | 0.933 (0.333) | 0.867 / 0.134 | 0.009 (0.600) | | |
| | survey | 0.800 (0.333) | 0.933 / 0.067 | 0.133 (0.600) | | |
| | travel | 0.800 (0.250) | 1.000 / 0.000 | 0.304 (0.600) | | |
| </details> | |
| Chance is the uniform-guess baseline: `1/k` accuracy for `choice`, 0.5 for `noul`, and the MAE of | |
| predicting the middle level for `score`. | |
| ## Intended use and caveats | |
| - Intended for experiments with schema-conditioned classification: routing, triage, moderation-style | |
| labelling, ordinal scoring, and yes/no propositions over short English texts. | |
| - Trained on synthetic data. Expect degraded accuracy far from the training domains, on long inputs, on | |
| non-English text, and on criteria that require world knowledge or reasoning across the text. | |
| - Probabilities are often very peaked (the grouped softmax is trained on one-hot targets); treat them as | |
| rankings rather than calibrated confidences. | |
| - Not a safety classifier. Do not use its `moderation`, `health`, or `security` outputs to make | |
| consequential decisions without human review. | |