mobarmg's picture
Upload scorer
fd858db verified
|
Raw History Blame
7.97 kB
---
language: en
license: mit
library_name: transformers
base_model: microsoft/deberta-v3-large
pipeline_tag: text-classification
inference: false
tags:
- deberta-v3
- schema-conditioned
- candidate-scoring
- zero-shot-classification
- structured-output
---
# Schema-conditioned candidate scorer (DeBERTa-v3-large)
One DeBERTa-v3-large encoder with a **single scalar head** scores `(state, question + candidate)` pairs.
Deterministic code groups the scalar logits per question and decodes them into three answer primitives:
| primitive | input schema | answer |
|---|---|---|
| `choice` | `criteria: {option_id: description}` | argmax option id + probabilities |
| `noul` | optional `criteria: {"true": ..., "false": ...}` | p(proposition is true) |
| `score` | `criteria: [level 0 description, level 1, ...]` | expected level index + per-level probabilities |
The question text, criteria and option ids are **read at inference time**, never baked into the weights,
so the same checkpoint answers new questions over new label sets without retraining.
Try it in the Space: **[mobarmg/jev-schema-scorer](https://huggingface.co/spaces/mobarmg/jev-schema-scorer)**.
## How it works
For every candidate of a question the model sees a sentence pair:
```
sequence_a = the state (free text, or a JSON object serialised)
sequence_b = {"candidate": {"id": "<option id>", "description": "<option description>"},
"type": "choice", "instructions": "...", "criteria": {...}}
```
The candidate sits right after `[SEP]`, so the only tokens that differ between a question's candidates
are where the encoder attends most easily. Each pair yields one logit; a softmax over the question's
candidates gives the answer distribution. Training minimises cross-entropy between that grouped softmax
and a target distribution (one-hot for `choice` / `score`, `[1-p, p]` for `noul`).
## Usage
The repo ships `schema_scorer.py` with the request compiler, decoder and a small adapter.
```python
from huggingface_hub import hf_hub_download
import importlib.util
repo = "mobarmg/jev-schema-scorer-deberta-v3-large"
spec = importlib.util.spec_from_file_location("schema_scorer", hf_hub_download(repo, "schema_scorer.py"))
schema_scorer = importlib.util.module_from_spec(spec); spec.loader.exec_module(schema_scorer)
scorer = schema_scorer.LocalSystemOne(repo)
scorer.system_one(
"Nine days of silence on a signed quote is a joke. Our launch event is on the 28th and we still "
"do not have the licence keys your sales team promised. Somebody pick up a phone.",
{
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or plan questions"}},
"frustration": {"type": "score", "instructions": "How frustrated the customer appears",
"criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]},
"is_urgent": {"type": "noul", "instructions": "The message conveys urgency"},
},
)
# {'answers': {'department': {'type': 'choice', 'choice': 'sales', 'probabilities': {...}},
# 'frustration': {'type': 'score', 'score': 2.0, 'probabilities': [...], 'legend': {...}},
# 'is_urgent': {'type': 'noul', 'noul': 0.99}}}
```
Without the helper, it is a plain `DebertaV2ForSequenceClassification` with `num_labels=1`: tokenize
`(state, serialised question + candidate)` pairs, take `logits[:, 0]`, and softmax over each question's
candidates.
Limits: state + schema + candidate must fit in 512 tokens; every question needs at least two candidates.
## Training data
Fine-tuned from `microsoft/deberta-v3-large` on 7,650 questions over 3,000 short English texts spanning
30 domains (support triage, moderation, product reviews, email routing, resume screening, banking,
insurance claims, telehealth messages, IT helpdesk, dating safety, and more). Each domain defines one
`choice`, one `score` and one `noul` question. The texts and labels are synthetic, written to a per-domain
spec. To make the model read the schema instead of memorising label ids, each training example draws a
random instruction wording and criteria wording, shuffles the `choice` options, and replaces the option
ids with opaque ids (`opt_a`, `k2`, `bravo`, ...) half of the time.
Held-out eval split: 1,350 questions (15% of records per domain, stratified over label combinations).
## Evaluation
Held-out eval split (1,350 questions, 45 per domain), measured with `bench_dataset.py`:
| primitive | metric | DeBERTa-v3-large scorer | chance |
|---|---|---|---|
| `choice` | accuracy | **0.889** | 0.255 |
| `noul` | accuracy | **0.940** | 0.500 |
| `noul` | Brier score (lower is better) | **0.052** | 0.250 |
| `score` | mean absolute error in levels (lower is better) | **0.183** | 0.700 |
<details>
<summary>Per-domain results</summary>
| domain | choice acc (chance) | noul acc / Brier | score MAE (uniform) |
|---|---|---|---|
| auto_service | 0.733 (0.200) | 0.933 / 0.067 | 0.105 (0.667) |
| banking | 0.867 (0.200) | 1.000 / 0.002 | 0.283 (0.600) |
| bug_report | 0.867 (0.250) | 0.933 / 0.067 | 0.097 (0.667) |
| dating_safety | 0.933 (0.250) | 0.933 / 0.067 | 0.176 (0.533) |
| ecommerce_order | 1.000 (0.250) | 1.000 / 0.000 | 0.133 (0.600) |
| education | 1.000 (0.250) | 1.000 / 0.000 | 0.064 (0.600) |
| email_routing | 0.800 (0.250) | 0.933 / 0.061 | 0.189 (0.733) |
| fitness_nutrition | 0.867 (0.250) | 1.000 / 0.000 | 0.024 (0.733) |
| gov_services | 0.800 (0.200) | 1.000 / 0.000 | 0.074 (0.667) |
| health_symptom | 1.000 (0.200) | 1.000 / 0.000 | 0.213 (1.100) |
| hr_workplace | 0.800 (0.200) | 1.000 / 0.000 | 0.041 (0.600) |
| insurance_claim | 1.000 (0.250) | 0.933 / 0.067 | 0.266 (0.733) |
| it_helpdesk | 0.800 (0.250) | 1.000 / 0.000 | 0.430 (0.533) |
| job_posting | 0.933 (0.250) | 0.933 / 0.066 | 0.116 (0.667) |
| legal_clause | 0.933 (0.250) | 0.800 / 0.205 | 0.421 (0.533) |
| mental_health | 0.667 (0.200) | 1.000 / 0.000 | 0.072 (0.667) |
| moderation | 1.000 (0.333) | 0.800 / 0.170 | 0.137 (0.667) |
| news | 1.000 (0.250) | 0.867 / 0.105 | 0.101 (0.667) |
| pharmacy | 0.933 (0.250) | 0.933 / 0.067 | 0.429 (0.600) |
| product_review | 0.867 (0.333) | 0.933 / 0.028 | 0.071 (1.133) |
| real_estate | 0.867 (0.250) | 1.000 / 0.000 | 0.000 (0.533) |
| restaurant_review | 0.933 (0.250) | 0.933 / 0.067 | 0.175 (1.267) |
| resume_screening | 0.933 (0.333) | 0.933 / 0.059 | 0.074 (0.667) |
| scientific_abstract | 1.000 (0.250) | 1.000 / 0.003 | 0.217 (0.733) |
| security_alert | 0.867 (0.200) | 0.800 / 0.134 | 0.814 (0.900) |
| smart_home | 0.933 (0.250) | 0.933 / 0.025 | 0.000 (0.733) |
| social_post | 0.800 (0.333) | 0.867 / 0.114 | 0.310 (0.667) |
| support_triage | 0.933 (0.333) | 0.867 / 0.134 | 0.009 (0.600) |
| survey | 0.800 (0.333) | 0.933 / 0.067 | 0.133 (0.600) |
| travel | 0.800 (0.250) | 1.000 / 0.000 | 0.304 (0.600) |
</details>
Chance is the uniform-guess baseline: `1/k` accuracy for `choice`, 0.5 for `noul`, and the MAE of
predicting the middle level for `score`.
## Intended use and caveats
- Intended for experiments with schema-conditioned classification: routing, triage, moderation-style
labelling, ordinal scoring, and yes/no propositions over short English texts.
- Trained on synthetic data. Expect degraded accuracy far from the training domains, on long inputs, on
non-English text, and on criteria that require world knowledge or reasoning across the text.
- Probabilities are often very peaked (the grouped softmax is trained on one-hot targets); treat them as
rankings rather than calibrated confidences.
- Not a safety classifier. Do not use its `moderation`, `health`, or `security` outputs to make
consequential decisions without human review.