File size: 7,974 Bytes
fd858db
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
---
language: en
license: mit
library_name: transformers
base_model: microsoft/deberta-v3-large
pipeline_tag: text-classification
inference: false
tags:
  - deberta-v3
  - schema-conditioned
  - candidate-scoring
  - zero-shot-classification
  - structured-output
---

# Schema-conditioned candidate scorer (DeBERTa-v3-large)

One DeBERTa-v3-large encoder with a **single scalar head** scores `(state, question + candidate)` pairs.
Deterministic code groups the scalar logits per question and decodes them into three answer primitives:

| primitive | input schema | answer |
|---|---|---|
| `choice` | `criteria: {option_id: description}` | argmax option id + probabilities |
| `noul` | optional `criteria: {"true": ..., "false": ...}` | p(proposition is true) |
| `score` | `criteria: [level 0 description, level 1, ...]` | expected level index + per-level probabilities |

The question text, criteria and option ids are **read at inference time**, never baked into the weights,
so the same checkpoint answers new questions over new label sets without retraining.

Try it in the Space: **[mobarmg/jev-schema-scorer](https://huggingface.co/spaces/mobarmg/jev-schema-scorer)**.

## How it works

For every candidate of a question the model sees a sentence pair:

```
sequence_a = the state (free text, or a JSON object serialised)
sequence_b = {"candidate": {"id": "<option id>", "description": "<option description>"},
              "type": "choice", "instructions": "...", "criteria": {...}}
```

The candidate sits right after `[SEP]`, so the only tokens that differ between a question's candidates
are where the encoder attends most easily. Each pair yields one logit; a softmax over the question's
candidates gives the answer distribution. Training minimises cross-entropy between that grouped softmax
and a target distribution (one-hot for `choice` / `score`, `[1-p, p]` for `noul`).

## Usage

The repo ships `schema_scorer.py` with the request compiler, decoder and a small adapter.

```python
from huggingface_hub import hf_hub_download
import importlib.util

repo = "mobarmg/jev-schema-scorer-deberta-v3-large"
spec = importlib.util.spec_from_file_location("schema_scorer", hf_hub_download(repo, "schema_scorer.py"))
schema_scorer = importlib.util.module_from_spec(spec); spec.loader.exec_module(schema_scorer)

scorer = schema_scorer.LocalSystemOne(repo)
scorer.system_one(
    "Nine days of silence on a signed quote is a joke. Our launch event is on the 28th and we still "
    "do not have the licence keys your sales team promised. Somebody pick up a phone.",
    {
        "department": {"type": "choice", "instructions": "Which team should handle this?",
                       "criteria": {"billing": "Payment or subscription issues",
                                    "technical": "Bugs or integration problems",
                                    "sales": "Pricing or plan questions"}},
        "frustration": {"type": "score", "instructions": "How frustrated the customer appears",
                        "criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]},
        "is_urgent": {"type": "noul", "instructions": "The message conveys urgency"},
    },
)
# {'answers': {'department': {'type': 'choice', 'choice': 'sales', 'probabilities': {...}},
#              'frustration': {'type': 'score', 'score': 2.0, 'probabilities': [...], 'legend': {...}},
#              'is_urgent': {'type': 'noul', 'noul': 0.99}}}
```

Without the helper, it is a plain `DebertaV2ForSequenceClassification` with `num_labels=1`: tokenize
`(state, serialised question + candidate)` pairs, take `logits[:, 0]`, and softmax over each question's
candidates.

Limits: state + schema + candidate must fit in 512 tokens; every question needs at least two candidates.

## Training data

Fine-tuned from `microsoft/deberta-v3-large` on 7,650 questions over 3,000 short English texts spanning
30 domains (support triage, moderation, product reviews, email routing, resume screening, banking,
insurance claims, telehealth messages, IT helpdesk, dating safety, and more). Each domain defines one
`choice`, one `score` and one `noul` question. The texts and labels are synthetic, written to a per-domain
spec. To make the model read the schema instead of memorising label ids, each training example draws a
random instruction wording and criteria wording, shuffles the `choice` options, and replaces the option
ids with opaque ids (`opt_a`, `k2`, `bravo`, ...) half of the time.

Held-out eval split: 1,350 questions (15% of records per domain, stratified over label combinations).

## Evaluation

Held-out eval split (1,350 questions, 45 per domain), measured with `bench_dataset.py`:

| primitive | metric | DeBERTa-v3-large scorer | chance |
|---|---|---|---|
| `choice` | accuracy | **0.889** | 0.255 |
| `noul` | accuracy | **0.940** | 0.500 |
| `noul` | Brier score (lower is better) | **0.052** | 0.250 |
| `score` | mean absolute error in levels (lower is better) | **0.183** | 0.700 |

<details>
<summary>Per-domain results</summary>

| domain | choice acc (chance) | noul acc / Brier | score MAE (uniform) |
|---|---|---|---|
| auto_service | 0.733 (0.200) | 0.933 / 0.067 | 0.105 (0.667) |
| banking | 0.867 (0.200) | 1.000 / 0.002 | 0.283 (0.600) |
| bug_report | 0.867 (0.250) | 0.933 / 0.067 | 0.097 (0.667) |
| dating_safety | 0.933 (0.250) | 0.933 / 0.067 | 0.176 (0.533) |
| ecommerce_order | 1.000 (0.250) | 1.000 / 0.000 | 0.133 (0.600) |
| education | 1.000 (0.250) | 1.000 / 0.000 | 0.064 (0.600) |
| email_routing | 0.800 (0.250) | 0.933 / 0.061 | 0.189 (0.733) |
| fitness_nutrition | 0.867 (0.250) | 1.000 / 0.000 | 0.024 (0.733) |
| gov_services | 0.800 (0.200) | 1.000 / 0.000 | 0.074 (0.667) |
| health_symptom | 1.000 (0.200) | 1.000 / 0.000 | 0.213 (1.100) |
| hr_workplace | 0.800 (0.200) | 1.000 / 0.000 | 0.041 (0.600) |
| insurance_claim | 1.000 (0.250) | 0.933 / 0.067 | 0.266 (0.733) |
| it_helpdesk | 0.800 (0.250) | 1.000 / 0.000 | 0.430 (0.533) |
| job_posting | 0.933 (0.250) | 0.933 / 0.066 | 0.116 (0.667) |
| legal_clause | 0.933 (0.250) | 0.800 / 0.205 | 0.421 (0.533) |
| mental_health | 0.667 (0.200) | 1.000 / 0.000 | 0.072 (0.667) |
| moderation | 1.000 (0.333) | 0.800 / 0.170 | 0.137 (0.667) |
| news | 1.000 (0.250) | 0.867 / 0.105 | 0.101 (0.667) |
| pharmacy | 0.933 (0.250) | 0.933 / 0.067 | 0.429 (0.600) |
| product_review | 0.867 (0.333) | 0.933 / 0.028 | 0.071 (1.133) |
| real_estate | 0.867 (0.250) | 1.000 / 0.000 | 0.000 (0.533) |
| restaurant_review | 0.933 (0.250) | 0.933 / 0.067 | 0.175 (1.267) |
| resume_screening | 0.933 (0.333) | 0.933 / 0.059 | 0.074 (0.667) |
| scientific_abstract | 1.000 (0.250) | 1.000 / 0.003 | 0.217 (0.733) |
| security_alert | 0.867 (0.200) | 0.800 / 0.134 | 0.814 (0.900) |
| smart_home | 0.933 (0.250) | 0.933 / 0.025 | 0.000 (0.733) |
| social_post | 0.800 (0.333) | 0.867 / 0.114 | 0.310 (0.667) |
| support_triage | 0.933 (0.333) | 0.867 / 0.134 | 0.009 (0.600) |
| survey | 0.800 (0.333) | 0.933 / 0.067 | 0.133 (0.600) |
| travel | 0.800 (0.250) | 1.000 / 0.000 | 0.304 (0.600) |

</details>

Chance is the uniform-guess baseline: `1/k` accuracy for `choice`, 0.5 for `noul`, and the MAE of
predicting the middle level for `score`.

## Intended use and caveats

- Intended for experiments with schema-conditioned classification: routing, triage, moderation-style
  labelling, ordinal scoring, and yes/no propositions over short English texts.
- Trained on synthetic data. Expect degraded accuracy far from the training domains, on long inputs, on
  non-English text, and on criteria that require world knowledge or reasoning across the text.
- Probabilities are often very peaked (the grouped softmax is trained on one-hot targets); treat them as
  rankings rather than calibrated confidences.
- Not a safety classifier. Do not use its `moderation`, `health`, or `security` outputs to make
  consequential decisions without human review.