stuntd support triage heads

Three decision heads that triage a support ticket in one request: which part of the product it is about, how soon it needs an answer, and whether a person has to read it. They were trained by stuntd on top of the frozen Laya encoder, and each one is about 50 MB.

This is stuntd's support demo: the tickets are generated and the teacher is a rule, so take it as a working example of what a head does, not as a production triage model. stuntd trains the same heads on your own traffic, with your own LLM as the teacher.

head answers
category billing, bug, feature, account, other
urgency 0 (can wait) to 3 (an outage, or stuck with a deadline)
needs_human true / false

Results

The tickets come from 20 templates, and a label depends only on the template and the sentences added to it, so the numbers depend on which tickets you test on. Each head answers on its own when its confidence is over the threshold stuntd picked for 0.99 agreement with the teacher; a ticket skips the big model only when all three heads are sure.

These heads were trained with stuntd 0.1.3 and carry the vectors of their training rows (embeddings.safetensors), so its novelty gate works on them: a request far from everything the heads trained on goes to the provider however confident the heads look.

1,000 test tickets answered by all three heads all three right when they answer
new states of the 20 trained templates, gate off 72.3% 97.2%
new states of the 20 trained templates, gate on 66.9% 97.6%
templates left out of training, gate off 40.0% 61.0%
templates left out of training, gate on 0% (all sent to the provider)

"What's the weather in Paris", asdf qwer zxcv and an empty state are stopped by the gate; without it the heads answered the weather question at confidence 1.00. Changing only channel and plan, which no rule reads, changes the answer of category on 2.1% of the tickets, needs_human on 3.5% and urgency on 10.2%.

The left-out rows come from heads trained on the other 15 templates, the others from these heads. Read the first rows as new wording of known tickets and the last rows as what happens on traffic that looks different from the training data.

Through a running stuntd serve, all three answers for one ticket come back at a p50 of 83 ms (200 tickets, RTX 5060 laptop). One head on CPU takes about 60 ms.

Both checks were suggested by Dipankar Sarkar in the discussion of these heads, who also found that the stuck-payout template tripped the "money back" refund cue. The wording is fixed and these heads are trained on the corrected rows.

Use them

pip install "stuntd[train]>=0.1.3"
hf download pollix/stuntd-support-triage --local-dir support-heads

Point stuntd at the folder in stuntd.toml (use the absolute path):

[training]
base_model = "convaiinnovations/laya"

[storage]
models_dir = "/absolute/path/to/support-heads"
stuntd enable category
stuntd enable urgency
stuntd enable needs_human
stuntd serve

Then ask over the Jev protocol (POST /v1/systemone) with three questions named category, urgency and needs_human. The state is a flat object:

{"channel": "email", "plan": "pro",
 "subject": "The renewal payment failed",
 "body": "Our card was declined at renewal and the team lost access this morning."}

examples/support/client.py sends tickets and prints the answers, the latency and the share answered by the heads.

Files

Each folder holds one head: head.safetensors (the trained head weights, fp16), embeddings.safetensors (the vectors of its 2,400 training rows, for the novelty gate) and meta.json (labels, the confidence threshold, the novelty cut-off, the temperature and the holdout numbers stuntd reports).

Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pollix/stuntd-support-triage

Adapter
(14)
this model

Collection including pollix/stuntd-support-triage