vrajnotviraj commited on
Commit
a00287d
·
verified ·
1 Parent(s): d228ae1

Single-file ONNX (fixes from_pretrained), better model card

Browse files
Files changed (3) hide show
  1. README.md +37 -6
  2. laya.onnx +2 -2
  3. laya.onnx.data +0 -3
README.md CHANGED
@@ -5,6 +5,7 @@ language:
5
  library_name: onnx
6
  pipeline_tag: zero-shot-classification
7
  base_model: jhu-clsp/ettin-encoder-150m
 
8
  tags:
9
  - intent-classification
10
  - intent-detection
@@ -22,6 +23,19 @@ tags:
22
  - ettin
23
  - laya
24
  - distillation
 
 
 
 
 
 
 
 
 
 
 
 
 
25
  datasets:
26
  - clinc_oos
27
  - PolyAI/banking77
@@ -48,13 +62,21 @@ model-index:
48
  - type: recall
49
  value: 0.986
50
  name: Out-of-scope recall
 
 
 
 
 
 
51
  ---
52
 
53
- # Laya Intent Router 150M (ONNX, int8)
 
 
54
 
55
- **Zero-shot intent classification that runs on a CPU in about 60 ms.** You give it a message and a list of intents written in plain English. It tells you which intent the message belongs to, or that it belongs to none of them.
56
 
57
- No training. No labelled data. You write the intents when you call it, and you can change them on every request.
58
 
59
  ```
60
  "i want to cancel my order 88213" -> cancel_order 0.97
@@ -66,6 +88,15 @@ No training. No labelled data. You write the intents when you call it, and you c
66
 
67
  That last line is why I built this. The original Laya sent `qwewqeqw` to `order_not_received` with 0.73 confidence. This one puts 0.98 on none.
68
 
 
 
 
 
 
 
 
 
 
69
  ## Try it
70
 
71
  ```bash
@@ -113,7 +144,7 @@ A SaaS support bot, 4 intents, nothing fine-tuned:
113
 
114
  The intents were `reset_password: "User can't log in or forgot their password"`, `billing: "Questions about invoices, charges or refunds"`, `bug_report: "Something in the app is broken or showing an error"` and `talk_to_human: "User wants to speak to a real person"`. That's the whole setup.
115
 
116
- ## How good is it
117
 
118
  I tested it on 1,636 hand-written messages across 10 routing setups: e-commerce, retail banking, insurance and telecom, plus an adversarial set of near-duplicates, typos, slang and "don't cancel, just tell me where it is" style traps. **Insurance and telecom were never seen in training.**
119
 
@@ -132,7 +163,7 @@ So you get the big fine-tuned model's accuracy at a third of its latency and hal
132
 
133
  It also holds up when I poke at it. A perturbation test (shuffled intent order, removed correct intent, distractor intents, rewritten messages, gibberish) scores 0.92 averaged over 3 seeds, against 0.73 for the original Laya.
134
 
135
- ## Big intent lists
136
 
137
  The model reads everything in one 512-token window, so it can only see so many intents at once. For lists longer than 4, a small embedder (`bge-small-en-v1.5`, bundled in `shortlist/`) picks the 4 closest intents first and the router decides between those and "none".
138
 
@@ -170,7 +201,7 @@ Everything trained locally on a 32 GB M2 Pro. Training code: [github.com/vrajnot
170
 
171
  | file | what |
172
  |---|---|
173
- | `laya.onnx`, `laya.onnx.data` | the router (int8 weights) |
174
  | `tokenizer.json`, `rl_agent_config.json` | tokenizer, prompt format, calibrated temperatures, threshold |
175
  | `shortlist/` | bge-small-en-v1.5 ONNX embedder for long intent lists |
176
  | `router.py` | the whole inference code, one file, no torch |
 
5
  library_name: onnx
6
  pipeline_tag: zero-shot-classification
7
  base_model: jhu-clsp/ettin-encoder-150m
8
+ base_model_relation: finetune
9
  tags:
10
  - intent-classification
11
  - intent-detection
 
23
  - ettin
24
  - laya
25
  - distillation
26
+ - knowledge-distillation
27
+ - onnxruntime
28
+ - zero-shot
29
+ - nlu
30
+ - intent-router
31
+ - semantic-router
32
+ - llm-router
33
+ - open-intent-detection
34
+ - out-of-distribution-detection
35
+ - customer-service
36
+ - banking
37
+ - e-commerce
38
+ - edge
39
  datasets:
40
  - clinc_oos
41
  - PolyAI/banking77
 
62
  - type: recall
63
  value: 0.986
64
  name: Out-of-scope recall
65
+ - type: accuracy
66
+ value: 0.943
67
+ name: Routing accuracy, menus of 20 to 45 intents
68
+ - type: accuracy
69
+ value: 0.92
70
+ name: Routing accuracy, 6 unseen domains
71
  ---
72
 
73
+ # Laya Intent Router 150M: zero-shot intent classification on CPU (ONNX, int8)
74
+
75
+ **A 150M-parameter zero-shot intent detection model with out-of-scope detection, 160 ms p95 on 4 CPU threads.** You give it a user message and a list of intents written in plain English. It tells you which intent the message belongs to, or that it belongs to none of them.
76
 
77
+ No training. No labelled data. You write the intents when you call it, and you can change them on every request. It's a small ONNX file you run with `onnxruntime` and `numpy`, no GPU and no torch.
78
 
79
+ On a held-out set of 1,636 messages it gets **94.9% routing accuracy and 98.6% out-of-scope recall**, up from 72.8% and 68.2% for the original [Laya](https://huggingface.co/convaiinnovations/laya).
80
 
81
  ```
82
  "i want to cancel my order 88213" -> cancel_order 0.97
 
88
 
89
  That last line is why I built this. The original Laya sent `qwewqeqw` to `order_not_received` with 0.73 confidence. This one puts 0.98 on none.
90
 
91
+ ## Use it for
92
+
93
+ - **Chatbot and voicebot intent routing**, where the menu of intents changes per flow or per customer.
94
+ - **Out-of-scope detection**: knowing when a message fits none of your intents, so you can fall back to a human, an LLM or an "I didn't get that".
95
+ - **A cheap semantic router in front of an LLM**, when you want a decision in 150 ms on CPU instead of a full LLM call.
96
+ - **Customer support triage** for e-commerce, banking, insurance, telecom and SaaS.
97
+
98
+ It's built for short English customer messages. It isn't a general text classifier for long documents, and it only speaks English.
99
+
100
  ## Try it
101
 
102
  ```bash
 
144
 
145
  The intents were `reset_password: "User can't log in or forgot their password"`, `billing: "Questions about invoices, charges or refunds"`, `bug_report: "Something in the app is broken or showing an error"` and `talk_to_human: "User wants to speak to a real person"`. That's the whole setup.
146
 
147
+ ## Benchmarks: 94.9% routing accuracy, 98.6% out-of-scope recall
148
 
149
  I tested it on 1,636 hand-written messages across 10 routing setups: e-commerce, retail banking, insurance and telecom, plus an adversarial set of near-duplicates, typos, slang and "don't cancel, just tell me where it is" style traps. **Insurance and telecom were never seen in training.**
150
 
 
163
 
164
  It also holds up when I poke at it. A perturbation test (shuffled intent order, removed correct intent, distractor intents, rewritten messages, gibberish) scores 0.92 averaged over 3 seeds, against 0.73 for the original Laya.
165
 
166
+ ## Big intent lists (up to 148 intents)
167
 
168
  The model reads everything in one 512-token window, so it can only see so many intents at once. For lists longer than 4, a small embedder (`bge-small-en-v1.5`, bundled in `shortlist/`) picks the 4 closest intents first and the router decides between those and "none".
169
 
 
201
 
202
  | file | what |
203
  |---|---|
204
+ | `laya.onnx` | the router, one file with int8 weights (302 MB) |
205
  | `tokenizer.json`, `rl_agent_config.json` | tokenizer, prompt format, calibrated temperatures, threshold |
206
  | `shortlist/` | bge-small-en-v1.5 ONNX embedder for long intent lists |
207
  | `router.py` | the whole inference code, one file, no torch |
laya.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:53e526532f19fff929506f3bffbd4faa0f6fdb684fa5842a2376e973e4091808
3
- size 2467583
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f4e29e431074a059977293de52408c4e18ce7f6e38d6b7b77e7e3e9406d67f34
3
+ size 302276803
laya.onnx.data DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:17cfafb27b977cd0537a61da24edecad8d8f50eac2d0eb9f3e14e5bc3116cf9c
3
- size 299825152