Zero-Shot Classification
ONNX
Safetensors
Laya
MLX
English
intent-classification
intent-detection
text-classification
chatbot
conversational-ai
customer-support
routing
out-of-scope-detection
int8
cpu
modernbert
ettin
distillation
knowledge-distillation
onnxruntime
apple-silicon
macos
metal
on-device
zero-shot
nlu
intent-router
semantic-router
llm-router
open-intent-detection
out-of-distribution-detection
customer-service
banking
e-commerce
edge
Eval Results (legacy)
Instructions to use vrajnotviraj/laya-intent-router-150m-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use vrajnotviraj/laya-intent-router-150m-onnx with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- MLX
How to use vrajnotviraj/laya-intent-router-150m-onnx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download vrajnotviraj/laya-intent-router-150m-onnx --local-dir laya-intent-router-150m-onnx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Single-file ONNX (fixes from_pretrained), better model card
Browse files- README.md +37 -6
- laya.onnx +2 -2
- laya.onnx.data +0 -3
README.md
CHANGED
|
@@ -5,6 +5,7 @@ language:
|
|
| 5 |
library_name: onnx
|
| 6 |
pipeline_tag: zero-shot-classification
|
| 7 |
base_model: jhu-clsp/ettin-encoder-150m
|
|
|
|
| 8 |
tags:
|
| 9 |
- intent-classification
|
| 10 |
- intent-detection
|
|
@@ -22,6 +23,19 @@ tags:
|
|
| 22 |
- ettin
|
| 23 |
- laya
|
| 24 |
- distillation
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
datasets:
|
| 26 |
- clinc_oos
|
| 27 |
- PolyAI/banking77
|
|
@@ -48,13 +62,21 @@ model-index:
|
|
| 48 |
- type: recall
|
| 49 |
value: 0.986
|
| 50 |
name: Out-of-scope recall
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 51 |
---
|
| 52 |
|
| 53 |
-
# Laya Intent Router 150M (ONNX, int8)
|
|
|
|
|
|
|
| 54 |
|
| 55 |
-
|
| 56 |
|
| 57 |
-
|
| 58 |
|
| 59 |
```
|
| 60 |
"i want to cancel my order 88213" -> cancel_order 0.97
|
|
@@ -66,6 +88,15 @@ No training. No labelled data. You write the intents when you call it, and you c
|
|
| 66 |
|
| 67 |
That last line is why I built this. The original Laya sent `qwewqeqw` to `order_not_received` with 0.73 confidence. This one puts 0.98 on none.
|
| 68 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
## Try it
|
| 70 |
|
| 71 |
```bash
|
|
@@ -113,7 +144,7 @@ A SaaS support bot, 4 intents, nothing fine-tuned:
|
|
| 113 |
|
| 114 |
The intents were `reset_password: "User can't log in or forgot their password"`, `billing: "Questions about invoices, charges or refunds"`, `bug_report: "Something in the app is broken or showing an error"` and `talk_to_human: "User wants to speak to a real person"`. That's the whole setup.
|
| 115 |
|
| 116 |
-
##
|
| 117 |
|
| 118 |
I tested it on 1,636 hand-written messages across 10 routing setups: e-commerce, retail banking, insurance and telecom, plus an adversarial set of near-duplicates, typos, slang and "don't cancel, just tell me where it is" style traps. **Insurance and telecom were never seen in training.**
|
| 119 |
|
|
@@ -132,7 +163,7 @@ So you get the big fine-tuned model's accuracy at a third of its latency and hal
|
|
| 132 |
|
| 133 |
It also holds up when I poke at it. A perturbation test (shuffled intent order, removed correct intent, distractor intents, rewritten messages, gibberish) scores 0.92 averaged over 3 seeds, against 0.73 for the original Laya.
|
| 134 |
|
| 135 |
-
## Big intent lists
|
| 136 |
|
| 137 |
The model reads everything in one 512-token window, so it can only see so many intents at once. For lists longer than 4, a small embedder (`bge-small-en-v1.5`, bundled in `shortlist/`) picks the 4 closest intents first and the router decides between those and "none".
|
| 138 |
|
|
@@ -170,7 +201,7 @@ Everything trained locally on a 32 GB M2 Pro. Training code: [github.com/vrajnot
|
|
| 170 |
|
| 171 |
| file | what |
|
| 172 |
|---|---|
|
| 173 |
-
| `laya.onnx`
|
| 174 |
| `tokenizer.json`, `rl_agent_config.json` | tokenizer, prompt format, calibrated temperatures, threshold |
|
| 175 |
| `shortlist/` | bge-small-en-v1.5 ONNX embedder for long intent lists |
|
| 176 |
| `router.py` | the whole inference code, one file, no torch |
|
|
|
|
| 5 |
library_name: onnx
|
| 6 |
pipeline_tag: zero-shot-classification
|
| 7 |
base_model: jhu-clsp/ettin-encoder-150m
|
| 8 |
+
base_model_relation: finetune
|
| 9 |
tags:
|
| 10 |
- intent-classification
|
| 11 |
- intent-detection
|
|
|
|
| 23 |
- ettin
|
| 24 |
- laya
|
| 25 |
- distillation
|
| 26 |
+
- knowledge-distillation
|
| 27 |
+
- onnxruntime
|
| 28 |
+
- zero-shot
|
| 29 |
+
- nlu
|
| 30 |
+
- intent-router
|
| 31 |
+
- semantic-router
|
| 32 |
+
- llm-router
|
| 33 |
+
- open-intent-detection
|
| 34 |
+
- out-of-distribution-detection
|
| 35 |
+
- customer-service
|
| 36 |
+
- banking
|
| 37 |
+
- e-commerce
|
| 38 |
+
- edge
|
| 39 |
datasets:
|
| 40 |
- clinc_oos
|
| 41 |
- PolyAI/banking77
|
|
|
|
| 62 |
- type: recall
|
| 63 |
value: 0.986
|
| 64 |
name: Out-of-scope recall
|
| 65 |
+
- type: accuracy
|
| 66 |
+
value: 0.943
|
| 67 |
+
name: Routing accuracy, menus of 20 to 45 intents
|
| 68 |
+
- type: accuracy
|
| 69 |
+
value: 0.92
|
| 70 |
+
name: Routing accuracy, 6 unseen domains
|
| 71 |
---
|
| 72 |
|
| 73 |
+
# Laya Intent Router 150M: zero-shot intent classification on CPU (ONNX, int8)
|
| 74 |
+
|
| 75 |
+
**A 150M-parameter zero-shot intent detection model with out-of-scope detection, 160 ms p95 on 4 CPU threads.** You give it a user message and a list of intents written in plain English. It tells you which intent the message belongs to, or that it belongs to none of them.
|
| 76 |
|
| 77 |
+
No training. No labelled data. You write the intents when you call it, and you can change them on every request. It's a small ONNX file you run with `onnxruntime` and `numpy`, no GPU and no torch.
|
| 78 |
|
| 79 |
+
On a held-out set of 1,636 messages it gets **94.9% routing accuracy and 98.6% out-of-scope recall**, up from 72.8% and 68.2% for the original [Laya](https://huggingface.co/convaiinnovations/laya).
|
| 80 |
|
| 81 |
```
|
| 82 |
"i want to cancel my order 88213" -> cancel_order 0.97
|
|
|
|
| 88 |
|
| 89 |
That last line is why I built this. The original Laya sent `qwewqeqw` to `order_not_received` with 0.73 confidence. This one puts 0.98 on none.
|
| 90 |
|
| 91 |
+
## Use it for
|
| 92 |
+
|
| 93 |
+
- **Chatbot and voicebot intent routing**, where the menu of intents changes per flow or per customer.
|
| 94 |
+
- **Out-of-scope detection**: knowing when a message fits none of your intents, so you can fall back to a human, an LLM or an "I didn't get that".
|
| 95 |
+
- **A cheap semantic router in front of an LLM**, when you want a decision in 150 ms on CPU instead of a full LLM call.
|
| 96 |
+
- **Customer support triage** for e-commerce, banking, insurance, telecom and SaaS.
|
| 97 |
+
|
| 98 |
+
It's built for short English customer messages. It isn't a general text classifier for long documents, and it only speaks English.
|
| 99 |
+
|
| 100 |
## Try it
|
| 101 |
|
| 102 |
```bash
|
|
|
|
| 144 |
|
| 145 |
The intents were `reset_password: "User can't log in or forgot their password"`, `billing: "Questions about invoices, charges or refunds"`, `bug_report: "Something in the app is broken or showing an error"` and `talk_to_human: "User wants to speak to a real person"`. That's the whole setup.
|
| 146 |
|
| 147 |
+
## Benchmarks: 94.9% routing accuracy, 98.6% out-of-scope recall
|
| 148 |
|
| 149 |
I tested it on 1,636 hand-written messages across 10 routing setups: e-commerce, retail banking, insurance and telecom, plus an adversarial set of near-duplicates, typos, slang and "don't cancel, just tell me where it is" style traps. **Insurance and telecom were never seen in training.**
|
| 150 |
|
|
|
|
| 163 |
|
| 164 |
It also holds up when I poke at it. A perturbation test (shuffled intent order, removed correct intent, distractor intents, rewritten messages, gibberish) scores 0.92 averaged over 3 seeds, against 0.73 for the original Laya.
|
| 165 |
|
| 166 |
+
## Big intent lists (up to 148 intents)
|
| 167 |
|
| 168 |
The model reads everything in one 512-token window, so it can only see so many intents at once. For lists longer than 4, a small embedder (`bge-small-en-v1.5`, bundled in `shortlist/`) picks the 4 closest intents first and the router decides between those and "none".
|
| 169 |
|
|
|
|
| 201 |
|
| 202 |
| file | what |
|
| 203 |
|---|---|
|
| 204 |
+
| `laya.onnx` | the router, one file with int8 weights (302 MB) |
|
| 205 |
| `tokenizer.json`, `rl_agent_config.json` | tokenizer, prompt format, calibrated temperatures, threshold |
|
| 206 |
| `shortlist/` | bge-small-en-v1.5 ONNX embedder for long intent lists |
|
| 207 |
| `router.py` | the whole inference code, one file, no torch |
|
laya.onnx
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f4e29e431074a059977293de52408c4e18ce7f6e38d6b7b77e7e3e9406d67f34
|
| 3 |
+
size 302276803
|
laya.onnx.data
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:17cfafb27b977cd0537a61da24edecad8d8f50eac2d0eb9f3e14e5bc3116cf9c
|
| 3 |
-
size 299825152
|
|
|
|
|
|
|
|
|
|
|
|