Laya Intent Router 150M: zero-shot intent classification on CPU (ONNX, int8) and Apple Silicon (MLX)

A 150M-parameter zero-shot intent detection model with out-of-scope detection, 160 ms p95 on 4 CPU threads. You give it a user message and a list of intents written in plain English. It tells you which intent the message belongs to, or that it belongs to none of them.

No training. No labelled data. You write the intents when you call it, and you can change them on every request. It's a small ONNX file you run with onnxruntime and numpy, no GPU and no torch. On an Apple Silicon Mac, an optional MLX backend runs the same model on the GPU in about 15 to 25 ms.

On a held-out set of 1,636 messages it gets 94.9% routing accuracy and 98.6% out-of-scope recall, up from 72.8% and 68.2% for the original Laya.

"i want to cancel my order 88213"         -> cancel_order     0.97
"hey where's my parcel, it's been 5 days" -> order_status     0.97
"ordered black shoes, got blue ones lol"  -> wrong_item       0.97
"what's the weather in paris"             -> none             (0.98 none)
"qwewqeqw"                                -> none             (0.98 none)

That last line is why I built this. The original Laya sent qwewqeqw to order_not_received with 0.73 confidence. This one puts 0.98 on none.

Use it for

  • Chatbot and voicebot intent routing, where the menu of intents changes per flow or per customer.
  • Out-of-scope detection: knowing when a message fits none of your intents, so you can fall back to a human, an LLM or an "I didn't get that".
  • A cheap semantic router in front of an LLM, when you want a decision in 150 ms on CPU instead of a full LLM call.
  • Customer support triage for e-commerce, banking, insurance, telecom and SaaS.

It's built for short English customer messages. It isn't a general text classifier for long documents, and it only speaks English.

Try it

pip install onnxruntime tokenizers numpy huggingface_hub
from huggingface_hub import hf_hub_download
import importlib.util, sys

spec = importlib.util.spec_from_file_location("laya_intent_router",
    hf_hub_download("vrajnotviraj/laya-intent-router-150m-onnx", "laya_intent_router.py"))
lir = importlib.util.module_from_spec(spec); spec.loader.exec_module(lir)

r = lir.LayaIntentRouter.from_pretrained()   # downloads ~440 MB once

intents = {
    "check_balance":  ["What's my account balance"],
    "transfer_money": ["Send money to someone", "Transfer funds between accounts"],
    "block_card":     ["Block my card", "My card was lost or stolen"],
    "loan_enquiry":   ["Questions about personal or home loans"],
}

r.route("i think someone stole my card", intents)
# {'match': 'block_card', 'score': 0.98, 'probabilities': {..., '__none__': 0.01}}

r.route("ok", intents)
# {'match': None, 'score': 0.0, 'probabilities': {..., '__none__': 0.98}}

Each intent is a key plus 1 or more example phrasings or a short description. match is None when the message fits nothing, or when the best score is under the threshold (0.625 by default; pass threshold= to change it).

Or clone the GitHub repo and use from laya_intent_router import LayaIntentRouter, or run python laya_intent_router.py "your message here".

Faster on a Mac: MLX backend

On Apple Silicon the same model can run on the GPU through MLX. Install mlx next to the usual packages and pass backend="mlx" or backend="auto":

pip install onnxruntime tokenizers numpy huggingface_hub mlx
r = lir.LayaIntentRouter.from_pretrained(backend="auto")   # MLX if it can, ONNX if not
r.backend                                              # 'mlx'
lir.LayaIntentRouter.from_pretrained(backend="mlx")      # force MLX, errors out if mlx is missing
lir.LayaIntentRouter.from_pretrained()                   # default: ONNX on CPU, never downloads mlx/

The default is ONNX, so nothing changes unless you ask for MLX. auto only picks MLX on an Apple Silicon Mac with mlx importable and the mlx/ weights in the repo. Everywhere else (Linux, Windows, Intel Macs, or a Mac without mlx) it runs ONNX on CPU, same as before. Each backend downloads only its own weights, so ONNX users never pull the 328 MB mlx/ folder and MLX users never pull laya.onnx. mlx is optional and never required.

One catch: a copy of laya_intent_router.py from before September 29, 2026 downloads the whole repo, mlx/ included. Grab the current file to skip that.

The MLX weights are converted from the same trained checkpoint as laya.onnx (fp16, with the Linear layers quantized to 8 bits at load). On 5,000 dev routing episodes they made the same match-or-none decision as the ONNX model 99.7% of the time, and accuracy and out-of-scope recall were within 0.001. The few that differ sit right on the 0.625 threshold, where ONNX's own int8 rounding tips them one way or the other.

Latency on an M2 Pro, 1 message at a time, warm, whole route() call including the shortlist (p50 / p95, 300 calls each):

intents ONNX, 4 CPU threads MLX (M2 Pro GPU)
2 66 / 107 ms 15 / 22 ms
7 107 / 213 ms 24 / 37 ms
20 150 / 219 ms 28 / 38 ms

So about 4 to 5x faster. The machine was busy during these runs (ONNX p95 is 160 ms on a quiet one), so treat the p95s as upper bounds and compare the columns with each other. It doesn't get under 10 ms: the cost is 22 encoder layers over a ~225-token prompt, and dropping from fp32 to fp16 or 8-bit saves at most about 8 ms. I've only measured an M2 Pro.

More examples

A SaaS support bot, 4 intents, nothing fine-tuned:

message match score
cant get into my account reset_password 0.95
why was i charged twice this month billing 0.98
the export button does nothing bug_report 0.73
can i talk to an actual human please talk_to_human 0.99

The intents were reset_password: "User can't log in or forgot their password", billing: "Questions about invoices, charges or refunds", bug_report: "Something in the app is broken or showing an error" and talk_to_human: "User wants to speak to a real person". That's the whole setup.

Benchmarks: 94.9% routing accuracy, 98.6% out-of-scope recall

I tested it on 1,636 hand-written messages across 10 routing setups: e-commerce, retail banking, insurance and telecom, plus an adversarial set of near-duplicates, typos, slang and "don't cancel, just tell me where it is" style traps. Insurance and telecom were never seen in training.

Laya (original, zero-shot) Laya-large fine-tuned (teacher) This model
Routing accuracy 0.728 0.950 0.949
Catches out-of-scope messages 0.682 0.959 0.986
Wrongly accepts out-of-scope 0.318 0.041 0.014
Big menus (20 to 45 intents) 0.704 0.938 0.943
6 brand new domains 0.777 0.920
166 extra test messages, written last 0.566 0.934 0.952
p95 latency, 4 CPU threads 252 ms 496 ms 160 ms
Size on disk 636 MB 636 MB 304 MB (+134 MB shortlist)

So you get the big fine-tuned model's accuracy at a third of its latency and half its size. Latency was measured on an Apple M2 Pro. Your server's vCPUs are probably slower, so measure there.

It also holds up when I poke at it. A perturbation test (shuffled intent order, removed correct intent, distractor intents, rewritten messages, gibberish) scores 0.92 averaged over 3 seeds, against 0.73 for the original Laya.

Big intent lists (up to 148 intents)

The model reads everything in one 512-token window, so it can only see so many intents at once. For lists longer than 4, a small embedder (bge-small-en-v1.5, bundled in shortlist/) picks the 4 closest intents first and the router decides between those and "none".

That's on by default and it's why accuracy stays flat as the list grows:

intents full list with shortlist
7 0.915 0.930
12 0.969 0.979
45 0.781 0.938
148 too long to run 0.889

The first call on a new intent list embeds all its phrasings (about 0.9 s for 50 intents), then it's cached. Later calls stay around 180 ms p95 whatever the list size. Pass shortlist_k=0 to turn it off.

How it was made

The base is Laya, a ModernBERT-large decision model that scores a list of options in one forward pass. Out of the box it was too eager to match, so:

  1. I fine-tuned Laya-large as a teacher on about 122k routing episodes built from 7 public intent datasets (CLINC150, BANKING77, HWU64, SNIPS and 3 Bitext customer-support sets) plus 76 synthetic workflows across 16 domains. About 40% of episodes had the right intent removed, so the model learns to say "none". Loss was soft cross-entropy plus Laya's RL objective.
  2. I rewrote the prompt format. Each intent went from a wordy wrapper to key: "phrasing 1" | "phrasing 2", with leftover token budget handed to intents that need it. That alone was worth about 3 points.
  3. I distilled the teacher into Ettin-150M with a fresh Laya head, mixing the teacher's probabilities 50/50 with the gold label. Checkpoints were picked by a perturbation-based intent score, since plain accuracy picked worse routers.
  4. I calibrated temperatures per menu size, then exported to ONNX with 8-bit weight quantization.

Everything trained locally on a 32 GB M2 Pro. Training code: github.com/vrajnotviraj/laya-intent-router.

Where it slips

  • English only. It hasn't seen other languages.
  • Filler on tiny menus. With only 2 intents like "Hi" and "Bye", words like "ok", "well" and "nice" can land on "Bye" (0.67 to 0.81). Give short closing intents a clear description, or raise the threshold for small menus.
  • It only knows what your phrasings say. "My card was stolen" won't hit a block_card intent whose only phrasing is "Block my card". Add 2 or 3 phrasings that cover how people actually talk.
  • Indirect requests and heavy typos are the weakest slices (about 0.83 to 0.86).
  • My test messages were written by the same process as the synthetic training data. Real user logs might be harder. Treat the numbers as a strong hint and run your own messages through it.

Files

file what
laya.onnx the router, one file with int8 weights (302 MB)
tokenizer.json, rl_agent_config.json tokenizer, prompt format, calibrated temperatures, threshold
shortlist/ bge-small-en-v1.5 ONNX embedder for long intent lists
laya_intent_router.py the whole inference code, one file, no torch
mlx/ the same model as MLX weights for Apple Silicon (fp16, 8-bit Linear at load, 328 MB), optional
laya_intent_router_mlx.py the MLX model code, only loaded when you pass backend="mlx" or "auto"

FAQ

Can I run it in Ollama or LM Studio?

No, and I don't think it's possible without breaking the model. Ollama and LM Studio run GGUF files through llama.cpp, which is built for text generation and plain embeddings. This model isn't either of those. After the encoder it has a small decision head: 2 extra transformer layers, a scorer that reads one position per intent (the [MASK] in front of each one) and a temperature per menu size. GGUF has nowhere to put that head, and llama.cpp has no way to run it. You'd get an embedding model and lose the part that says "none of these".

If you want it local and fast on a Mac, use the MLX backend above. It runs the whole graph, head included, in-process.

License and credits

Apache-2.0. Built on convaiinnovations/laya (Apache-2.0, ModernBERT-large), jhu-clsp/ettin-encoder-150m (MIT) and BAAI/bge-small-en-v1.5 (MIT).

Training data: CLINC150 (CC BY 3.0, Larson et al. 2019), BANKING77 (CC BY 4.0, Casanueva et al. 2020), HWU64 (CC BY 4.0, Liu et al. 2019), SNIPS (CC0) and the Bitext customer-support, retail e-commerce and retail banking datasets (CDLA-Sharing-1.0). No training data is included here.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vrajnotviraj/laya-intent-router-150m-onnx

Finetuned
(28)
this model

Datasets used to train vrajnotviraj/laya-intent-router-150m-onnx

Evaluation results

  • Routing accuracy at threshold 0.625 on Held-out routing suite (1,636 messages, 10 workflows)
    self-reported
    0.949
  • Out-of-scope recall on Held-out routing suite (1,636 messages, 10 workflows)
    self-reported
    0.986
  • Routing accuracy, menus of 20 to 45 intents on Held-out routing suite (1,636 messages, 10 workflows)
    self-reported
    0.943
  • Routing accuracy, 6 unseen domains on Held-out routing suite (1,636 messages, 10 workflows)
    self-reported
    0.920