laya-onnx
ONNX Runtime for Laya β typed
System-1 decisions, no generated tokens. PC sibling of
laya-coreml; same choice / score /
noul contract.
- Source code: https://github.com/Geoking2104/laya-onnx
- Weights (not hosted in this repo, β1.6 GB fp32): https://huggingface.co/receptron/laya-onnx
- Benchmark Β· calibration Β· checksums: https://huggingface.co/datasets/gdelatournelle/laya-onnx-bench
This Hub repository mirrors the project's source tree (runtime, CLI, Snake demo, benchmarks, tests, docs). Model weights are not included here.
Model description
| Task | Typed decision β choice, score, noul |
| Output | per-question label/score + probability; no generated tokens (output_tokens: 0) |
| Backbone | Laya (ModernBERT-style encoder + RL decision head) |
| Runtime | ONNX Runtime β CPU, Intel OpenVINO, CUDA |
| Weights | receptron/laya-onnx (~1.6 GB, fp32) |
| Precision | fp32; laya-onnx optimize --precision int8 on Intel CPU |
| License | Apache-2.0 (following Laya) |
laya-onnx answers typed questions about a state and returns probabilities β
it never writes text. predict() batches every question for a state into one
call and decides by argmax; predict_argmax() ignores the calibrated temperature
for a reproducible path.
Install
git clone https://github.com/Geoking2104/laya-onnx && cd laya-onnx
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -U pip && pip install -e ".[demo,dev]"
pip install -e ".[openvino]" # Intel CPU
pip install -e ".[ultrafast]" # optional: Playwright DOM loop
Model weights are not in git. The first load("receptron/laya-onnx") (or CLI
run) fetches them from the Hub.
Usage
from laya_onnx import load
agent = load("receptron/laya-onnx", providers="cpu")
print(agent.predict(
"The customer requests a refund of a duplicate payment.",
{"refund": {"type": "noul", "instructions": "Does the customer request a refund?"}},
))
Deterministic path (single thread, no padding, temperature ignored):
agent = load("receptron/laya-onnx", providers="cpu", deterministic=True)
print(agent.predict_argmax(state, questions))
CLI
laya-onnx predict --state "..." --questions q.json --model ./onnx
laya-onnx convert --model convaiinnovations/laya --output onnx --precision fp32
laya-onnx optimize ./onnx --precision int8 # Intel CPU
laya-onnx verify --model ./onnx # SHA-256 check vs bundled manifest
laya-onnx-ultrafast --dry-run --fixture examples/ultrafast_page.json --goal "..."
laya-onnx-snake --model ./onnx # terminal Snake
Published benchmark
CPU run (Windows 11, Intel i7-1255U, onnxruntime), bundle receptron/laya-onnx
@ 68f27dfe:
| question type | n | accuracy | MAE | ECE |
|---|---|---|---|---|
| choice | 4 | 1.00 | β | 0.456 |
| noul | 4 | 1.00 | β | 0.129 |
| score | 2 | 0.50 | 0.530 | 0.112 |
| overall | 10 | 0.90 | β | 0.256 |
Latency per predict() call: mean 2808 ms / p50 1760 ms / p95 7418 ms. Small
self-check set (benchmarks/eval/decisions.jsonl), not a leaderboard β see
the bench dataset
for the method and caveats.
Calibration
Confidence is calibrated with per-question-type temperatures (plus per
option-count buckets); load() warns and clamps any temperature outside the
accepted range. Calibration is reported as ECE in the benchmark above.
Verify a download
laya-onnx verify --model ~/.cache/huggingface/hub/laya-onnx-bundles/receptron--laya-onnx
Hashes every bundle file against the SHA-256 manifest shipped in
laya_onnx/checksums/.
Repository map
laya_onnx/β runtime (agent,hub,inputs,tokenizer,convert,optimize,eval,verify,cli)laya_onnx/snake/Β·laya_onnx/ultrafast/β demos (terminal Snake; Playwright DOM loop)benchmarks/βevaluate.py,pc_benchmark.py,eval/decisions.jsonl, publishedresults.*docs/βDETERMINISTIC.md,ULTRAFAST.mdexamples/β quickstart, sample state/questions,snake.htmlbrowser mock
Attribution
Apache-2.0. Derived from Laya (Apache-2.0) and laya-mlx. Ultrafast loop design
from Browser Use. See LICENSE and NOTICE.
Model tree for gdelatournelle/laya-onnx
Base model
convaiinnovations/laya