laya-onnx

ONNX Runtime for Laya β€” typed System-1 decisions, no generated tokens. PC sibling of laya-coreml; same choice / score / noul contract.

This Hub repository mirrors the project's source tree (runtime, CLI, Snake demo, benchmarks, tests, docs). Model weights are not included here.

Model description

Task Typed decision β€” choice, score, noul
Output per-question label/score + probability; no generated tokens (output_tokens: 0)
Backbone Laya (ModernBERT-style encoder + RL decision head)
Runtime ONNX Runtime β€” CPU, Intel OpenVINO, CUDA
Weights receptron/laya-onnx (~1.6 GB, fp32)
Precision fp32; laya-onnx optimize --precision int8 on Intel CPU
License Apache-2.0 (following Laya)

laya-onnx answers typed questions about a state and returns probabilities β€” it never writes text. predict() batches every question for a state into one call and decides by argmax; predict_argmax() ignores the calibrated temperature for a reproducible path.

Install

git clone https://github.com/Geoking2104/laya-onnx && cd laya-onnx
python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -U pip && pip install -e ".[demo,dev]"
pip install -e ".[openvino]"   # Intel CPU
pip install -e ".[ultrafast]"  # optional: Playwright DOM loop

Model weights are not in git. The first load("receptron/laya-onnx") (or CLI run) fetches them from the Hub.

Usage

from laya_onnx import load

agent = load("receptron/laya-onnx", providers="cpu")
print(agent.predict(
    "The customer requests a refund of a duplicate payment.",
    {"refund": {"type": "noul", "instructions": "Does the customer request a refund?"}},
))

Deterministic path (single thread, no padding, temperature ignored):

agent = load("receptron/laya-onnx", providers="cpu", deterministic=True)
print(agent.predict_argmax(state, questions))

CLI

laya-onnx predict  --state "..." --questions q.json --model ./onnx
laya-onnx convert  --model convaiinnovations/laya --output onnx --precision fp32
laya-onnx optimize ./onnx --precision int8           # Intel CPU
laya-onnx verify   --model ./onnx                     # SHA-256 check vs bundled manifest
laya-onnx-ultrafast --dry-run --fixture examples/ultrafast_page.json --goal "..."
laya-onnx-snake --model ./onnx                        # terminal Snake

Published benchmark

CPU run (Windows 11, Intel i7-1255U, onnxruntime), bundle receptron/laya-onnx @ 68f27dfe:

question type n accuracy MAE ECE
choice 4 1.00 – 0.456
noul 4 1.00 – 0.129
score 2 0.50 0.530 0.112
overall 10 0.90 – 0.256

Latency per predict() call: mean 2808 ms / p50 1760 ms / p95 7418 ms. Small self-check set (benchmarks/eval/decisions.jsonl), not a leaderboard β€” see the bench dataset for the method and caveats.

Calibration

Confidence is calibrated with per-question-type temperatures (plus per option-count buckets); load() warns and clamps any temperature outside the accepted range. Calibration is reported as ECE in the benchmark above.

Verify a download

laya-onnx verify --model ~/.cache/huggingface/hub/laya-onnx-bundles/receptron--laya-onnx

Hashes every bundle file against the SHA-256 manifest shipped in laya_onnx/checksums/.

Repository map

  • laya_onnx/ β€” runtime (agent, hub, inputs, tokenizer, convert, optimize, eval, verify, cli)
  • laya_onnx/snake/ Β· laya_onnx/ultrafast/ β€” demos (terminal Snake; Playwright DOM loop)
  • benchmarks/ β€” evaluate.py, pc_benchmark.py, eval/decisions.jsonl, published results.*
  • docs/ β€” DETERMINISTIC.md, ULTRAFAST.md
  • examples/ β€” quickstart, sample state/questions, snake.html browser mock

Attribution

Apache-2.0. Derived from Laya (Apache-2.0) and laya-mlx. Ultrafast loop design from Browser Use. See LICENSE and NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for gdelatournelle/laya-onnx

Finetuned
(132)
this model