Instructions to use cklxx/laya-browser with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cklxx/laya-browser with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="cklxx/laya-browser", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("cklxx/laya-browser", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("cklxx/laya-browser", trust_remote_code=True, device_map="auto")laya-browser — laya fine-tuned as a browser-agent decision head (drop-in replacement for TypeSafe Jev)
laya (convaiinnovations/laya) is a non-autoregressive "System 1" decision model:
one bidirectional encoder pass answers several typed questions (choice / score / noul) with calibrated probabilities, no text
generation. This repo fine-tunes it into the decision head of browser-use/jev-ultrafast,
whose /v1/systemone request is exactly laya's predict(state, questions): every step, one forward pass (~22 ms at full GPU clock) picks the operation
(CLICK / TYPE_TEXT / SELECT / PRESS_ENTER / SCROLL / DONE …) and its target element. Trained and evaluated locally on one RTX 4070
Ti SUPER (16 GB), no paid API; a local Qwen3-8B-AWQ only writes the text that TYPE_TEXT types and picks dropdown values.
One model, at the repo root (v19s). Earlier checkpoints are in the commit history. This repo is updated only when a new version is clearly better.
Results (v19s, mmBERT-base 322M)
All numbers below were measured with the harness in code/jev-ultrafast.patch; the v17s column was measured with the harness it
shipped with, so part of the gain is the harness (see "What changed").
| evaluation | v17s | v19s |
|---|---|---|
| Suite C: 27 multi-step tasks on held-out real sites (search + filters + sort + open), ×2 | 2/54 | 11–14/54 (two ×2 runs; single ×1 runs: 5–7/27) |
| Suite B: 18 tasks on 18 held-out sites, ×3 | 54/54 | 54/54 |
| Suite A: 16 original real-site tasks, ×3 | 41/48 | 39/48 |
| webgym held-out synthetic forms, 7 kinds × 10 (flight, hotel, shop, car, restaurant, signup, filter) | 15/70 | 30/70 |
| held-out decisions on sites unseen in training (WebChain): click operation / target top-1 | — | 0.94 / 0.48 |
- Suite C is the honest headline: real multi-step tasks on domains absent from every training source. v19s reaches ~20–26 %; single runs vary by ±4 tasks (site load times, popups), so treat differences smaller than that as noise.
- Suite A lost
books-open-book(0/3), gained nothing else; suite B unchanged.
Latency. One decision is 22–27 ms on an RTX 4070 Ti SUPER while the GPU is at full clock (requests back to back).
An agent waits seconds between decisions for pages to load, and the GPU drops to idle clocks (P5/P8, 210–850 MHz) within a
second or two: the next decision then takes 75–270 ms (the demo above shows these live numbers). Locking the minimum SM
clock removes that — sudo nvidia-smi -lgc 2100,3135 (undo: sudo nvidia-smi -rgc; costs some idle power). The first request
of each input-length bucket also compiles a kernel once (~100–500 ms); warm up with a few requests after starting the server.
Use
transformers (the model ships its own code; only torch + transformers are needed):
from transformers import AutoModel
m = AutoModel.from_pretrained("cklxx/laya-browser", trust_remote_code=True).to("cuda") # CPU works too
page = {"url": ..., "title": ..., "text": ..., # the page: interactive elements + visible text
"actions": [{"id": "e1", "kind": "fill", "node": 1, "label": "Search packages", "role": "searchbox"},
{"id": "e2", "kind": "click", "node": 2, "label": "argo", "role": "link"}, ...]}
d = m.decide(page, goal="Search packages for 'json' and open the package 'argo'.", history=[])
d["operation"], d["action"], d["confidence"] # "CLICK", {"id": "e2", ...}, 0.88
decide builds the request exactly in the format the model was trained on (instructions, compact option strings, form-field
summary, 1200 chars of page text, split of choices wider than 60) and runs one forward pass. m.systemone(body) answers a
jev-ultrafast /v1/systemone request body. The same checkpoint also loads with the upstream laya package
(laya.load("cklxx/laya-browser")); laya_browser.py in this repo wraps it the same way and adds a server:
As a TypeSafe replacement for browser-use/jev-ultrafast:
python laya_browser.py serve --port 8791 # --model <local dir> to use a downloaded copy
TYPESAFE_BASE_URL=http://127.0.0.1:8791 TYPESAFE_API_KEY=local <run jev-ultrafast as usual>
It answers jev's /v1/systemone requests identically to the evaluation server (checked on 21 real steps: same operation and
target on all 21). The harness improvements behind the suite numbers are in code/jev-ultrafast.patch (apply to
jev-ultrafast 1231850); the full evaluation / training setup is in code/ (uv sync --extra fast, code/verify.py).
What changed since v17s
Harness (code/jev-ultrafast.patch, all measured case by case on failures):
- wait for the navigation an Enter / click starts before observing (the agent used to see the old page and think Enter did nothing);
- an "undo guard": for 3 steps after a filter / sort / radio click that changed the page, that control (and its "remove filter" chip, and other options of the same dropdown) is not offered again — the agent used to toggle filters on and off until its budget ran out;
- the text model gets each field's placeholder / type / pattern (dates in
MM/DD/YYYYfields) and picks the value of a dropdown the policy chose (ages, times, countries that differ by a digit); - options scrolled out of view inside an open list are offered and scrolled into view before clicking.
Data (v19s = mmBERT-base v17s continued, full fine-tune):
- WebChain (CC-BY-4.0): 3,000 human trajectories on real sites, 10.9k
steps after dropping unlabeled / duplicate-label targets; the gold element is recovered exactly from each step's DOM snapshot
via its CSS selector (
code/finetune/convert_webchain.py); - webgym grown to 7 task kinds, plus DAgger on webgym (the model drives, the scripted expert labels every visited state);
- format v5 (above).
What still fails
- Stopping too early is ~60 % of real-site failures: after a search the model often says DONE on the results page when the
task asks to open a result or a sub-page ("find the recipe page of X", "open its episode list"). More DONE data (Go-Browse),
cost-sensitive training, a noul completion head, goal-contrast twins and a run-time sub-goal planner were all tried; none moved
suite C beyond noise (details in
results/v19s/). A 322M single-pass policy does not reliably read these goal distinctions. - Pages whose target is several screens down behind many links; sites that block headless Chromium (roughly a third of the Online-Mind2Web sites); logins (jev hides password fields by design).
Changelog
| version | date | change |
|---|---|---|
| v10s | 2026-09-21 | first usable model: format v3, 17–23 ms per step |
| v14s | 2026-09-26 | harness fixes, Enter / scroll / select data, NNetNav; suite B 100 % |
| v17s | 2026-09-27 | dropdown option names kept, real category links; suite A 85 %, webgym 47 % (3 kinds) |
| v19s | 2026-09-29 | format v5, WebChain real-site trajectories, webgym ×7 + DAgger, harness fixes; suite C 4 % → ~20–26 %, webgym 21 % → 43 % (7 kinds) |
Files
model.safetensors, encoder/, tokenizer/, rl_agent_config.json the model (v19s; a laya checkpoint dir)
modeling_laya_browser.py, config.json transformers remote code (AutoModel + trust_remote_code)
laya_browser.py the same on top of the laya package, plus a TypeSafe-compatible server
code/ server, suites A/B/C, Online-Mind2Web runner + judge, finetune pipeline (webgym, WebChain / Go-Browse converters),
TileLang kernels, jev-ultrafast patch, verify.py
results/ suite JSONs, traces and logs behind the numbers above (results/v19s/, older versions in their folders)
assets/ demo video
License
Apache-2.0, same as laya. Mind2Web, NNetNav and WebChain are used under their own licenses for training only.
- Downloads last month
- -

# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="cklxx/laya-browser", trust_remote_code=True)