--- license: apache-2.0 base_model: convaiinnovations/laya language: - en - multilingual tags: - laya - system-1 - browser-agent - web-navigation - decision-model - mmbert - mind2web - tilelang datasets: - osunlp/Mind2Web - stanfordnlp/nnetnav-live - webagentlab/webchain --- # laya-browser — laya fine-tuned as a browser-agent decision head (drop-in replacement for TypeSafe Jev) ![laya driving a real browser on held-out sites: search, filters, open a result](assets/laya_browser_demo.gif) **laya** ([convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya)) is a non-autoregressive "System 1" decision model: one bidirectional encoder pass answers several typed questions (`choice` / `score` / `noul`) with calibrated probabilities, no text generation. This repo fine-tunes it into the decision head of [browser-use/jev-ultrafast](https://github.com/browser-use/jev-ultrafast), whose `/v1/systemone` request is exactly laya's `predict(state, questions)`: every step, one forward pass (~22 ms at full GPU clock) picks the operation (CLICK / TYPE_TEXT / SELECT / PRESS_ENTER / SCROLL / DONE …) and its target element. Trained and evaluated locally on one RTX 4070 Ti SUPER (16 GB), no paid API; a local Qwen3-8B-AWQ only writes the text that TYPE_TEXT types and picks dropdown values. **One model, `v19s/`.** Earlier checkpoints are in the commit history. This repo is updated only when a new version is clearly better. ## Results (v19s, mmBERT-base 322M) All numbers below were measured with the harness in `code/jev-ultrafast.patch`; the v17s column was measured with the harness it shipped with, so part of the gain is the harness (see "What changed"). | evaluation | v17s | **v19s** | |---|---|---| | **Suite C: 27 multi-step tasks on held-out real sites** (search + filters + sort + open), ×2 | 2/54 | **11–14/54** (two ×2 runs; single ×1 runs: 5–7/27) | | Suite B: 18 tasks on 18 held-out sites, ×3 | 54/54 | **54/54** | | Suite A: 16 original real-site tasks, ×3 | 41/48 | 39/48 | | webgym held-out synthetic forms, 7 kinds × 10 (flight, hotel, shop, car, restaurant, signup, filter) | 15/70 | **30/70** | | held-out decisions on sites unseen in training (WebChain): click operation / target top-1 | — | 0.94 / 0.48 | - **Suite C** is the honest headline: real multi-step tasks on domains absent from every training source. v19s reaches ~20–26 %; single runs vary by ±4 tasks (site load times, popups), so treat differences smaller than that as noise. - Suite A lost `books-open-book` (0/3), gained nothing else; suite B unchanged. **Latency.** One decision is **22–27 ms** on an RTX 4070 Ti SUPER *while the GPU is at full clock* (requests back to back). An agent waits seconds between decisions for pages to load, and the GPU drops to idle clocks (P5/P8, 210–850 MHz) within a second or two: the next decision then takes **75–270 ms** (the demo above shows these live numbers). Locking the minimum SM clock removes that — `sudo nvidia-smi -lgc 2100,3135` (undo: `sudo nvidia-smi -rgc`; costs some idle power). The first request of each input-length bucket also compiles a kernel once (~100–500 ms); warm up with a few requests after starting the server. ## Use ```bash huggingface-cli download cklxx/laya-browser --local-dir laya-browser cd laya-browser/code && uv sync --extra fast # pinned uv.lock (Python 3.12, torch 2.11, tilelang 0.1.14) uv run python verify.py # downloads v19s, answers one recorded browser step ``` As a TypeSafe replacement for jev-ultrafast (apply `code/jev-ultrafast.patch` to jev-ultrafast `1231850`): ```bash python code/apps/systemone_server.py 8791 /path/to/laya-browser/v19s 60 # 60 = split choices wider than 60 options # jev-ultrafast: TYPESAFE_BASE_URL=http://127.0.0.1:8791 ``` The checkpoint records `laya_fmt` (**v5**) and `head_max_len_train` (768); the server applies the matching input format: option labels without the duplicated `[key]`, `