--- license: apache-2.0 base_model: convaiinnovations/laya tags: - laya - decision-model - mind2web - web-agent - browser-automation datasets: - osunlp/Mind2Web language: en --- # Laya Browser (Mind2Web fine-tune) A [Laya](https://huggingface.co/convaiinnovations/laya) (421M, ModernBERT-large encoder + typed decision heads) checkpoint fine-tuned to pick the next browser action — an operation (`CLICK` / `TYPE_TEXT` / `SELECT`) and a target element — from a page's DOM and a natural-language goal. Trained on [Mind2Web](https://huggingface.co/datasets/osunlp/Mind2Web) (Deng et al., NeurIPS 2023 Datasets and Benchmarks Track, CC-BY-4.0). Built as the policy for [laya-system-use](https://github.com/HiteshS08/laya-system-use), a fork of [browser-use/jev-ultrafast](https://github.com/browser-use/jev-ultrafast) that replaces the hosted TypeSafe Jev API with this open-weight model. **This is a research/hobby-project checkpoint, not a production model.** Read the Limitations section before using it. ## What it does Given a page's interactive elements (label, role, current value) and a goal, it answers two typed questions in one forward pass: which operation to perform, and which element to act on. It does not generate text — a separate small local LLM fills in typed values (`TYPE_TEXT`), and the element ranking is done by an untrained lexical shortlister before this model ever sees the candidates (see Methodology). ## Methodology ### Data pipeline 1. **Parsing.** Mind2Web's raw HTML/candidate format is converted into a neutral `Candidate(id, label, role, value, ops)` record. Labels come from the DOM (`aria-label`, text content, `alt`, `title`, `placeholder`, `value`, `name`, in that order), falling back to the element's role when none exist — matching what the live browser snapshot exports, so training and serving never see different inputs. 2. **Shortlisting.** Mind2Web pages average ~140-580 candidate elements; the fine-tuned model's context budget only fits ~20. An untrained lexical ranker (rare-word-weighted overlap between the goal/history and each candidate's label, plus a small role prior and page-position prior) keeps the top 20 per operation. This ranker was tuned on Mind2Web's *training* data only: it raised recall@20 from 0.686 to 0.755 (train) / 0.662 to 0.790 (dev) before any model training began. 3. **Reclaiming non-interactive gold elements.** ~15% of steps have a labeled target that isn't itself clickable/typeable (an icon inside a button, a `