# Delegate a complete browser workflow Runtime 1.1 adds named targets, staged execution, bounded recovery and a compact result. The browser-head-v5 weights and inference schema are unchanged. A planning agent translates the user's instruction into a short request; NanoJev-Web is still an action selector, not a free-text chat agent or autonomous website planner. ## One request With the local model running, adapt [browser-workflow.json](../examples/browser-workflow.json) and run: ```bash ./run.command --task examples/browser-workflow.json ``` The example requires your own authorized page. It does not launch a demo site. `name` is optional; allowed origins default to the initial URL's origin. For permitted navigation across origins, supply `allowedOrigins` explicitly. ```json { "url": "http://127.0.0.1:9000/contacts", "steps": [ {"click": "Edit"}, {"fill": {"Name": "Test Contact", "Address": "10 Example Street"}}, {"select": {"Plan": "Team"}}, {"check": {"Consent": true}}, {"click": "Save"} ], "expect": {"visibleText": "Saved"} } ``` `click` identifies a named button/link. `fill` assigns text fields; `select` uses exact option labels in a native select (an array selects multiple options); `check` assigns booleans to checkboxes. One field step may contain several fields, and the next stage starts after all assigned values match. Click stages execute once and then move to the next stage. A final click is classified as `submit` by default; an explicit `purpose` can override that classification. Native date/time fields use the existing keyboard helpers. Targets are resolved again after every transition, so a modal's fields do not need precomputed CSS selectors. Matching uses exact normalized labels, aria names, placeholders or button text, with an active modal as the default scope. This is a deliberately bounded DOM resolver, not complete W3C accessible-name computation, fuzzy language understanding or screenshot reasoning. It does not invent missing values or select an arbitrary match. For a duplicate label, narrow the target explicitly: ```json {"fill": [{"target": {"name": "Name", "within": "#editor"}, "value": "Test Contact"}]} ``` A target object also accepts `role` or an explicit `selector`. `within` is a CSS selector identifying one visible container. The existing `controls` contract remains available for custom widgets, conditional actions, frames and other advanced operations. Do not combine `steps` and `controls` in a single task. ## Model and hybrid modes The default is `hybrid`. Explicit `--mode model` sends each action choice to NanoJev-Web. Workflow-stage advancement and target resolution are executor responsibilities. Default `hybrid` execution (also selectable with `--mode hybrid`) executes explicitly assigned, visible native field operations deterministically after a fresh observation. Clicks, scrolling, waits and termination remain model decisions. There is no second model, cloud fallback or extra inference provider. Reports label every executed action's source and count deterministic actions separately. Hybrid performance describes the combined system; it must not be reported as all actions having been chosen by the model. The observer combines origin, targets, values, visibility and motion checks into one read-only DOM request, with a separate active-tab check. It no longer takes a full accessibility snapshot on every observation; temporal input helpers may still request focused snapshots. A model's decision is checked against a new observation before execution. Hybrid input actions execute immediately from their just-read state without a redundant second observation. ## Attach to an owned session ```bash ./run.command --task workflow.json --session my-test --attach ``` Attach requires an already running, explicitly named agent-browser session. It preserves the active tab, omits navigation and viewport setup, and leaves the session open. `url` may be omitted in attach mode; the active tab's URL supplies the default origin. An integration that owns exact role/profile flags should continue using its own bound adapter; the standalone CLI does not read third-party session manifests. `--stdin` accepts a JSON request from standard input instead of `--task`. `--keep-open` retains a newly created session. Standard output contains one JSON result; `--verbose` writes progress to standard error. A caller can also import `runTask` from `runtime/browser.mjs` and supply its owned `Browser` instance with `setup: false`. ## Animation and BLOCKED recovery The observer samples target geometry 25 ms apart and waits up to 1,200 ms for relevant finite animations to finish and geometry to stop changing. An unsettled page remains busy. This is a bounded stability check, not a promise that every possible JavaScript animation has finished. On a model `BLOCKED` decision, the runner permits **two retries, 500 ms apart**. Each retry observes the page again and sends a new decision request. A third consecutive refusal returns `model_blocked`. A successful browser action resets the refusal counter. Existing decision/time limits still apply. `--blocked-retries` and `--retry-delay-ms` configure this policy; zero retries restores immediate handoff. Repeating an identical request is not expected to fix a deterministic model decision. Recovery is useful when the page changes during the interval. Incorrect targeting or classification must be repaired separately. Use `interaction` for an Edit/open-details button; `reveal` is reserved for generated scroll actions. Legacy controls marked `reveal` are normalized to their input/interaction purpose. Explicit `advance`, `submit` and `dismiss` purposes remain intact. Retries do not replay failed clicks or uncertain writes. An execution error after an action is attempted returns `action_uncertain`. A submit control is issued at most once in a run. If success is still missing, inspect/reconcile the result rather than submitting again. ## Result and verification The compact result includes status, completed stages, model decisions, deterministic actions, retry count, local input tokens, loop timings, browser calls, visible success, missing targets and evidence paths. Full traces retain exact model requests, answers and fresh observations, including unsuccessful retries, in private `.local/runs/` files. Assigned field values may appear in these local traces. They are not public release artifacts. Loop time excludes initial navigation/setup and final evidence capture. Browser loop metrics use the same scope. Local API cost is zero; hardware and the planning agent's usage are not included. Input tokens count separately encoded candidate paths, including repeated context. A visible success message is a UI checkpoint. The planning agent still verifies saved data and business rules independently. Ambiguous targets, unsupported controls and unresolved states return control to that agent. ## Checks `./test.command` covers request validation, retry bounds, stale observations, uncertain writes, submit-once behavior and workflow-stage handling. With the local model and agent-browser installed, run the optional isolated browser regression: ```bash node tests/browser-live.mjs ``` It starts a loopback synthetic form, exercises a 300 ms sliding modal, named fields, native selection, scrolling and saving through agent-browser, independently checks server receipts, and verifies that attach preserves the session. It does not access a real application or call cloud models. [Recorded runtime validation](RUNTIME_1_1_REPORT.md).