# Runtime 1.1 validation This release improves delegation and recovery around the unchanged NanoJev-Web browser-head-v5 checkpoint. It does not add training, change the model API schema, or replace the recorded 15-field demonstration. ## Real browser regression The synthetic workflow opens a card, waits for a 300 ms sliding dialog, scrolls to Edit, opens the editor, fills Name/Address/Email, selects Plan by option label, checks Consent, and saves. Element IDs are generated anew. Task requests use labels rather than CSS selectors. All UI actions use agent-browser; an independent loopback server checks the exact five assigned values and records one submission per completed attempt. Hardware: Apple M3 Max, 36 GB unified memory; macOS 26.5.2; Node.js 24.19.0; agent-browser 0.37.1; local PyTorch MPS, FP32. No hosted model calls. The package's separately reported M2 Pro / 16 GB compatibility check is unchanged and is not the source of these timings. Execution order was model, hybrid, hybrid, model, model, hybrid. All six runs passed. One additional CLI run passed while preserving its already-open session and active tab. There were seven exact server receipts in total. | Mode | Individual loop times (s) | Median (s) | Model decisions | Browser actions | Deterministic actions | Browser loop calls | Local input tokens | |---|---|---:|---:|---:|---:|---:|---:| | model | 6.997, 4.594, 4.640 | 4.640 | 10 | 9 | 0 | 55 | 7,074 | | hybrid | 3.319, 3.358, 3.377 | 3.358 | 5 | 9 | 5 | 45 | 3,175 | On this workflow, hybrid median loop time was **27.6% lower** and model requests fell from 10 to 5. Hybrid explicitly delegates five input operations to deterministic code; these are not model-selected actions. It is a measurement of the combined runtime, not a claim that the neural model itself became faster. Three runs on one fixture are limited evidence, and the slower first model run is retained above. Elapsed time excludes initial navigation/setup and final evidence capture. Browser loop metrics use the same scope. Candidate-path input tokens include repeated context. There are zero autoregressive output tokens and zero hosted API charges; host-agent tokens, hardware and electricity were not measured. These figures must not be compared directly with the earlier 15-field benchmark, which uses a different task and contract. ## Recovery and correctness - Six Python inference-request contract checks and eleven Node runtime checks passed. - Persistent BLOCKED returns after three model refusals, with two measured 500 ms pauses and new observations. This policy is tested with controlled model responses; normal live runs did not require retries. - A transient refusal can recover; a failed or uncertain click is not replayed. A submit control is issued at most once even when visible success is absent. - Stale observations cannot execute a queued model action. A preexisting Saved banner cannot skip pending workflow stages. - The live animation observation settled after approximately 267 ms; it did not act on a moving target. - A real dialog with duplicate Name labels was rejected before input. Legacy reveal classification on Edit was normalized to interaction, with a scroll candidate available for the offscreen button. - Exact native-select label matching was checked with options whose displayed labels differ from their stored values. - The CLI returned one compact JSON result and left its attached tab/session intact. A reconstructed delayed-dialog fixture is not proof of the cause of any earlier unrecorded failure. Future private traces retain the exact requests, responses and fresh observations needed for diagnosis. Run `node tests/browser-live.mjs` with the local model running to repeat this regression. It uses an isolated loopback fixture and agent-browser sessions, not a production application or a comparison stand. Raw local traces and screenshots stay outside the public release. [Machine-readable measurements](../evals/runtime-1.1.json) ยท [Workflow interface](WORKFLOWS.md). ## Published default The published 1.1 runtime defaults to `hybrid`; `--mode model` remains available. The six timing runs above explicitly selected their mode and are unchanged. A separate browser regression omitted `--mode` entirely: the CLI used hybrid, performed five deterministic input actions, preserved its attached session, and produced one exact server receipt. Unit checks cover both the default and the model-only override. This changes executor configuration, not checkpoint weights.