Recorded browser demonstration: Jev, NanoJev-Web and GPT-6 Astra High
This is supporting documentation for NanoJev-Web. The comparison application used to record it is separate from the model distribution and is not included here. Local model inference and browser execution do not require a comparison provider.
Observed result: NanoJev-Web and Astra completed and saved all 15 assigned values. Jev filled four values, then alternated between two scroll targets until the harness stopped the repeated state/action cycle.
This is the manually started English Workshop booking test on 2026-09-21, seed 43129, attempt 20260921T120532066-91fc95. It is one attempt per model, not an average or a multi-seed reliability claim. No model was rerun to replace its result.
Recording
Recording excerpt, seed 43129 (26.51 seconds of GIF playback). The clip ends with Astra showing “Booking saved” while its final decision is still pending. Final verified outcomes and elapsed times come from the saved run logs and the result screenshot; GIF playback duration is not benchmark execution time.
Final verified result:
Test conditions
| Item | Recorded condition |
|---|---|
| Hardware | Apple M3 Max, 36 GB unified memory |
| Operating system | macOS 26.5.2 |
| Browser execution | agent-browser 0.37.1, three isolated sessions |
| Viewport | 1000 × 820 pixels per session |
| Execution order | Three models run concurrently on separate copies of the same form |
| Local inference | NanoJev-Web browser-head-v5, PyTorch MPS, FP32 |
| Jev identity | typesafe/jev-1.13-20260917 |
| Astra identity | openai/gpt-6-astra, high reasoning |
| Display pauses | 0 ms; no artificial action delay in this attempt |
| Start | Manual Run click, followed by a six-second countdown |
| Limits | 100 decisions; 300 seconds checked between steps; 60-second model-request timeout |
| Repair policy | No scripted repair, teacher fallback, substitute model or automatic retry |
All three models started with identical decision inputs, verified by hashes. Each model received a stateless abstract page state and the same action-description format. Candidate sets then evolved independently as their actions changed their own sessions.
The host supplied selectors, assigned form values and success conditions. These remained in the executor; models did not receive the full DOM, screenshots, field labels, URLs or literal values. The field names used in this report identify the host-side targets. The decision inputs described actions using operation, purpose, current match state, assignment, viewport, enabled and blocked flags. No previous-action history was included.
This evaluates the model-plus-executor system's ability to choose and perform bounded actions. It does not measure autonomous website discovery or the ability to infer form values from an open-ended request.
Results
| Model | Outcome | Elapsed | Valid fields | Exact assigned values | Decisions / executed actions | Input / output tokens | API cost (USD) |
|---|---|---|---|---|---|---|---|
| Jev | action_cycle; not saved |
8.7733 s | 6/15 (40.0%) | 4/15 (26.7%) | 11 / 10 | 13,307 / 1,916 | $0.000558894 |
| NanoJev-Web | completed; saved |
19.8501 s | 15/15 (100%) | 15/15 (100%) | 21 / 20 | 52,743 / 0 | $0 local API |
| GPT-6 Astra High | completed; saved |
79.1565 s | 15/15 (100%) | 15/15 (100%) | 20 / 19 | 14,661 / 812 | $0.187210000 |
All scored model requests have recorded usage; there are no unknown-usage requests in this attempt. Combined reported hosted API cost was $0.187768894. This excludes the coordinating assistant, preparation, hardware, electricity and earlier runs. NanoJev-Web encodes candidate paths separately, so its token count includes repeated context and is not directly comparable to provider billing. Astra's 812 output tokens include 505 reported reasoning tokens; these are not additional tokens to add to that total.
Elapsed time measures the action loop, including observation, model requests, execution and status updates. Navigation/setup, countdown and final audit/screenshots are excluded. With zero display pauses, elapsed and “without pauses” are equal. Failed-run duration is time until stopping, not time to successful completion.
| Component | Jev | NanoJev-Web | Astra High |
|---|---|---|---|
| Model HTTP request wall time | 5.1236 s | 13.0616 s | 71.8182 s |
| Browser command wall time in the loop | 3.6104 s | 6.7261 s | 7.2716 s |
| Other loop overhead | 0.0393 s | 0.0624 s | 0.0668 s |
| Browser calls in the loop | 85 | 170 | 163 |
HTTP time includes service/network overhead and is not a pure neural inference measurement. All sessions and live streaming shared the same machine; this was not an isolated hardware latency benchmark.
The exact Jev cycle
Jev first filled Notes → Contact name → Project website → Time correctly. From decision 5 onward, it repeatedly chose scrolling over available input actions:
| Decision | Selected action | Target location before action | Selected probability | Execution |
|---|---|---|---|---|
| 5 | scroll to Status notifications (#notifications) |
below | 0.38 | Executed |
| 6 | scroll to I confirm these are test data (#consent) |
above | 0.25 | Executed |
| 7 | scroll to Status notifications (#notifications) |
below | 0.29 | Executed |
| 8 | scroll to I confirm these are test data (#consent) |
above | 0.24 | Executed |
| 9 | scroll to Status notifications (#notifications) |
below | 0.39 | Executed |
| 10 | scroll to I confirm these are test data (#consent) |
above | 0.28 | Executed |
| 11 | scroll to Status notifications (#notifications) |
below | 0.31 | Rejected before execution |
The observed loop was:
Upper form: consent is visible and can be checked
↓ scroll to #notifications
Lower form: notifications are visible and can be enabled
↑ scroll back to #consent
Repeat without changing either control
There were six executed scroll actions, forming three complete down/up loops. At decision 11, Jev selected the fourth downward scroll to #notifications. The repeated-state guard rejected that action before calling the browser. Therefore the trace contains 11 model decisions but only 10 executed actions: four fills plus six scrolls.
The guard counts visits to a key formed from the fresh host state, operation, target selector, target match state and viewport position. It stops when a key is selected more than three times, including nonconsecutive repetitions. It is not a “three consecutive identical clicks” rule. The successful terminal DONE decisions for the other models are also not browser actions.
The useful actions were available
At decisions 5, 7, 9 and 11, the upper-page inputs and ordered candidates were exactly identical. Candidate a2 offered check on the visible, enabled, unblocked consent checkbox. Jev instead selected a0, the scroll to the offscreen notifications switch.
At decisions 6, 8 and 10, the lower-page inputs and ordered candidates were exactly identical. Candidate a0 offered click on the visible, enabled, unblocked notifications switch. Jev instead selected a2, the scroll back to consent above the viewport.
For example, at decision 6 the returned distribution assigned 0.25 to scrolling up and 0.07 to clicking the notifications switch. The harness used the selected choice without a confidence cutoff, consistently with the common comparison policy. These reported probabilities are not calibrated estimates of task success, and action IDs such as a0 are local to each observation.
The trace establishes that scrolling changed which controls were in view, the corresponding local input actions remained available, and no additional values were completed during the loop. There was no provider error, input-parser error, recorded keyboard error, invalid submission or permission block. Jev did not attempt to click either loop target and fail; it kept choosing another scroll.
What caused it, and what is still uncertain
The immediate observed failure was action selection: the returned rankings favored another offscreen target over useful controls already in view. The interface was stateless and did not provide action history, so the two identical recurring inputs did not carry an explicit reminder that this navigation had already failed to make progress. That is a plausible contributor, not a proven explanation of the model's internal reasoning. The trace does not show private model reasoning.
This does not establish that Jev is unable to operate checkboxes, switches or forms in general. NanoJev-Web and Astra completed the same fixture through the same executor. A separate experiment could test action-history summaries or scroll-after-arrival constraints, but neither was added to rescue this measured attempt. The weights, prompt and action policy were not changed in response to this failure.
Field-by-field verification
All values below are synthetic test data. An exact value match is stricter than HTML validity.
| Field | Assigned value | Jev final value | Jev valid / exact | NanoJev-Web | Astra High |
|---|---|---|---|---|---|
| Project website | https://form-632.example.test | https://form-632.example.test | Yes / Yes | Exact, valid | Exact, valid |
| I confirm these are test data | true | false | No / No | Exact, valid | Exact, valid |
| Notes | Synthetic booking 35548. Prepare the workshop materials. | Synthetic booking 35548. Prepare the workshop materials. | Yes / Yes | Exact, valid | Exact, valid |
| Support team | coral | (empty) | No / No | Exact, valid | Exact, valid |
| Contact name | Test Contact Gamma / sample 185 | Test Contact Gamma / sample 185 | Yes / Yes | Exact, valid | Exact, valid |
| Time | 18:00 | 18:00 | Yes / Yes | Exact, valid | Exact, valid |
| Priority | low | (empty) | No / No | Exact, valid | Exact, valid |
| Extras | recording, screen | (none selected) | No / No | Exact, valid | Exact, valid |
| Date | 2030-11-16 | (empty) | No / No | Exact, valid | Exact, valid |
| Receive marketing emails | false | true | Yes / No | Exact, valid | Exact, valid |
| Email address | form.2652@example.test | (empty) | No / No | Exact, valid | Exact, valid |
| Location | north | (empty) | No / No | Exact, valid | Exact, valid |
| Phone number | +12025550197 | (empty) | No / No | Exact, valid | Exact, valid |
| Status notifications | true | false | Yes / No | Exact, valid | Exact, valid |
| Participants | 37 | (empty) | No / No | Exact, valid | Exact, valid |
Jev's two additional valid fields were unchanged defaults: marketing emails remained true and notifications remained false. Both were valid boolean states, but both contradicted the assigned values. This explains 6 valid fields versus only 4 exact values; the report does not count valid defaults as completed task data.
Persistence and workflow audit
| Check | Jev | NanoJev-Web | Astra High |
|---|---|---|---|
| Server submissions | 0 | 1 | 1 |
| Exactly one submission with all assigned values | No | Yes | Yes |
| Server accepted all field constraints | No submission | Yes | Yes |
| Visible “Booking saved” success | No | Yes | Yes |
| Optional/reset/cancel actions executed | 0 | 0 | 0 |
| Invalid advance/review/submission attempts | 0 | 0 | 0 |
| Recorded unidentified keyboard events | 0 | 0 | 0 |
| Whole workflow passed | No | Yes | Yes |
A model's DONE selection alone is not sufficient to pass. Both successful attempts passed the DOM match checks, independent server validation, exact saved-data comparison, single-submission check and workflow-event checks.
The bench also recorded a Stop request at the overall job level. Individual attempt outcomes are preserved exactly as recorded: Jev action_cycle, NanoJev-Web completed, and Astra completed. The job-level flag is not treated as a new model failure or as the cause of Jev's earlier cycle.
Complete decision sequence
Each cell identifies the model-selected operation and host-side target. DONE confirms visible success and is not an executed browser input. Jev's last selected scroll was rejected by the cycle guard.
| Decision | Jev | NanoJev-Web | Astra High |
|---|---|---|---|
| 1 | fill Notes |
fill Notes |
fill Notes |
| 2 | fill Contact name |
open Support team |
check I confirm these are test data |
| 3 | fill Project website |
check I confirm these are test data |
fill Contact name |
| 4 | fill Time |
fill Project website |
open Support team |
| 5 | scroll Status notifications |
fill Contact name |
fill Project website |
| 6 | scroll I confirm these are test data |
fill Time |
fill Time |
| 7 | scroll Status notifications |
option Coral |
option Coral |
| 8 | scroll I confirm these are test data |
radio Low |
radio Low |
| 9 | scroll Status notifications |
scroll Extras |
scroll Status notifications |
| 10 | scroll I confirm these are test data |
multiselect Extras |
click Status notifications |
| 11 | scroll Status notifications (blocked by cycle guard) |
fill Date |
fill Email address |
| 12 | — | fill Email address |
uncheck Receive marketing emails |
| 13 | — | uncheck Receive marketing emails |
fill Date |
| 14 | — | fill Phone number |
multiselect Extras |
| 15 | — | click Status notifications |
select Location |
| 16 | — | select Location |
fill Participants |
| 17 | — | scroll Participants |
fill Phone number |
| 18 | — | fill Participants |
click Review booking |
| 19 | — | click Review booking |
click Save booking |
| 20 | — | click Save booking |
DONE — visible success |
| 21 | — | DONE — visible success |
— |
Recorded protocol and evidence
The recording used seed 43129, no display pauses, three fresh browser sessions and the same initial decision inputs. The original fixture, comparison server and cloud adapters are not bundled with this model release. Repeating this exact three-model experiment requires that separate demonstration application; the release itself provides local inference and a runner for your own page contracts.
Fixture SHA-256: 03e4713e07448df49bd9c5f56830ef41275e20c32299a64ed5beb9b638cb898b. Protocol SHA-256: 6fa80a82af021e45f8446d0980636af5b619941601862b41b582e6c15e12dd6c. Weight SHA-256: 0a93cfcf7121112fdf90336bda1198ad67a3bb45247da6eedf0dd8672d4d53b5.
The two repeated Jev input hashes are:
- Upper-page decisions 5/7/9/11:
805e8da7365a5b6529811bd317b36cedf7b51c2d00bdbcceceb44c27a224d4bd. - Lower-page decisions 6/8/10:
0447597c79b267d77718dee96eef6b43a40ec43e516c0cddf67df33989444600.
Input hashes use UTF-8 JSON with preserved key order and compact separators, matching JSON.stringify for these payloads. The detailed machine-readable report includes all 52 decisions, model-facing inputs, selected actions, usage and independent final audits. The compact result contains the headline measurements. The repeated-state guard used for this attempt is described above; executable comparison code is excluded from this release.
This release report contains only synthetic fixture data and allowlisted evidence. Local filesystem paths, credentials, browser profiles and private provider identifiers are omitted. The comparison is a single specialized action-selection test and does not establish general browser-agent superiority. Earlier comparisons are not merged into this recorded attempt.

