NanoJev-Web / docs /SINGLE_FORM_REPORT.md
candypunk's picture
Release NanoJev-Web browser action model
a0a9254 verified
|
Raw
History Blame Contribute Delete
15.9 kB

Recorded browser demonstration: Jev, NanoJev-Web and GPT-6 Astra High

This is supporting documentation for NanoJev-Web. The comparison application used to record it is separate from the model distribution and is not included here. Local model inference and browser execution do not require a comparison provider.

Observed result: NanoJev-Web and Astra completed and saved all 15 assigned values. Jev filled four values, then alternated between two scroll targets until the harness stopped the repeated state/action cycle.

This is the manually started English Workshop booking test on 2026-09-21, seed 43129, attempt 20260921T120532066-91fc95. It is one attempt per model, not an average or a multi-seed reliability claim. No model was rerun to replace its result.

Recording

Live recording of Jev, NanoJev-Web and GPT-6 Astra High on the seed-43129 form

Recording excerpt, seed 43129 (26.51 seconds of GIF playback). The clip ends with Astra showing “Booking saved” while its final decision is still pending. Final verified outcomes and elapsed times come from the saved run logs and the result screenshot; GIF playback duration is not benchmark execution time.

Final verified result:

Final state: Jev stopped, NanoJev-Web passed, and GPT-6 Astra passed

Test conditions

Item Recorded condition
Hardware Apple M3 Max, 36 GB unified memory
Operating system macOS 26.5.2
Browser execution agent-browser 0.37.1, three isolated sessions
Viewport 1000 × 820 pixels per session
Execution order Three models run concurrently on separate copies of the same form
Local inference NanoJev-Web browser-head-v5, PyTorch MPS, FP32
Jev identity typesafe/jev-1.13-20260917
Astra identity openai/gpt-6-astra, high reasoning
Display pauses 0 ms; no artificial action delay in this attempt
Start Manual Run click, followed by a six-second countdown
Limits 100 decisions; 300 seconds checked between steps; 60-second model-request timeout
Repair policy No scripted repair, teacher fallback, substitute model or automatic retry

All three models started with identical decision inputs, verified by hashes. Each model received a stateless abstract page state and the same action-description format. Candidate sets then evolved independently as their actions changed their own sessions.

The host supplied selectors, assigned form values and success conditions. These remained in the executor; models did not receive the full DOM, screenshots, field labels, URLs or literal values. The field names used in this report identify the host-side targets. The decision inputs described actions using operation, purpose, current match state, assignment, viewport, enabled and blocked flags. No previous-action history was included.

This evaluates the model-plus-executor system's ability to choose and perform bounded actions. It does not measure autonomous website discovery or the ability to infer form values from an open-ended request.

Results

Model Outcome Elapsed Valid fields Exact assigned values Decisions / executed actions Input / output tokens API cost (USD)
Jev action_cycle; not saved 8.7733 s 6/15 (40.0%) 4/15 (26.7%) 11 / 10 13,307 / 1,916 $0.000558894
NanoJev-Web completed; saved 19.8501 s 15/15 (100%) 15/15 (100%) 21 / 20 52,743 / 0 $0 local API
GPT-6 Astra High completed; saved 79.1565 s 15/15 (100%) 15/15 (100%) 20 / 19 14,661 / 812 $0.187210000

All scored model requests have recorded usage; there are no unknown-usage requests in this attempt. Combined reported hosted API cost was $0.187768894. This excludes the coordinating assistant, preparation, hardware, electricity and earlier runs. NanoJev-Web encodes candidate paths separately, so its token count includes repeated context and is not directly comparable to provider billing. Astra's 812 output tokens include 505 reported reasoning tokens; these are not additional tokens to add to that total.

Elapsed time measures the action loop, including observation, model requests, execution and status updates. Navigation/setup, countdown and final audit/screenshots are excluded. With zero display pauses, elapsed and “without pauses” are equal. Failed-run duration is time until stopping, not time to successful completion.

Component Jev NanoJev-Web Astra High
Model HTTP request wall time 5.1236 s 13.0616 s 71.8182 s
Browser command wall time in the loop 3.6104 s 6.7261 s 7.2716 s
Other loop overhead 0.0393 s 0.0624 s 0.0668 s
Browser calls in the loop 85 170 163

HTTP time includes service/network overhead and is not a pure neural inference measurement. All sessions and live streaming shared the same machine; this was not an isolated hardware latency benchmark.

The exact Jev cycle

Jev first filled Notes → Contact name → Project website → Time correctly. From decision 5 onward, it repeatedly chose scrolling over available input actions:

Decision Selected action Target location before action Selected probability Execution
5 scroll to Status notifications (#notifications) below 0.38 Executed
6 scroll to I confirm these are test data (#consent) above 0.25 Executed
7 scroll to Status notifications (#notifications) below 0.29 Executed
8 scroll to I confirm these are test data (#consent) above 0.24 Executed
9 scroll to Status notifications (#notifications) below 0.39 Executed
10 scroll to I confirm these are test data (#consent) above 0.28 Executed
11 scroll to Status notifications (#notifications) below 0.31 Rejected before execution

The observed loop was:

Upper form: consent is visible and can be checked
    ↓ scroll to #notifications
Lower form: notifications are visible and can be enabled
    ↑ scroll back to #consent
Repeat without changing either control

There were six executed scroll actions, forming three complete down/up loops. At decision 11, Jev selected the fourth downward scroll to #notifications. The repeated-state guard rejected that action before calling the browser. Therefore the trace contains 11 model decisions but only 10 executed actions: four fills plus six scrolls.

The guard counts visits to a key formed from the fresh host state, operation, target selector, target match state and viewport position. It stops when a key is selected more than three times, including nonconsecutive repetitions. It is not a “three consecutive identical clicks” rule. The successful terminal DONE decisions for the other models are also not browser actions.

The useful actions were available

At decisions 5, 7, 9 and 11, the upper-page inputs and ordered candidates were exactly identical. Candidate a2 offered check on the visible, enabled, unblocked consent checkbox. Jev instead selected a0, the scroll to the offscreen notifications switch.

At decisions 6, 8 and 10, the lower-page inputs and ordered candidates were exactly identical. Candidate a0 offered click on the visible, enabled, unblocked notifications switch. Jev instead selected a2, the scroll back to consent above the viewport.

For example, at decision 6 the returned distribution assigned 0.25 to scrolling up and 0.07 to clicking the notifications switch. The harness used the selected choice without a confidence cutoff, consistently with the common comparison policy. These reported probabilities are not calibrated estimates of task success, and action IDs such as a0 are local to each observation.

The trace establishes that scrolling changed which controls were in view, the corresponding local input actions remained available, and no additional values were completed during the loop. There was no provider error, input-parser error, recorded keyboard error, invalid submission or permission block. Jev did not attempt to click either loop target and fail; it kept choosing another scroll.

What caused it, and what is still uncertain

The immediate observed failure was action selection: the returned rankings favored another offscreen target over useful controls already in view. The interface was stateless and did not provide action history, so the two identical recurring inputs did not carry an explicit reminder that this navigation had already failed to make progress. That is a plausible contributor, not a proven explanation of the model's internal reasoning. The trace does not show private model reasoning.

This does not establish that Jev is unable to operate checkboxes, switches or forms in general. NanoJev-Web and Astra completed the same fixture through the same executor. A separate experiment could test action-history summaries or scroll-after-arrival constraints, but neither was added to rescue this measured attempt. The weights, prompt and action policy were not changed in response to this failure.

Field-by-field verification

All values below are synthetic test data. An exact value match is stricter than HTML validity.

Field Assigned value Jev final value Jev valid / exact NanoJev-Web Astra High
Project website https://form-632.example.test https://form-632.example.test Yes / Yes Exact, valid Exact, valid
I confirm these are test data true false No / No Exact, valid Exact, valid
Notes Synthetic booking 35548. Prepare the workshop materials. Synthetic booking 35548. Prepare the workshop materials. Yes / Yes Exact, valid Exact, valid
Support team coral (empty) No / No Exact, valid Exact, valid
Contact name Test Contact Gamma / sample 185 Test Contact Gamma / sample 185 Yes / Yes Exact, valid Exact, valid
Time 18:00 18:00 Yes / Yes Exact, valid Exact, valid
Priority low (empty) No / No Exact, valid Exact, valid
Extras recording, screen (none selected) No / No Exact, valid Exact, valid
Date 2030-11-16 (empty) No / No Exact, valid Exact, valid
Receive marketing emails false true Yes / No Exact, valid Exact, valid
Email address form.2652@example.test (empty) No / No Exact, valid Exact, valid
Location north (empty) No / No Exact, valid Exact, valid
Phone number +12025550197 (empty) No / No Exact, valid Exact, valid
Status notifications true false Yes / No Exact, valid Exact, valid
Participants 37 (empty) No / No Exact, valid Exact, valid

Jev's two additional valid fields were unchanged defaults: marketing emails remained true and notifications remained false. Both were valid boolean states, but both contradicted the assigned values. This explains 6 valid fields versus only 4 exact values; the report does not count valid defaults as completed task data.

Persistence and workflow audit

Check Jev NanoJev-Web Astra High
Server submissions 0 1 1
Exactly one submission with all assigned values No Yes Yes
Server accepted all field constraints No submission Yes Yes
Visible “Booking saved” success No Yes Yes
Optional/reset/cancel actions executed 0 0 0
Invalid advance/review/submission attempts 0 0 0
Recorded unidentified keyboard events 0 0 0
Whole workflow passed No Yes Yes

A model's DONE selection alone is not sufficient to pass. Both successful attempts passed the DOM match checks, independent server validation, exact saved-data comparison, single-submission check and workflow-event checks.

The bench also recorded a Stop request at the overall job level. Individual attempt outcomes are preserved exactly as recorded: Jev action_cycle, NanoJev-Web completed, and Astra completed. The job-level flag is not treated as a new model failure or as the cause of Jev's earlier cycle.

Complete decision sequence

Each cell identifies the model-selected operation and host-side target. DONE confirms visible success and is not an executed browser input. Jev's last selected scroll was rejected by the cycle guard.

Decision Jev NanoJev-Web Astra High
1 fill Notes fill Notes fill Notes
2 fill Contact name open Support team check I confirm these are test data
3 fill Project website check I confirm these are test data fill Contact name
4 fill Time fill Project website open Support team
5 scroll Status notifications fill Contact name fill Project website
6 scroll I confirm these are test data fill Time fill Time
7 scroll Status notifications option Coral option Coral
8 scroll I confirm these are test data radio Low radio Low
9 scroll Status notifications scroll Extras scroll Status notifications
10 scroll I confirm these are test data multiselect Extras click Status notifications
11 scroll Status notifications (blocked by cycle guard) fill Date fill Email address
12 — fill Email address uncheck Receive marketing emails
13 — uncheck Receive marketing emails fill Date
14 — fill Phone number multiselect Extras
15 — click Status notifications select Location
16 — select Location fill Participants
17 — scroll Participants fill Phone number
18 — fill Participants click Review booking
19 — click Review booking click Save booking
20 — click Save booking DONE — visible success
21 — DONE — visible success —

Recorded protocol and evidence

The recording used seed 43129, no display pauses, three fresh browser sessions and the same initial decision inputs. The original fixture, comparison server and cloud adapters are not bundled with this model release. Repeating this exact three-model experiment requires that separate demonstration application; the release itself provides local inference and a runner for your own page contracts.

Fixture SHA-256: 03e4713e07448df49bd9c5f56830ef41275e20c32299a64ed5beb9b638cb898b. Protocol SHA-256: 6fa80a82af021e45f8446d0980636af5b619941601862b41b582e6c15e12dd6c. Weight SHA-256: 0a93cfcf7121112fdf90336bda1198ad67a3bb45247da6eedf0dd8672d4d53b5.

The two repeated Jev input hashes are:

  • Upper-page decisions 5/7/9/11: 805e8da7365a5b6529811bd317b36cedf7b51c2d00bdbcceceb44c27a224d4bd.
  • Lower-page decisions 6/8/10: 0447597c79b267d77718dee96eef6b43a40ec43e516c0cddf67df33989444600.

Input hashes use UTF-8 JSON with preserved key order and compact separators, matching JSON.stringify for these payloads. The detailed machine-readable report includes all 52 decisions, model-facing inputs, selected actions, usage and independent final audits. The compact result contains the headline measurements. The repeated-state guard used for this attempt is described above; executable comparison code is excluded from this release.

This release report contains only synthetic fixture data and allowlisted evidence. Local filesystem paths, credentials, browser profiles and private provider identifiers are omitted. The comparison is a single specialized action-selection test and does not establish general browser-agent superiority. Earlier comparisons are not merged into this recorded attempt.