File size: 15,900 Bytes
a0a9254 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 | # Recorded browser demonstration: Jev, NanoJev-Web and GPT-6 Astra High
This is supporting documentation for NanoJev-Web. The comparison application used to record it is separate from the model distribution and is not included here. Local model inference and browser execution do not require a comparison provider.
**Observed result:** NanoJev-Web and Astra completed and saved all 15 assigned values. Jev filled four values, then alternated between two scroll targets until the harness stopped the repeated state/action cycle.
This is the manually started English **Workshop booking** test on **2026-09-21**, seed **43129**, attempt `20260921T120532066-91fc95`. It is one attempt per model, not an average or a multi-seed reliability claim. No model was rerun to replace its result.
## Recording

*Recording excerpt, seed 43129 (26.51 seconds of GIF playback). The clip ends with Astra showing “Booking saved” while its final decision is still pending. Final verified outcomes and elapsed times come from the saved run logs and the result screenshot; GIF playback duration is not benchmark execution time.*
**Final verified result:**

## Test conditions
| Item | Recorded condition |
|---|---|
| Hardware | Apple M3 Max, 36 GB unified memory |
| Operating system | macOS 26.5.2 |
| Browser execution | agent-browser 0.37.1, three isolated sessions |
| Viewport | 1000 × 820 pixels per session |
| Execution order | Three models run concurrently on separate copies of the same form |
| Local inference | NanoJev-Web browser-head-v5, PyTorch MPS, FP32 |
| Jev identity | `typesafe/jev-1.13-20260917` |
| Astra identity | `openai/gpt-6-astra`, high reasoning |
| Display pauses | 0 ms; no artificial action delay in this attempt |
| Start | Manual Run click, followed by a six-second countdown |
| Limits | 100 decisions; 300 seconds checked between steps; 60-second model-request timeout |
| Repair policy | No scripted repair, teacher fallback, substitute model or automatic retry |
All three models started with **identical decision inputs**, verified by hashes. Each model received a stateless abstract page state and the same action-description format. Candidate sets then evolved independently as their actions changed their own sessions.
The host supplied selectors, assigned form values and success conditions. These remained in the executor; models did not receive the full DOM, screenshots, field labels, URLs or literal values. The field names used in this report identify the host-side targets. The decision inputs described actions using operation, purpose, current match state, assignment, viewport, enabled and blocked flags. No previous-action history was included.
This evaluates the model-plus-executor system's ability to choose and perform bounded actions. It does not measure autonomous website discovery or the ability to infer form values from an open-ended request.
## Results
| Model | Outcome | Elapsed | Valid fields | Exact assigned values | Decisions / executed actions | Input / output tokens | API cost (USD) |
|---|---|---:|---:|---:|---:|---:|---:|
| Jev | `action_cycle`; not saved | 8.7733 s | 6/15 (40.0%) | 4/15 (26.7%) | 11 / 10 | 13,307 / 1,916 | $0.000558894 |
| NanoJev-Web | `completed`; saved | 19.8501 s | 15/15 (100%) | 15/15 (100%) | 21 / 20 | 52,743 / 0 | $0 local API |
| GPT-6 Astra High | `completed`; saved | 79.1565 s | 15/15 (100%) | 15/15 (100%) | 20 / 19 | 14,661 / 812 | $0.187210000 |
All scored model requests have recorded usage; there are no unknown-usage requests in this attempt. Combined reported hosted API cost was **$0.187768894**. This excludes the coordinating assistant, preparation, hardware, electricity and earlier runs. NanoJev-Web encodes candidate paths separately, so its token count includes repeated context and is not directly comparable to provider billing. Astra's 812 output tokens include **505 reported reasoning tokens**; these are not additional tokens to add to that total.
Elapsed time measures the action loop, including observation, model requests, execution and status updates. Navigation/setup, countdown and final audit/screenshots are excluded. With zero display pauses, elapsed and “without pauses” are equal. Failed-run duration is time until stopping, not time to successful completion.
| Component | Jev | NanoJev-Web | Astra High |
|---|---:|---:|---:|
| Model HTTP request wall time | 5.1236 s | 13.0616 s | 71.8182 s |
| Browser command wall time in the loop | 3.6104 s | 6.7261 s | 7.2716 s |
| Other loop overhead | 0.0393 s | 0.0624 s | 0.0668 s |
| Browser calls in the loop | 85 | 170 | 163 |
HTTP time includes service/network overhead and is not a pure neural inference measurement. All sessions and live streaming shared the same machine; this was not an isolated hardware latency benchmark.
## The exact Jev cycle
Jev first filled **Notes → Contact name → Project website → Time** correctly. From decision 5 onward, it repeatedly chose scrolling over available input actions:
| Decision | Selected action | Target location before action | Selected probability | Execution |
|---:|---|---|---:|---|
| 5 | `scroll` to **Status notifications** (`#notifications`) | below | 0.38 | Executed |
| 6 | `scroll` to **I confirm these are test data** (`#consent`) | above | 0.25 | Executed |
| 7 | `scroll` to **Status notifications** (`#notifications`) | below | 0.29 | Executed |
| 8 | `scroll` to **I confirm these are test data** (`#consent`) | above | 0.24 | Executed |
| 9 | `scroll` to **Status notifications** (`#notifications`) | below | 0.39 | Executed |
| 10 | `scroll` to **I confirm these are test data** (`#consent`) | above | 0.28 | Executed |
| 11 | `scroll` to **Status notifications** (`#notifications`) | below | 0.31 | **Rejected before execution** |
The observed loop was:
```text
Upper form: consent is visible and can be checked
↓ scroll to #notifications
Lower form: notifications are visible and can be enabled
↑ scroll back to #consent
Repeat without changing either control
```
There were **six executed scroll actions**, forming **three complete down/up loops**. At decision **11**, Jev selected the fourth downward scroll to `#notifications`. The repeated-state guard rejected that action before calling the browser. Therefore the trace contains **11 model decisions but only 10 executed actions**: four fills plus six scrolls.
The guard counts visits to a key formed from the fresh host state, operation, target selector, target match state and viewport position. It stops when a key is selected more than three times, including nonconsecutive repetitions. It is not a “three consecutive identical clicks” rule. The successful terminal `DONE` decisions for the other models are also not browser actions.
### The useful actions were available
At decisions **5, 7, 9 and 11**, the upper-page inputs and ordered candidates were exactly identical. Candidate `a2` offered `check` on the visible, enabled, unblocked consent checkbox. Jev instead selected `a0`, the scroll to the offscreen notifications switch.
At decisions **6, 8 and 10**, the lower-page inputs and ordered candidates were exactly identical. Candidate `a0` offered `click` on the visible, enabled, unblocked notifications switch. Jev instead selected `a2`, the scroll back to consent above the viewport.
For example, at decision 6 the returned distribution assigned **0.25** to scrolling up and **0.07** to clicking the notifications switch. The harness used the selected choice without a confidence cutoff, consistently with the common comparison policy. These reported probabilities are not calibrated estimates of task success, and action IDs such as `a0` are local to each observation.
The trace establishes that scrolling changed which controls were in view, the corresponding local input actions remained available, and no additional values were completed during the loop. There was **no provider error, input-parser error, recorded keyboard error, invalid submission or permission block**. Jev did not attempt to click either loop target and fail; it kept choosing another scroll.
### What caused it, and what is still uncertain
The immediate observed failure was **action selection**: the returned rankings favored another offscreen target over useful controls already in view. The interface was stateless and did not provide action history, so the two identical recurring inputs did not carry an explicit reminder that this navigation had already failed to make progress. That is a plausible contributor, not a proven explanation of the model's internal reasoning. The trace does not show private model reasoning.
This does not establish that Jev is unable to operate checkboxes, switches or forms in general. NanoJev-Web and Astra completed the same fixture through the same executor. A separate experiment could test action-history summaries or scroll-after-arrival constraints, but neither was added to rescue this measured attempt. The weights, prompt and action policy were not changed in response to this failure.
## Field-by-field verification
All values below are synthetic test data. An exact value match is stricter than HTML validity.
| Field | Assigned value | Jev final value | Jev valid / exact | NanoJev-Web | Astra High |
|---|---|---|---|---|---|
| Project website | https://form-632.example.test | https://form-632.example.test | Yes / Yes | Exact, valid | Exact, valid |
| I confirm these are test data | true | false | No / No | Exact, valid | Exact, valid |
| Notes | Synthetic booking 35548. Prepare the workshop materials. | Synthetic booking 35548. Prepare the workshop materials. | Yes / Yes | Exact, valid | Exact, valid |
| Support team | coral | (empty) | No / No | Exact, valid | Exact, valid |
| Contact name | Test Contact Gamma / sample 185 | Test Contact Gamma / sample 185 | Yes / Yes | Exact, valid | Exact, valid |
| Time | 18:00 | 18:00 | Yes / Yes | Exact, valid | Exact, valid |
| Priority | low | (empty) | No / No | Exact, valid | Exact, valid |
| Extras | recording, screen | (none selected) | No / No | Exact, valid | Exact, valid |
| Date | 2030-11-16 | (empty) | No / No | Exact, valid | Exact, valid |
| Receive marketing emails | false | true | Yes / No | Exact, valid | Exact, valid |
| Email address | form.2652@example.test | (empty) | No / No | Exact, valid | Exact, valid |
| Location | north | (empty) | No / No | Exact, valid | Exact, valid |
| Phone number | +12025550197 | (empty) | No / No | Exact, valid | Exact, valid |
| Status notifications | true | false | Yes / No | Exact, valid | Exact, valid |
| Participants | 37 | (empty) | No / No | Exact, valid | Exact, valid |
Jev's two additional valid fields were unchanged defaults: marketing emails remained **true** and notifications remained **false**. Both were valid boolean states, but both contradicted the assigned values. This explains **6 valid fields versus only 4 exact values**; the report does not count valid defaults as completed task data.
## Persistence and workflow audit
| Check | Jev | NanoJev-Web | Astra High |
|---|---:|---:|---:|
| Server submissions | 0 | 1 | 1 |
| Exactly one submission with all assigned values | No | Yes | Yes |
| Server accepted all field constraints | No submission | Yes | Yes |
| Visible “Booking saved” success | No | Yes | Yes |
| Optional/reset/cancel actions executed | 0 | 0 | 0 |
| Invalid advance/review/submission attempts | 0 | 0 | 0 |
| Recorded unidentified keyboard events | 0 | 0 | 0 |
| Whole workflow passed | No | Yes | Yes |
A model's `DONE` selection alone is not sufficient to pass. Both successful attempts passed the DOM match checks, independent server validation, exact saved-data comparison, single-submission check and workflow-event checks.
The bench also recorded a **Stop request at the overall job level**. Individual attempt outcomes are preserved exactly as recorded: Jev `action_cycle`, NanoJev-Web `completed`, and Astra `completed`. The job-level flag is not treated as a new model failure or as the cause of Jev's earlier cycle.
## Complete decision sequence
Each cell identifies the model-selected operation and host-side target. `DONE` confirms visible success and is not an executed browser input. Jev's last selected scroll was rejected by the cycle guard.
| Decision | Jev | NanoJev-Web | Astra High |
|---:|---|---|---|
| 1 | `fill` Notes | `fill` Notes | `fill` Notes |
| 2 | `fill` Contact name | `open` Support team | `check` I confirm these are test data |
| 3 | `fill` Project website | `check` I confirm these are test data | `fill` Contact name |
| 4 | `fill` Time | `fill` Project website | `open` Support team |
| 5 | `scroll` Status notifications | `fill` Contact name | `fill` Project website |
| 6 | `scroll` I confirm these are test data | `fill` Time | `fill` Time |
| 7 | `scroll` Status notifications | `option` Coral | `option` Coral |
| 8 | `scroll` I confirm these are test data | `radio` Low | `radio` Low |
| 9 | `scroll` Status notifications | `scroll` Extras | `scroll` Status notifications |
| 10 | `scroll` I confirm these are test data | `multiselect` Extras | `click` Status notifications |
| 11 | `scroll` Status notifications **(blocked by cycle guard)** | `fill` Date | `fill` Email address |
| 12 | — | `fill` Email address | `uncheck` Receive marketing emails |
| 13 | — | `uncheck` Receive marketing emails | `fill` Date |
| 14 | — | `fill` Phone number | `multiselect` Extras |
| 15 | — | `click` Status notifications | `select` Location |
| 16 | — | `select` Location | `fill` Participants |
| 17 | — | `scroll` Participants | `fill` Phone number |
| 18 | — | `fill` Participants | `click` Review booking |
| 19 | — | `click` Review booking | `click` Save booking |
| 20 | — | `click` Save booking | `DONE` — visible success |
| 21 | — | `DONE` — visible success | — |
## Recorded protocol and evidence
The recording used seed **43129**, no display pauses, three fresh browser sessions and the same initial decision inputs. The original fixture, comparison server and cloud adapters are not bundled with this model release. Repeating this exact three-model experiment requires that separate demonstration application; the release itself provides local inference and a runner for your own page contracts.
Fixture SHA-256: `03e4713e07448df49bd9c5f56830ef41275e20c32299a64ed5beb9b638cb898b`. Protocol SHA-256: `6fa80a82af021e45f8446d0980636af5b619941601862b41b582e6c15e12dd6c`. Weight SHA-256: `0a93cfcf7121112fdf90336bda1198ad67a3bb45247da6eedf0dd8672d4d53b5`.
The two repeated Jev input hashes are:
- Upper-page decisions 5/7/9/11: `805e8da7365a5b6529811bd317b36cedf7b51c2d00bdbcceceb44c27a224d4bd`.
- Lower-page decisions 6/8/10: `0447597c79b267d77718dee96eef6b43a40ec43e516c0cddf67df33989444600`.
Input hashes use UTF-8 JSON with preserved key order and compact separators, matching `JSON.stringify` for these payloads. The [detailed machine-readable report](demo/decisions.json) includes all 52 decisions, model-facing inputs, selected actions, usage and independent final audits. The [compact result](demo/measurements.json) contains the headline measurements. The repeated-state guard used for this attempt is described above; executable comparison code is excluded from this release.
This release report contains only synthetic fixture data and allowlisted evidence. Local filesystem paths, credentials, browser profiles and private provider identifiers are omitted. The comparison is a single specialized action-selection test and does not establish general browser-agent superiority. Earlier comparisons are not merged into this recorded attempt.
|