--- language: - en license: other license_name: mit-and-apache-2.0 license_link: LICENSE library_name: pytorch base_model: C-Tianyu/NanoJev base_model_relation: finetune tags: - browser-automation - agent-browser - action-selection - local-inference - apple-silicon - safetensors - nanojev - jev --- # NanoJev-Web **Open-weight browser hands for a larger planning agent.** NanoJev-Web is a lightweight browser-action model derived from NanoJev. It runs locally and selects the next action from a host-provided set of permitted browser operations. The included **agent-browser** runtime executes the selected action and checks the resulting page state. Use it as an execution component for form filling and bounded web workflows. A larger model or your application supplies the goal, selectors, values and success conditions; NanoJev-Web handles action selection within that contract. The weights, inference code and browser runtime are available for local integration. No hosted model service or API key is required. ## How it fits into an agent ```text Planning agent / application supplies the task, allowed controls, values and success conditions ↓ Host observer → abstract page state + available action candidates ↓ NanoJev-Web → selected action ID + candidate probabilities ↓ Browser executor → performs the action → observes the new page state ↺ Host verifies the outcome or handles a blocked workflow ``` The model scores supplied alternatives without generating text. Selectors and literal form values stay with the supplied executor; it sends abstract action descriptions and page-state flags to the model. Planning, page understanding and independent outcome verification belong to the host. [Integration guide and API](docs/INTEGRATION.md). ## Inheritance and browser specialization NanoJev-Web inherits its decision architecture and tokenizer from [C-Tianyu/NanoJev](https://huggingface.co/C-Tianyu/NanoJev), using the **Qwen3-0.6B** backbone. All **596,049,920 backbone parameters remain frozen**. The **200,578-parameter decision head** was further trained on synthetic browser action states. The released weights are the verified `browser-head-v5` checkpoint. Packaging it as NanoJev-Web did not introduce another training run. Game interfaces, game datasets, game training commands and teacher-service integrations are excluded. Removing those components does not erase knowledge from the inherited weights. [Training details](docs/TRAINING.md) · [Provenance](PROVENANCE.json). | Property | Value | |---|---| | Total parameters | 596,250,498 | | Output | One supplied action ID and probabilities over the alternatives | | Storage | Full FP32 safetensors checkpoint; no quantization or adapter | | Checkpoint size | 2,385,039,280 bytes | | Candidate context limit | 768 tokens per separately encoded candidate path; overflows rejected | | Local inference | PyTorch MPS on Apple Silicon; explicit CPU option | | Tested hardware | MacBook Pro, Apple M2 Pro, 16 GB unified memory | ## Available browser actions The supplied executor supports: - **Inputs:** `fill`, `type`, `clear`, `select`, `multiselect`, `check`, `uncheck`, `radio`, `upload`. - **Interactions:** `click`, `hover`, `dblclick`, `drag`, `focus`, `press`. - **Structured controls and navigation:** `open`, `option`, `add`, `remove`, `link`, `tab`. - **Visibility:** observer-generated `scroll` to reveal an allowed target. - **Control decisions:** `DONE`, `WAIT`, `BLOCKED`. These are executor capabilities exposed as permitted choices. The model does not invent commands. Same-origin frame targeting is available through a control's `frame` property. See the [complete action reference](docs/ACTIONS.md) for required data, semantics and boundaries. ## Run locally The supplied installer targets **native Apple Silicon, macOS 14+, Node.js 24+ with npm, and uv**. It installs pinned Python 3.12.13 dependencies and agent-browser 0.37.1. Installation downloads dependencies; model inference then uses local files. Other Apple Silicon generations have not all been benchmarked. **Tested on MacBook Pro with Apple M2 Pro and 16 GB unified memory.** Successful operation was reported in a separate functional test. [Hardware specifications and test scope](docs/INTEGRATION.md#apple-m2-pro-compatibility). ```bash cd NanoJev-Web chmod +x *.command ./setup.command ./start.command ``` The model listens on **127.0.0.1:8774**. In another terminal: ```bash curl http://127.0.0.1:8774/api/health curl -H 'Content-Type: application/json' \ --data-binary @examples/choice-request.json \ http://127.0.0.1:8774/api/evaluate ``` For browser execution, adapt [examples/browser-task.json](examples/browser-task.json) to your own page, selectors and desired values, then run: ```bash ./run.command --task your-task.json ``` The runner creates an isolated agent-browser session and stores its local trace and final screenshot in `.local/runs/`. The example describes a page contract; it does not start or include a demo website. For a Python-only model setup or your own executor, see [integration](docs/INTEGRATION.md). This is a custom decision checkpoint. Load it with the supplied inference code; the repository root is not a standard `AutoModelForCausalLM` chat checkpoint or a configured Hugging Face Inference widget. ## Evaluation and limits On held-out synthetic decision states, the browser head selected an acceptable action in **228/228** cases. Development scored **191/192** and calibration **189/192**. These splits share the training schema; they measure action selection within that schema, not success on arbitrary websites. Probabilities have not been certified as calibrated confidence. [Data and methodology](docs/TRAINING.md). The current model and runtime target bounded workflows with supplied controls. They do not independently discover a site's workflow, invent form values, interpret arbitrary screenshots, manage credentials or make business decisions. The host must define success and verify any required saved data; a `DONE` choice alone does not establish business correctness. Changing the observation/action schema can change model behavior and requires evaluation. ## Browser demonstration The following recording illustrates the model in use. In one 15-field form attempt, NanoJev-Web completed and saved all assigned values in **19.85 seconds**, using **20 browser actions and 21 decisions** on Apple M3 Max / 36 GB. Jev and GPT-6 Astra High ran alongside it through the same browser executor for context. ![Final recorded browser demonstration: Jev stopped, NanoJev-Web and GPT-6 Astra passed](docs/assets/browser-demo.png) ![Browser demonstration with NanoJev-Web, Jev and GPT-6 Astra High](docs/assets/seed-43129.gif) *Recording excerpt, seed 43129 (26.51 seconds of GIF playback). The clip ends with Astra showing “Booking saved” while its final decision is still pending. Final verified outcomes and elapsed times come from the saved run logs and result screenshot; GIF playback duration is not execution time.* This is one observed demonstration with supplied selectors and values, not a general browser reliability claim. The [detailed recording report](docs/SINGLE_FORM_REPORT.md) retains the comparison, timing, token accounting and verification evidence. The comparison application and its cloud-provider integrations are separate from this model distribution; only static documentation assets are included. ## Release contents and privacy The distribution contains the checkpoint, tokenizer, inference API, agent-browser executor, integration examples, synthetic browser-head data, model checks, documentation and license notices. The model's inherited **MIT and Apache-2.0** terms are described in [LICENSE](LICENSE). Private machine paths, credentials, browser profiles and raw local run files are excluded. The safetensors header has no author metadata, and head training used abstract synthetic states. This audit does not establish that inherited pretrained weights contain no personal information. [Privacy audit](docs/PRIVACY.md) · [Publishing instructions](docs/PUBLISHING.md).