NanoJev-Web / docs /INTEGRATION.md
candypunk's picture
Release runtime 1.1 with named workflows and hybrid by default
d786697 verified
|
Raw History Blame Contribute Delete
7.79 kB

Integrating NanoJev-Web

NanoJev-Web chooses among supplied actions. Your application or planning model owns the workflow, page interpretation, named targets or selectors, assigned values and verification. Runtime 1.1 adds a short staged request, exact-name target resolution, existing-session attachment and bounded BLOCKED recovery. You can use the bundled browser runner or call the local inference API from your own executor.

Use the browser runner

After setup.command, start the local model with start.command. Copy browser-task.json and replace the example URL, origins, selectors and values with those for your own authorized page. Then run:

./run.command --task your-task.json

The task contract has these fields:

Field Responsibility
name A simple run name, used for local report filenames
url Initial page to open
allowedOrigins Origins authorized for this task
controls Explicit targets, operations, desired values and optional conditions
success A condition describing the expected final page state

Each control has a unique id, an op, a label, and either a selector or a named target. Input actions also carry the required value or desired state. when makes a control conditional; expect can describe its postcondition. The action reference lists supported operations and conditions.

The observer reads the page through agent-browser, removes satisfied or non-executable candidates and offers scrolling for allowed offscreen targets. The model receives abstract state and action descriptions. The executor rechecks the observation before acting, handles permitted input mechanics, and stops on completion, blocking, action cycles or execution limits. These checks are part of the host runtime, not learned model behavior.

The runner uses the local endpoint by default. An explicit alternative local port can be selected with --endpoint http://127.0.0.1:PORT. A JSON trace and final screenshot are written under .local/runs/. They may contain the task's values and must remain private for private tasks. The runner checks visible success; your application must separately verify persistence and business rules where needed.

Call the model directly

GET /api/health reports readiness, device, precision and checkpoint identity. POST /api/evaluate accepts a JSON object containing states. Each state has a unique id, a state description and named questions.

Every question must contain:

  • type: exactly choice.
  • instructions: a nonempty string explaining the decision.
  • criteria: a mapping from candidate IDs to nonempty action descriptions.

Use choice-request.json as a concrete request:

curl -H 'Content-Type: application/json' \
  --data-binary @examples/choice-request.json \
  http://127.0.0.1:8774/api/evaluate

Read the selected ID at states[0].answers.next.choice and the distribution at states[0].answers.next.probabilities. Candidate IDs are local to each request. The response also contains checkpoint identity and execution accounting. execution.input_tokens sums separately encoded candidate paths, including repeated context; execution.output_tokens is zero because there is no autoregressive decoding.

The default limit is 768 tokens per complete candidate path, including state, instructions, candidate text and formatting. Oversized paths are rejected instead of truncated. A request supports 1–32 states, 2–255 alternatives per question, at most 96 questions and at most 256 total candidate paths. The HTTP server rejects request bodies over 2,000,000 bytes. The choice schema is broader than the browser training distribution: supplying arbitrary text does not establish reliable general-purpose reasoning.

For an existing agent framework, map its observed page state and permitted actions into the browser schema, call the model, execute the returned choice, then observe again. Preserve the semantics of busy/success state, unfinished fields/interactions, visibility, assignment and blocking flags. runtime/browser.mjs contains the reference state formatter and action descriptions. Keeping literal values and selectors in the host is a feature of this runtime; the low-level API does not itself remove private text supplied by a caller.

Apple M2 Pro compatibility

Tested on MacBook Pro with Apple M2 Pro and 16 GB unified memory. A contributor confirmed successful local operation of NanoJev-Web on this configuration. This is a reported functional check; no detailed trace or timing measurements were supplied. The CPU/GPU core counts and macOS version were not recorded.

Apple lists the following M2 Pro configurations for the 2023 MacBook Pro. These describe the hardware family, not the exact configuration of the reported test machine. Apple technical specifications.

Component M2 Pro specification
CPU / GPU 10-core CPU with 16-core GPU, or 12-core CPU with 19-core GPU
Unified memory 16 GB or 32 GB
Memory bandwidth 200 GB/s
OS requirement for this package macOS 14+

The supplied runtime uses the GPU through PyTorch MPS. Apple Silicon and macOS 14+ meet the documented platform requirements. The functional check covers the reported machine; it does not establish performance on every configuration. Apple's PyTorch/MPS guidance · PyTorch MPS documentation.

The full FP32 checkpoint occupies about 2.39 GB on disk, while inference also needs activations, temporary buffers, framework memory and browser processes. Candidate count and path length affect memory consumption. A 32 GB configuration provides more headroom than 16 GB, but no M2 Pro peak-memory measurement is available from this functional check. Running a larger planning model on the same Mac adds its own requirements.

The recorded demonstration was measured on M3 Max / 36 GB; its completion times remain attributed to that machine. The M2 Pro functional result is recorded separately in release-checks.json. After installation, use the package checks and a representative page contract to verify your environment. The installer checks torch.backends.mps.is_available() before declaring setup complete.

Python-only inference setup

If your application already supplies browser execution, the model server does not require Node.js or agent-browser. On native Apple Silicon with macOS 14+ and uv, from the release root:

export UV_PYTHON_INSTALL_DIR="$PWD/.runtime/python"
export UV_CACHE_DIR="$PWD/.runtime/cache"
uv python install 3.12.13 --no-bin
uv venv --managed-python --python 3.12.13 .venv
uv pip sync --python .venv/bin/python requirements-lock.txt
./start.command

The server loads the included checkpoint once and binds to loopback. To use CPU explicitly, run ./start.command --device cpu; CPU latency has not been benchmarked. No external model calls or teacher fallback are used. Installation needs network access for dependencies; inference uses local model files.

Verify the model package

./test.command

This runs the six request-contract tests and verifies the public file manifest. With the model server running, the frozen v5 reference decisions can also be checked:

.venv/bin/python tests/parity.py

Parity compares action choices and probabilities with saved reference outputs at a tolerance of 1e-4. These are model checks; they do not contact comparison models or run a comparison website.