NanoJev-Web / docs /TRAINING.md
candypunk's picture
Release NanoJev-Web browser action model
a0a9254 verified
|
Raw
History Blame Contribute Delete
3.65 kB

Model lineage and browser head training

Lineage

  • Decision model: C-Tianyu/NanoJev, revision unified-games-v1.
  • Source: TianyuCodings/NanoJev, commit 618cea6d906d54e128360786d12f703fff2b1245.
  • Backbone: Qwen/Qwen3-0.6B, revision c1899de289a04d12100db370d81485cdf75e47ca.
  • Original checkpoint SHA-256: f68c47d66998231b86b7e91b4ed5e82ae23acf104c8b7cd6d165c3ac7b7ffe1b.
  • Released browser checkpoint SHA-256: 0a93cfcf7121112fdf90336bda1198ad67a3bb45247da6eedf0dd8672d4d53b5.

Each candidate path combines state, question, candidate text and a terminal EOS token. The backbone's final candidate representation is normalized and scored, with an attention-based set head mixing candidate information. Inference applies softmax across the supplied alternatives. There is no autoregressive decoding and no generated text to parse in the NanoJev-Web API.

The set head uses a 1024-wide LayerNorm, a scalar projection, a 1025-to-128 set projection, four-head attention with width 128, and a final 128-to-1 projection. All 200,578 head parameters were trainable; all 596,049,920 backbone parameters were frozen. The frozen-backbone comparison and changed-head audit are retained in weight-audit.json.

Synthetic browser data

The shipped JSONL files contain 1,863 abstract synthetic states, not collected user forms, browsing histories, names, email addresses, passwords or conversations. Labels define acceptable next actions in a known action schema. Some states admit more than one acceptable action.

Split Rows Generator seed Correct acceptable-action choices
Train 1,251 731 Not reported as a held-out result
Development 192 1709 191/192 (99.48%)
Calibration 192 2081 189/192 (98.44%)
Test 228 3253 228/228 (100%)

The development process included 51 feedback-derived examples and 240 small-page curriculum examples. Exact cross-split input duplicates were absent in the recorded data audit. The splits share a synthetic state grammar and action families; they do not test unfamiliar website reasoning. Development feedback influenced iteration. “Calibration” is a dataset split name; release probabilities have not been certified as calibrated confidence estimates.

Optimization recipe

  • Frozen backbone feature extraction; head optimization on cached representations.
  • AdamW, learning rate 0.0007, weight decay 0.005, batch size 24.
  • Gradient norm clipping at 1.0, seed 42, 200 epochs.
  • Soft-target cross-entropy over candidate choices.
  • Best checkpoint chosen by development accuracy, then development cross-entropy; selected epoch 191.
  • Final held-out test evaluated after checkpoint selection.

The recorded 58.14 seconds cover cached head optimization, not feature extraction, previous experiments or total development. Full numeric results are in training-results.json. Training data and the inference architecture are published for inspection; this package does not include the old game trainer or its feature caches. It is a runnable inference release, not a claim of a fully reproduced upstream training environment.

Distribution cleanup did not change the weights. The reference browser interface explicitly distinguishes unfinished inputs from available advance/submit actions and supplies only executable candidates. A recorded browser demonstration is documented separately from model-only synthetic accuracy.