Model lineage and browser head training
Lineage
- Decision model: C-Tianyu/NanoJev, revision
unified-games-v1. - Source: TianyuCodings/NanoJev, commit
618cea6d906d54e128360786d12f703fff2b1245. - Backbone: Qwen/Qwen3-0.6B, revision
c1899de289a04d12100db370d81485cdf75e47ca. - Original checkpoint SHA-256:
f68c47d66998231b86b7e91b4ed5e82ae23acf104c8b7cd6d165c3ac7b7ffe1b. - Released browser checkpoint SHA-256:
0a93cfcf7121112fdf90336bda1198ad67a3bb45247da6eedf0dd8672d4d53b5.
Each candidate path combines state, question, candidate text and a terminal EOS token. The backbone's final candidate representation is normalized and scored, with an attention-based set head mixing candidate information. Inference applies softmax across the supplied alternatives. There is no autoregressive decoding and no generated text to parse in the NanoJev-Web API.
The set head uses a 1024-wide LayerNorm, a scalar projection, a 1025-to-128 set projection, four-head attention with width 128, and a final 128-to-1 projection. All 200,578 head parameters were trainable; all 596,049,920 backbone parameters were frozen. The frozen-backbone comparison and changed-head audit are retained in weight-audit.json.
Synthetic browser data
The shipped JSONL files contain 1,863 abstract synthetic states, not collected user forms, browsing histories, names, email addresses, passwords or conversations. Labels define acceptable next actions in a known action schema. Some states admit more than one acceptable action.
| Split | Rows | Generator seed | Correct acceptable-action choices |
|---|---|---|---|
| Train | 1,251 | 731 | Not reported as a held-out result |
| Development | 192 | 1709 | 191/192 (99.48%) |
| Calibration | 192 | 2081 | 189/192 (98.44%) |
| Test | 228 | 3253 | 228/228 (100%) |
The development process included 51 feedback-derived examples and 240 small-page curriculum examples. Exact cross-split input duplicates were absent in the recorded data audit. The splits share a synthetic state grammar and action families; they do not test unfamiliar website reasoning. Development feedback influenced iteration. “Calibration” is a dataset split name; release probabilities have not been certified as calibrated confidence estimates.
Optimization recipe
- Frozen backbone feature extraction; head optimization on cached representations.
- AdamW, learning rate
0.0007, weight decay0.005, batch size24. - Gradient norm clipping at
1.0, seed42,200epochs. - Soft-target cross-entropy over candidate choices.
- Best checkpoint chosen by development accuracy, then development cross-entropy; selected epoch
191. - Final held-out test evaluated after checkpoint selection.
The recorded 58.14 seconds cover cached head optimization, not feature extraction, previous experiments or total development. Full numeric results are in training-results.json. Training data and the inference architecture are published for inspection; this package does not include the old game trainer or its feature caches. It is a runnable inference release, not a claim of a fully reproduced upstream training environment.
Distribution cleanup did not change the weights. The reference browser interface explicitly distinguishes unfinished inputs from available advance/submit actions and supplies only executable candidates. A recorded browser demonstration is documented separately from model-only synthetic accuracy.