# Model lineage and browser head training ## Lineage - Decision model: [C-Tianyu/NanoJev](https://huggingface.co/C-Tianyu/NanoJev), revision `unified-games-v1`. - Source: [TianyuCodings/NanoJev](https://github.com/TianyuCodings/NanoJev), commit `618cea6d906d54e128360786d12f703fff2b1245`. - Backbone: [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B), revision `c1899de289a04d12100db370d81485cdf75e47ca`. - Original checkpoint SHA-256: `f68c47d66998231b86b7e91b4ed5e82ae23acf104c8b7cd6d165c3ac7b7ffe1b`. - Released browser checkpoint SHA-256: `0a93cfcf7121112fdf90336bda1198ad67a3bb45247da6eedf0dd8672d4d53b5`. Each candidate path combines state, question, candidate text and a terminal EOS token. The backbone's final candidate representation is normalized and scored, with an attention-based set head mixing candidate information. Inference applies softmax across the supplied alternatives. There is no autoregressive decoding and no generated text to parse in the NanoJev-Web API. The set head uses a 1024-wide LayerNorm, a scalar projection, a 1025-to-128 set projection, four-head attention with width 128, and a final 128-to-1 projection. All **200,578 head parameters** were trainable; all **596,049,920 backbone parameters** were frozen. The frozen-backbone comparison and changed-head audit are retained in [weight-audit.json](../evals/weight-audit.json). ## Synthetic browser data The shipped JSONL files contain **1,863 abstract synthetic states**, not collected user forms, browsing histories, names, email addresses, passwords or conversations. Labels define acceptable next actions in a known action schema. Some states admit more than one acceptable action. | Split | Rows | Generator seed | Correct acceptable-action choices | |---|---:|---:|---:| | Train | 1,251 | 731 | Not reported as a held-out result | | Development | 192 | 1709 | 191/192 (99.48%) | | Calibration | 192 | 2081 | 189/192 (98.44%) | | Test | 228 | 3253 | 228/228 (100%) | The development process included 51 feedback-derived examples and 240 small-page curriculum examples. Exact cross-split input duplicates were absent in the recorded data audit. The splits share a synthetic state grammar and action families; they do not test unfamiliar website reasoning. Development feedback influenced iteration. “Calibration” is a dataset split name; release probabilities have not been certified as calibrated confidence estimates. ## Optimization recipe - Frozen backbone feature extraction; head optimization on cached representations. - AdamW, learning rate `0.0007`, weight decay `0.005`, batch size `24`. - Gradient norm clipping at `1.0`, seed `42`, `200` epochs. - Soft-target cross-entropy over candidate choices. - Best checkpoint chosen by development accuracy, then development cross-entropy; selected epoch `191`. - Final held-out test evaluated after checkpoint selection. The recorded **58.14 seconds** cover cached head optimization, not feature extraction, previous experiments or total development. Full numeric results are in [training-results.json](../evals/training-results.json). Training data and the inference architecture are published for inspection; this package does not include the old game trainer or its feature caches. It is a runnable inference release, not a claim of a fully reproduced upstream training environment. Distribution cleanup did not change the weights. The reference browser interface explicitly distinguishes unfinished inputs from available advance/submit actions and supplies only executable candidates. A recorded browser demonstration is documented separately from model-only synthetic accuracy.