# Results Generated by `scripts/make_results.py` from the JSON written by `train.py`, `evaluate.py`, `probe.py` and `play.py`. Do not edit by hand. ## Baseline: what NanoJev reports for itself From `docs/DEVELOPMENT_RESULTS.md` in the NanoJev repository. | NanoJev task | reported acc | always-true control | above control | | --- | ---: | ---: | :---: | | `scaled_maze` (test) | 0.5625 | 0.5625 | no | | `scaled_maze` (ood) | 0.5625 | 0.5625 | no | | `snake_one_step_safety_v4` | 0.9220 | 0.9220 | no | Closed-loop play, same source, 128-step limit: | maze | Jev | NanoJev | reference | | --- | ---: | ---: | ---: | | all three mazes | 0/128 | 0/128 | solved 1/16, 1/24, 1/96 | > "The learned direct-action models and Jev do not solve any of these > three mazes within the shared 128-step limit." Every accuracy in this file is therefore printed next to its constant-prediction control, and gameplay is reported as solve rate rather than per-question accuracy. ## Does the planner actually compute distances? Ridge probe from the planner's per-cell features to true BFS distance, on held-out mazes. `diameter` is the longest true shortest-path in those boards; `T` is the iteration budget. `descent` is the fraction of cells whose lowest-predicted neighbour lies on a true shortest path. `reach` is the fraction from which greedy descent actually arrives at the goal, and it is the one that predicts solve rate: a field can be right 70% of the time per step and still trap most walks in a spurious basin, so `descent` is an upper bound on nothing in particular. R² scores absolute values; only `reach` scores the behaviour. `read-out R²` scores `predict_field()`, the field this controller actually descends -- the min-plus recurrence when the run enabled it, the conv head otherwise. The training curve's `conv field R²` scores a different head; see the note there. | size | cells | probe R² | probe MAE (cells) | read-out R² | descent | reach | diameter | T | | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | 11 | 215 | **0.0937** | 7.97 | 1.0000 | 1.0000 | **1.0000** | 48 | 82 | | 21 | 901 | **0.0420** | 29.71 | 1.0000 | 1.0000 | **1.0000** | 152 | 262 | | 31 | 2011 | **0.0316** | 63.06 | 1.0000 | 1.0000 | **1.0000** | 310 | 542 | | 51 | 5679 | **0.0068** | 145.23 | 1.0000 | 1.0000 | **1.0000** | 672 | 1402 | ## Question-level accuracy `uniform` is the accuracy of guessing uniformly over the same candidates. A model that matches its control has learned nothing, which is the failure NanoJev's own log records for its maze model. `boards` is how many distinct boards the row's questions came from, and it is there so that `n` can be trusted. Questions that share a board are not independent tests of whether the model can read a board -- they share walls, goal and distance field, so they succeed or fail together -- and a row reporting `n` alone cannot tell you whether that is happening. Here it is not: the generator draws a fresh board per state, so `boards` tracks `n` on every split and `n` is the honest denominator. A row where the two diverge is one whose precision should be read off `boards`. ### val — 1325 questions, ECE 0.0865 | group | n | boards | accuracy | uniform control | TVD | Brier | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | `maze/action` | 97 | 97 | **1.0000** | 0.4905 | 0.0188 | 0.0023 | | `maze/boolean` | 530 | 106 | **1.0000** | 0.5000 | 0.0007 | 0.0000 | | `maze/choice` | 97 | 97 | **1.0000** | 0.4905 | 0.0188 | 0.0023 | | `maze/clear` | 424 | 106 | **1.0000** | 0.5000 | 0.0004 | 0.0000 | | `maze/distance` | 106 | 106 | **0.3962** | 0.1429 | 0.7764 | 0.7640 | | `maze/score` | 106 | 106 | **0.3962** | 0.1429 | 0.7764 | 0.7640 | | `maze/solvable` | 106 | 106 | **1.0000** | 0.5000 | 0.0018 | 0.0000 | | `snake/action` | 74 | 74 | **0.9865** | 0.4955 | 0.3500 | 0.3148 | | `snake/boolean` | 444 | 74 | **0.8829** | 0.5000 | 0.2338 | 0.2022 | | `snake/choice` | 74 | 74 | **0.9865** | 0.4955 | 0.3500 | 0.3148 | | `snake/escape` | 222 | 74 | **0.8108** | 0.5000 | 0.2773 | 0.2891 | | `snake/room` | 74 | 74 | **0.2838** | 0.2000 | 0.7990 | 0.7981 | | `snake/safe` | 222 | 74 | **0.9550** | 0.5000 | 0.1903 | 0.1152 | | `snake/score` | 74 | 74 | **0.2838** | 0.2000 | 0.7990 | 0.7981 | ### test — 1765 questions, ECE 0.0590 | group | n | boards | accuracy | uniform control | TVD | Brier | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | `maze/action` | 141 | 140 | **1.0000** | 0.4882 | 0.0396 | 0.0079 | | `maze/boolean` | 740 | 147 | **1.0000** | 0.5000 | 0.0007 | 0.0000 | | `maze/choice` | 141 | 140 | **1.0000** | 0.4882 | 0.0396 | 0.0079 | | `maze/clear` | 592 | 147 | **1.0000** | 0.5000 | 0.0004 | 0.0000 | | `maze/distance` | 148 | 147 | **0.2973** | 0.1429 | 0.7847 | 0.7796 | | `maze/score` | 148 | 147 | **0.2973** | 0.1429 | 0.7847 | 0.7796 | | `maze/solvable` | 148 | 147 | **1.0000** | 0.5000 | 0.0019 | 0.0000 | | `snake/action` | 92 | 92 | **0.9674** | 0.4783 | 0.3107 | 0.2770 | | `snake/boolean` | 552 | 92 | **0.8895** | 0.5000 | 0.2223 | 0.1938 | | `snake/choice` | 92 | 92 | **0.9674** | 0.4783 | 0.3107 | 0.2770 | | `snake/escape` | 276 | 92 | **0.8188** | 0.5000 | 0.2628 | 0.2747 | | `snake/room` | 92 | 92 | **0.2500** | 0.2000 | 0.7993 | 0.7986 | | `snake/safe` | 276 | 92 | **0.9601** | 0.5000 | 0.1818 | 0.1128 | | `snake/score` | 92 | 92 | **0.2500** | 0.2000 | 0.7993 | 0.7986 | ### ood — 1167 questions, ECE 0.0540 | group | n | boards | accuracy | uniform control | TVD | Brier | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | `maze/action` | 107 | 107 | **1.0000** | 0.5132 | 0.2699 | 0.1504 | | `maze/boolean` | 550 | 110 | **1.0000** | 0.5000 | 0.0007 | 0.0000 | | `maze/choice` | 107 | 107 | **1.0000** | 0.5132 | 0.2699 | 0.1504 | | `maze/clear` | 440 | 110 | **1.0000** | 0.5000 | 0.0004 | 0.0000 | | `maze/distance` | 110 | 110 | **0.0545** | 0.1429 | 0.8542 | 0.9160 | | `maze/score` | 110 | 110 | **0.0545** | 0.1429 | 0.8542 | 0.9160 | | `maze/solvable` | 110 | 110 | **1.0000** | 0.5000 | 0.0020 | 0.0000 | | `snake/action` | 50 | 50 | **0.9800** | 0.4867 | 0.3574 | 0.3199 | | `snake/boolean` | 300 | 50 | **0.8333** | 0.5000 | 0.2636 | 0.2621 | | `snake/choice` | 50 | 50 | **0.9800** | 0.4867 | 0.3574 | 0.3199 | | `snake/escape` | 150 | 50 | **0.6800** | 0.5000 | 0.3561 | 0.4471 | | `snake/room` | 50 | 50 | **0.2000** | 0.2000 | 0.7989 | 0.7978 | | `snake/safe` | 150 | 50 | **0.9867** | 0.5000 | 0.1710 | 0.0771 | | `snake/score` | 50 | 50 | **0.2000** | 0.2000 | 0.7989 | 0.7978 | ## Closed-loop play Read `model` and `model-field` as answering different questions. `model` is the network's action head alone -- no search, no memory, no visited set -- and is the apples-to-apples comparison with NanoJev. `model-field` descends a min-plus recurrence whose fixed point is a shortest-path distance by construction, so it arrives from anywhere; for a maze it does so at initialisation, before any training, because uniform cost is already the right answer. Its solve rate is a property of the architecture, not a measurement of what this run learned. ### Maze `deadlock` is the fraction of episodes in which the controller spent a tenth of its step budget walking into the same wall, and `cells` is how many distinct squares it stood on. They separate two failures a solve rate of 0.00 reports identically: standing still and touring the board. Deadlock should now read 0.00 for every controller, because only legal moves are offered as candidates -- it is kept as a regression guard, not a finding. The failure that remains is cycling: the planner never sees the agent, so a greedy policy that steps A->B->A has no state with which to notice, and `cells` is what exposes it. | controller | solve rate | mean steps | efficiency | collisions | deadlock | cells | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | `maze_model` | 1.00 | 93.2 | 1.000 | 0.0 | 0.00 | 94.2 | | `maze_model_sampled` | 1.00 | 106.1 | 0.926 | 0.0 | 0.00 | 95.4 | | `maze_model_field` | 1.00 | 93.2 | 1.000 | 0.0 | 0.00 | 94.2 | | `maze_model_memory` | 1.00 | 93.2 | 1.000 | 0.0 | 0.00 | 94.2 | | `maze_random_memory` | 0.83 | 689.7 | 0.363 | 0.0 | 0.00 | 188.5 | | `maze_reference` | 1.00 | 93.2 | 1.000 | 0.0 | 0.00 | 94.2 | | `maze_random` | 0.14 | 560.4 | 0.068 | 0.0 | 0.00 | 92.2 | ### Snake | controller | mean food | max food | mean steps | survival | | --- | ---: | ---: | ---: | ---: | | `snake_model` | 4.50 | 16 | 288.0 | 1.00 | | `snake_reference` | 23.50 | 31 | 288.0 | 1.00 | | `snake_random` | 0.67 | 2 | 33.8 | 0.00 | ## Training curve Frozen validation split, never sampled during training. `conv field R²` scores `supervised_field()` -- the conv head the field loss trains. It is **not** the field the controller descends. Under `--min-plus-field` those are different objects: the read-out is the min-plus recurrence, scored as `read-out R²` in the probe table above, and nothing at inference reads the conv head at all. It stays in the loss as an auxiliary task on the shared trunk, so a negative value here means that auxiliary head has stopped tracking the trunk the action head and the recurrence are shaping -- it does not mean the planner is wrong, and `reach` is where that would show up. | step | loss | acc | choice | boolean | score | conv field R² | maze action | snake action | | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | 500 | 0.6430 | 0.8310 | 0.6806 | 0.9469 | 0.3500 | 0.5048 | 0.5773 | 0.8108 | | 1000 | 0.6269 | 0.8626 | 0.8469 | 0.9469 | 0.4222 | -2.7876 | 0.7732 | 0.9459 | | 1500 | 0.6161 | 0.8393 | 0.6665 | 0.9469 | 0.4222 | 0.0461 | 0.7835 | 0.5135 | | 2000 | 0.5041 | 0.8703 | 0.9833 | 0.9469 | 0.3500 | -0.1305 | 1.0000 | 0.9595 | | 2500 | 0.4842 | 0.8531 | 0.8499 | 0.9469 | 0.3500 | 0.1714 | 1.0000 | 0.6486 | | 3000 | 0.4821 | 0.8718 | 0.9944 | 0.9469 | 0.3500 | -0.2088 | 1.0000 | 0.9865 | | 3500 | 0.4725 | 0.8718 | 0.9944 | 0.9469 | 0.3500 | 0.0103 | 1.0000 | 0.9865 | | 4000 | 0.4797 | 0.8702 | 0.9828 | 0.9469 | 0.3500 | 0.2466 | 0.9794 | 0.9865 | | 4500 | 0.4765 | 0.8710 | 0.9889 | 0.9469 | 0.3500 | 0.2097 | 0.9897 | 0.9865 | | 5000 | 0.4733 | 0.8718 | 0.9944 | 0.9469 | 0.3500 | 0.1148 | 1.0000 | 0.9865 | | 5500 | 0.4667 | 0.8725 | 0.9944 | 0.9469 | 0.3556 | -0.2881 | 1.0000 | 0.9865 | | 6000 | 0.4579 | 0.8673 | 0.9595 | 0.9469 | 0.3500 | -0.1574 | 1.0000 | 0.9054 | | 6500 | 0.4543 | 0.8695 | 0.9767 | 0.9469 | 0.3500 | 0.2626 | 1.0000 | 0.9459 | | 7000 | 0.4258 | 0.8860 | 0.9889 | 0.9673 | 0.3500 | 0.0325 | 1.0000 | 0.9730 | | 7500 | 0.4088 | 0.8831 | 0.9368 | 0.9673 | 0.3778 | -0.4171 | 1.0000 | 0.8514 | | 8000 | 0.4030 | 0.8860 | 0.9889 | 0.9673 | 0.3500 | 0.0676 | 1.0000 | 0.9730 | | 8500 | 0.3999 | 0.8853 | 0.9889 | 0.9673 | 0.3444 | 0.2261 | 1.0000 | 0.9730 | | 9000 | 0.4021 | 0.8958 | 0.9889 | 0.9673 | 0.4222 | 0.1083 | 1.0000 | 0.9730 |