72 commanded action sequences over 6 start frames, for scoring a tactile
world model's rollouts against ground truth that is geometric, not photometric.
Nobody performed these motions, so there is no ground-truth future image — what
is ground truth is where the sensor would be, and its projection into each
camera. Full method and usage in the
README; per-probe clips are under
/probes/.
Yellow: commanded
ground truth. Red: a deliberately wrong rollout, offset 25 mm in world x —
it reads 18–19 px against a ~6 px noise floor. Dimmed: the hand that
must stay still.
72probes · 6 start frames
0.11–0.38 mtranslation · 23–90° rotation
p40–p79speed vs the dataset
~6 pxworst reprojection noise floor
0.133 mclosest the hands come (rule 0.12)
Start frames are drawn from held-out intervals of splits.json — never training frames. Eligible sessions: 2026-05-10, 2026-05-11, 2026-05-19. This draw happened to select 2026-05-10, 2026-05-11; 2026-05-19 was not sampled this time, which is chance, not a judgement about it. 2026-05-19 had its OptiTrack world redefined mid-collection; the release applies a translation-only correction and the residual yaw about the table normal is unmeasured (about 16 px at the workspace). It is included with that stated in world_residual rather than dropped — a bounded, declared error is not a reason to discard a fifth of the sessions.
run0 — 2026-05-11/episode_007, rows 1322–1325 moving right, holding left