Results

Both reconstructions through one protocol: half the frames in each group fit a 5-feature least squares, half are scored; pooled ρ, five seeds, beside a within-group label shuffle. The second number is the comparable one — the share of the floor-to-perfect distance the arm covered, so a high floor cannot cap it. Only the image→gradient step differs.

Presses the sensor images whole

datasetnLUT
ρ / of what was left
calibration-free
ρ / of what was left
shuffle floorgroups fitted
GelSight Mini, CNC presses, 0-20 N605
of 6,219
0.993
+0.992
0.998
+0.998
+0.067 / +0.0896/6
FoTa cnc_Mini604
of 3,351
0.900
+0.900
0.938
+0.936
-0.040 / +0.0356/6
FEATS (MARKER gel)2,000
of 16,276
0.642
+0.607
0.706
+0.677
+0.089 / +0.0898/8
Sparsh / Meta, 10 gel pads2,000
of 129,389
0.959
+0.957
0.986
+0.986
+0.050 / +0.0304/4
FeelAnyForce
floor-dominated
2,000
of 13,892
0.939
+0.764
0.949
+0.803
+0.742 / +0.74113/14

The one marker gel is the one low row. FEATS is the only gel here with a printed dot lattice — 63 dots counted on its references against 1 or fewer on the other 4 — and it holds the lowest calibration-free ρ. Its panel shows why: the dots emboss themselves into the reconstructed surface. One markered dataset is one data point, so this is an observation, not a controlled comparison.

truncated presses
A press is truncated when its contact core reaches a border: the free-boundary solve then runs off the edge with nothing to stop the ramp, and its depth is not identifiable from the image. 15.9% reconstruct deeper than the 4.25 mm gel, against 0.0% of whole presses (441 and 59 frames).

Excluding them is what the headline buys: calibration-free scores ρ 0.998 on whole presses against 0.851 with truncated frames mixed in. Two datasets cannot reach the 2,000 this table samples; they have no more presses to give.

All frames

datasetnLUT
ρ / of what was left
calibration-free
ρ / of what was left
shuffle floorgroups fitted
GelSight Mini, CNC presses, 0-20 N2,376
of 6,219
0.707
+0.661
0.851
+0.827
+0.135 / +0.1386/6
FoTa cnc_Mini2,239
of 3,351
0.442
+0.427
0.482
+0.464
+0.026 / +0.0336/6
FEATS (MARKER gel)2,956
of 16,276
0.707
+0.642
0.702
+0.629
+0.183 / +0.1968/8
Sparsh / Meta, 10 gel pads3,091
of 129,389
0.669
+0.656
0.698
+0.685
+0.040 / +0.0414/4
FeelAnyForce3,378
of 13,892
0.905
+0.843
0.935
+0.891
+0.396 / +0.40414/14

React's production number adds a fitted position gain field and lives on the method page.

predicted vs ground-truth force
Held-out prediction against ground truth, shared axes per row, each panel annotated as the table is.
cross-dataset transfer
Fit on one dataset, predict on every other. Read each cell against the random-weight baseline under its column: the features are collinear and monotone in contact size, so on an easy target almost any direction ranks.

FoTa cnc_Mini→FEATS, Sparsh→FEATS are ≥99 % extrapolation. Their MAE measures extrapolation, not prediction. FeelAnyForce's row goes negative: collinear features let least squares cancel opposite-sign terms (method). Non-negative weights fix it: off-diagonal ρ 0.574 → 0.731, negative cells 3 → 0 of 20, costing 0.010 on the diagonal. The deployed estimator is unchanged: on React both agree at ρ 0.989 (1.8 % of frames outside the rig's range), and 15 held-out seeds differ by +0.002 ± 0.014 ρ.

Which reconstruction for React's force channel?

React's own calibration objects cannot answer this: calibration-free scores ρ 0.781 against the LUT's 0.763 on 158 held-out presses, but a paired bootstrap puts the margin at 95% CI [-0.081, +0.120] — a coin flip. The table is ahead on all five once each row's floor is corrected for, but that is a ranking over five other sensors, not a test on React. It ships because it needs no per-sensor lookup table.

The two agree at ρ = 0.925 over 2,400 React frames, mean difference 0.86 N. Published to yxma/React: this channel across all 72 sides of 36 episodes (480,080 frames).

Error analysis

The ten worst held-out frames reconstruct as well as the five best — same dipoles, same compact depth, no ramping. The residual is in the depth→force fit, not in image→depth, so a better reconstruction will not move them.

Relative error is |pred−true| over the dataset's force span.

cnc_mini_26 errors
GelSight Mini CNC, span 19.50 N — median 0.7%, p90 2.7%, worst 14.1%.
cnc errors
FoTa cnc_Mini, span 4.01 N — median 5.0%, p90 13.8%, worst 49.5%.
feats errors
FEATS (marker), span 32.74 N — median 5.4%, p90 15.0%, worst 79.7%.
sparsh errors
Sparsh / Meta, span 1.03 N — median 1.9%, p90 7.2%, worst 37.0%.
faf errors
FeelAnyForce, span 17.44 N — median 1.8%, p90 9.0%, worst 37.3%.