Spaces:
Running
Running
| <html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>Results</title><style> | |
| :root{--bg:#0b1020;--fg:#e8eefb;--dim:#8ea0c2;--line:#1e2a45;--card:#111a2e; | |
| --accent:#ffc46b;--ok:#7be0a0;--bad:#ff8f7a; | |
| --s0:12px;--s1:14px;--s2:16px;--s3:20px;--s4:28px;--s5:40px} | |
| *{box-sizing:border-box} | |
| body{margin:0;background:var(--bg);color:var(--fg);font-size:var(--s2); | |
| font-family:'IBM Plex Sans',system-ui,sans-serif;line-height:1.6} | |
| .wrap{max-width:1100px;margin:0 auto;padding:0 20px 72px} | |
| a{color:var(--accent)} | |
| h1{font-size:var(--s5);line-height:1.15;margin:40px 0 8px;font-weight:650} | |
| h2{font-size:var(--s4);margin:44px 0 10px;font-weight:600} | |
| h3{font-size:var(--s3);margin:28px 0 6px;font-weight:600} | |
| p{margin:10px 0;max-width:70ch} | |
| .dim{color:var(--dim);font-size:var(--s1)} | |
| .bad{color:var(--bad);font-size:var(--s0)} | |
| nav{display:flex;gap:8px;flex-wrap:wrap;margin:18px 0 4px} | |
| nav a{display:inline-block;padding:10px 16px;min-height:44px;line-height:24px; | |
| border:1px solid var(--line);border-radius:999px;text-decoration:none; | |
| background:var(--card)} | |
| nav a[aria-current]{border-color:var(--accent)} | |
| /* built by other scripts and published beside this site */ | |
| nav a.ext{border-style:dashed} | |
| .cards{display:grid;grid-template-columns:repeat(auto-fit,minmax(200px,1fr)); | |
| gap:12px;margin:20px 0} | |
| .card{background:var(--card);border:1px solid var(--line);border-radius:12px; | |
| padding:16px} | |
| .card b{display:block;font-size:var(--s4);color:var(--accent);font-weight:650} | |
| .card span{color:var(--dim);font-size:var(--s1)} | |
| figure{margin:24px 0} | |
| figure img{width:100%;border-radius:10px;border:1px solid var(--line); | |
| background:#fff} | |
| figcaption{color:var(--dim);font-size:var(--s1);margin-top:8px;max-width:80ch} | |
| table{border-collapse:collapse;width:100%;margin:14px 0;font-size:var(--s1)} | |
| th,td{border-bottom:1px solid var(--line);padding:9px 10px;text-align:right} | |
| th:first-child,td:first-child{text-align:left} | |
| th{color:var(--dim);font-weight:500} | |
| td b{color:var(--ok)} | |
| details{background:var(--card);border:1px solid var(--line);border-radius:10px; | |
| padding:12px 16px;margin:16px 0} | |
| summary{cursor:pointer;color:var(--dim);font-size:var(--s1);min-height:44px; | |
| display:flex;align-items:center} | |
| code{background:#0d1526;padding:2px 6px;border-radius:5px;font-size:var(--s1)} | |
| /* Inline links in prose measured 47x16 and 54x16 at 375 px — a tap target, | |
| because a finger does not know it is "only prose". Padding alone would | |
| break the line box, so the height comes from an inline-block with the | |
| line-height carrying it. */ | |
| p a,figcaption a{display:inline-block;min-height:44px;line-height:44px; | |
| padding:0 2px} | |
| /* The <pre> pipeline diagram and the results table are the two things wider | |
| than a phone. Let each scroll inside its own box rather than pushing the | |
| document sideways — a horizontally scrolling PAGE hides content with no | |
| affordance, a scrolling code block is a known idiom. */ | |
| pre{background:var(--card);border:1px solid var(--line);border-radius:10px; | |
| padding:14px 16px;overflow-x:auto;font-size:var(--s1);max-width:100%} | |
| .tablewrap{overflow-x:auto;-webkit-overflow-scrolling:touch} | |
| .tablewrap table{min-width:640px} | |
| img{max-width:100%;height:auto} | |
| /* Subscripts default to a fraction of the parent and rendered at 11.7px — | |
| off the --s0..--s5 scale the audit counts. Pinned to the smallest step. */ | |
| sub,sup{font-size:var(--s0);line-height:0} | |
| </style></head><body><div class="wrap"><nav><a href="index.html">overview</a><a href="method.html">method</a><a href="results.html" aria-current="page">results</a><a href="sensors.html">sensors</a><a href="force.html">force</a><a href="gallery.html">gallery</a><a href="workbench.html">3D workbench</a><a href="probes/index.html" class="ext">probe clips</a><a href="testset/index.html" class="ext">test set</a><a href="sim/index.html" class="ext">simulator</a></nav> | |
| <h1>Results</h1> | |
| <p>Both reconstructions through <b>one</b> protocol: half the frames in each | |
| group fit a 5-feature least squares, half are scored; pooled ρ, five seeds, | |
| beside a within-group label shuffle. The second number is the comparable one — | |
| the share of the floor-to-perfect distance the arm covered, so a high floor | |
| cannot cap it. Only the image→gradient step differs.</p> | |
| <h3>Presses the sensor images whole</h3> | |
| <div class="tablewrap"><table><thead><tr><th>dataset</th><th>n</th><th>LUT<br>ρ / of what was left</th><th>calibration-free<br>ρ / of what was left</th><th>shuffle floor</th><th>groups fitted</th></tr></thead><tbody> | |
| <tr><td>GelSight Mini, CNC presses, 0-20 N</td><td>605<br><span class='dim'>of 6,219</span><td>0.993<br><span class='dim'>+0.992</span></td><td><b>0.998<br><span class='dim'>+0.998</span></b></td><td class='dim'>+0.067 / +0.089</td><td class='dim'>6/6</td></tr><tr><td>FoTa cnc_Mini</td><td>604<br><span class='dim'>of 3,351</span><td>0.900<br><span class='dim'>+0.900</span></td><td><b>0.938<br><span class='dim'>+0.936</span></b></td><td class='dim'>-0.040 / +0.035</td><td class='dim'>6/6</td></tr><tr><td>FEATS (MARKER gel)</td><td>2,000<br><span class='dim'>of 16,276</span><td>0.642<br><span class='dim'>+0.607</span></td><td><b>0.706<br><span class='dim'>+0.677</span></b></td><td class='dim'>+0.089 / +0.089</td><td class='dim'>8/8</td></tr><tr><td>Sparsh / Meta, 10 gel pads</td><td>2,000<br><span class='dim'>of 129,389</span><td>0.959<br><span class='dim'>+0.957</span></td><td><b>0.986<br><span class='dim'>+0.986</span></b></td><td class='dim'>+0.050 / +0.030</td><td class='dim'>4/4</td></tr><tr><td>FeelAnyForce<br><span class='bad'>floor-dominated</span></td><td>2,000<br><span class='dim'>of 13,892</span><td>0.939<br><span class='dim'>+0.764</span></td><td><b>0.949<br><span class='dim'>+0.803</span></b></td><td class='dim'>+0.742 / +0.741</td><td class='dim'>13/14</td></tr></tbody></table></div> | |
| <p><b>The one marker gel is the one low row.</b> FEATS is the only | |
| gel here with a printed dot lattice — 63 dots counted on | |
| its references against 1 | |
| or fewer on the other 4 — and it holds the lowest calibration-free ρ. | |
| Its panel shows why: the dots emboss themselves into the reconstructed surface. | |
| One markered dataset is one data point, so this is an observation, not a | |
| controlled comparison.</p> | |
| <figure><img src="assets/truncation.png" alt="truncated presses"> | |
| <figcaption>A press is <b>truncated</b> when its contact core reaches a | |
| border: the free-boundary solve then runs off the edge with nothing to stop the | |
| ramp, and its depth is not identifiable from the image. | |
| 15.9% reconstruct deeper than the | |
| 4.25 mm gel, against | |
| 0.0% of whole presses (441 and | |
| 59 frames).</figcaption></figure> | |
| <p>Excluding them is what the headline buys: calibration-free scores | |
| ρ 0.998 on whole presses | |
| against 0.851 with truncated frames | |
| mixed in. Two datasets cannot reach the 2,000 this table samples; they have no | |
| more presses to give.</p> | |
| <h3>All frames</h3> | |
| <div class="tablewrap"><table><thead><tr><th>dataset</th><th>n</th><th>LUT<br>ρ / of what was left</th><th>calibration-free<br>ρ / of what was left</th><th>shuffle floor</th><th>groups fitted</th></tr></thead><tbody> | |
| <tr><td>GelSight Mini, CNC presses, 0-20 N</td><td>2,376<br><span class='dim'>of 6,219</span><td>0.707<br><span class='dim'>+0.661</span></td><td><b>0.851<br><span class='dim'>+0.827</span></b></td><td class='dim'>+0.135 / +0.138</td><td class='dim'>6/6</td></tr><tr><td>FoTa cnc_Mini</td><td>2,239<br><span class='dim'>of 3,351</span><td>0.442<br><span class='dim'>+0.427</span></td><td><b>0.482<br><span class='dim'>+0.464</span></b></td><td class='dim'>+0.026 / +0.033</td><td class='dim'>6/6</td></tr><tr><td>FEATS (MARKER gel)</td><td>2,956<br><span class='dim'>of 16,276</span><td><b>0.707<br><span class='dim'>+0.642</span></b></td><td>0.702<br><span class='dim'>+0.629</span></td><td class='dim'>+0.183 / +0.196</td><td class='dim'>8/8</td></tr><tr><td>Sparsh / Meta, 10 gel pads</td><td>3,091<br><span class='dim'>of 129,389</span><td>0.669<br><span class='dim'>+0.656</span></td><td><b>0.698<br><span class='dim'>+0.685</span></b></td><td class='dim'>+0.040 / +0.041</td><td class='dim'>4/4</td></tr><tr><td>FeelAnyForce</td><td>3,378<br><span class='dim'>of 13,892</span><td>0.905<br><span class='dim'>+0.843</span></td><td><b>0.935<br><span class='dim'>+0.891</span></b></td><td class='dim'>+0.396 / +0.404</td><td class='dim'>14/14</td></tr></tbody></table></div> | |
| <p class="dim">React's production number adds a fitted position gain field and | |
| lives on the <a href="method.html">method</a> page.</p> | |
| <figure><img src="assets/pred_vs_gt.png" alt="predicted vs ground-truth force"> | |
| <figcaption>Held-out prediction against ground truth, shared axes per row, | |
| each panel annotated as the table is.</figcaption> | |
| </figure> | |
| <figure><img src="assets/cross_dataset.png" alt="cross-dataset transfer"> | |
| <figcaption>Fit on one dataset, predict on every other. Read each cell against | |
| the random-weight baseline under its column: the features are collinear and | |
| monotone in contact size, so on an easy target almost any direction | |
| ranks.</figcaption></figure> | |
| <p>FoTa cnc_Mini→FEATS, Sparsh→FEATS are ≥99 % extrapolation. Their MAE measures extrapolation, not prediction. | |
| FeelAnyForce's row goes <i>negative</i>: collinear features let least squares | |
| cancel opposite-sign terms (<a href="method.html">method</a>). Non-negative | |
| weights fix it: off-diagonal ρ | |
| 0.574 → 0.731, negative | |
| cells 3 → 0 of | |
| 20, costing 0.010 on the diagonal. | |
| <b>The deployed estimator is unchanged</b>: on React both agree at | |
| ρ 0.989 (1.8 % of | |
| frames outside the rig's range), and 15 held-out seeds differ by | |
| +0.002 ± 0.014 ρ.</p> | |
| <h2>Which reconstruction for React's force channel?</h2> | |
| <p>React's own calibration objects <b>cannot answer this</b>: calibration-free | |
| scores ρ 0.781 against the LUT's | |
| 0.763 on 158 held-out presses, but | |
| a paired bootstrap puts the margin at 95% CI | |
| [-0.081, | |
| +0.120] — a coin flip. The table is | |
| ahead on all five once each row's floor is corrected for, but that is a ranking | |
| over five other sensors, not a test on React. It ships because it needs no | |
| per-sensor lookup table.</p> | |
| <p>The two agree at ρ = 0.925 over 2,400 React | |
| frames, mean difference 0.86 N. Published to <code>yxma/React</code>: this channel across all 72 sides of 36 episodes (480,080 frames).</p> | |
| <h2>Error analysis</h2> | |
| <p>The ten worst held-out frames reconstruct as well as the five best — same | |
| dipoles, same compact depth, no ramping. The residual is in the depth→force | |
| fit, not in image→depth, so a better reconstruction will not move | |
| them.</p> | |
| <p class='dim'>Relative error is |pred−true| over the dataset's force span.</p><figure><img src="assets/errors_cnc_mini_26.png" alt="cnc_mini_26 errors"><figcaption>GelSight Mini CNC, span 19.50 N — median 0.7%, p90 2.7%, worst 14.1%.</figcaption></figure><figure><img src="assets/errors_cnc.png" alt="cnc errors"><figcaption>FoTa cnc_Mini, span 4.01 N — median 5.0%, p90 13.8%, worst 49.5%.</figcaption></figure><figure><img src="assets/errors_feats.png" alt="feats errors"><figcaption>FEATS (marker), span 32.74 N — median 5.4%, p90 15.0%, worst 79.7%.</figcaption></figure><figure><img src="assets/errors_sparsh.png" alt="sparsh errors"><figcaption>Sparsh / Meta, span 1.03 N — median 1.9%, p90 7.2%, worst 37.0%.</figcaption></figure><figure><img src="assets/errors_faf.png" alt="faf errors"><figcaption>FeelAnyForce, span 17.44 N — median 1.8%, p90 9.0%, worst 37.3%.</figcaption></figure> | |
| </div></body></html> |