yxma commited on
Commit
a33871a
·
verified ·
1 Parent(s): ee95e84

force recovery: methods, evaluation, debug log

Browse files
.gitattributes CHANGED
@@ -118,3 +118,5 @@ assets/recon/glowtact_16.png filter=lfs diff=lfs merge=lfs -text
118
  assets/recon/glowtact_17.png filter=lfs diff=lfs merge=lfs -text
119
  assets/recon/glowtact_18.png filter=lfs diff=lfs merge=lfs -text
120
  assets/recon/glowtact_19.png filter=lfs diff=lfs merge=lfs -text
 
 
 
118
  assets/recon/glowtact_17.png filter=lfs diff=lfs merge=lfs -text
119
  assets/recon/glowtact_18.png filter=lfs diff=lfs merge=lfs -text
120
  assets/recon/glowtact_19.png filter=lfs diff=lfs merge=lfs -text
121
+ assets/feats_marker_removal.png filter=lfs diff=lfs merge=lfs -text
122
+ assets/mnist_examples.png filter=lfs diff=lfs merge=lfs -text
assets/feats_marker_removal.png ADDED

Git LFS Details

  • SHA256: 5d700dd274a336ef3a96db6bebda23f42d3636fe6b906a881bb889ab63c81e6a
  • Pointer size: 132 Bytes
  • Size of remote file: 1.14 MB
assets/gallery/clip_motherboard_episode_003_left.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1969dde8cddca1b6dee4d022fe7158048f8448d2b376cd0229a57d0275f40998
3
- size 167859
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2b63045cce3dfc98bf332526f7b0325395315935ba13ac6167c2403c7ed8416a
3
+ size 165343
assets/gallery/clip_motherboard_episode_008_left.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a54d767bc4ca423cfcc861974cb6c9de8dca53c07e19540e6baa3583f952dc5e
3
- size 169872
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c3e6f513bd3edd081ea1a41906c15a886ff76bfb5fe07b67ce6beaec3965d5fd
3
+ size 166428
assets/gallery/clip_motherboard_episode_011_left.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d9b222a941bc914f0ab034b32c50bc90b21b0b82c10fd51e08031d19ea230800
3
- size 156408
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:111be423f21bef84cfe63e6fff218dc35e15ae3330d128240084e93a7b971ee6
3
+ size 156306
assets/gallery/clip_motherboard_episode_017_left.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5acf2de8e26e6efed7ef6a35e931458674e3fb5af88f1bf3ef4fe3495646e0be
3
- size 218522
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d5d44f81ca19728e9480b8a67f2d3cffb1625e48be5833d83d30e127f28a4807
3
+ size 214487
assets/gallery/clip_pushT_episode_001_right.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:217af6d88a5bf8088573c0c9e7938b243a7f6c1e1e543d924a2139cdd1531c1a
3
- size 148854
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6009cd9b75d7054d76010a1030fd24ec0b130c3104865e2ccacbca038777e66a
3
+ size 149351
assets/gallery/feats_00.png CHANGED

Git LFS Details

  • SHA256: 22a2dea734cca49b974e87e66a0baf374e5708fd8f9fe1aa62942844fa909944
  • Pointer size: 131 Bytes
  • Size of remote file: 350 kB

Git LFS Details

  • SHA256: e91b5862b8caa17bb55fb47a7bb5f18f23a4875c5e2785724b3db9eb77247664
  • Pointer size: 131 Bytes
  • Size of remote file: 344 kB
assets/gallery/feats_01.png CHANGED

Git LFS Details

  • SHA256: 1f218eab046e2bb66f44d742c56ddce4bfb2e4d7e0fc885d398e5245c55ce2dc
  • Pointer size: 131 Bytes
  • Size of remote file: 403 kB

Git LFS Details

  • SHA256: bba37732572c5d00f0d89528a88b2666db590d70fe8e340175dc085cd94187d4
  • Pointer size: 131 Bytes
  • Size of remote file: 403 kB
assets/gallery/feats_02.png CHANGED

Git LFS Details

  • SHA256: 98da12982bfc1b15756bab87b8fd601e9f61a6f886cb03ee32e3ca7ffdc4f39d
  • Pointer size: 131 Bytes
  • Size of remote file: 389 kB

Git LFS Details

  • SHA256: 04e7dd4a5603b86cda3f74e6f5fb50433209b120fc34ed011779446c7c20bee9
  • Pointer size: 131 Bytes
  • Size of remote file: 370 kB
assets/gallery/feats_03.png CHANGED

Git LFS Details

  • SHA256: 7fd856e2a0c3d77c94d3194af36252f1adf77fe51cee7a9fc1f72fe4a73dfde1
  • Pointer size: 131 Bytes
  • Size of remote file: 362 kB

Git LFS Details

  • SHA256: cf64741efe2485129723224a23ddebacb2f373d95b2f0d9f219703df0b482d3c
  • Pointer size: 131 Bytes
  • Size of remote file: 338 kB
assets/gallery/feats_04.png CHANGED

Git LFS Details

  • SHA256: 36cf62e522397eee8f91e46ba401725b24df8efd881fc431ce23a7b9533cfc37
  • Pointer size: 131 Bytes
  • Size of remote file: 398 kB

Git LFS Details

  • SHA256: d77db98d50f86549be93e9b2cb598e55783dfc2b26f297bd3c455ba2481166b0
  • Pointer size: 131 Bytes
  • Size of remote file: 379 kB
assets/gallery/feats_05.png CHANGED

Git LFS Details

  • SHA256: 7b05ff9e855e0c3d5c291cddc74d2d02796e3957d646e5ccbac6c369b7061f48
  • Pointer size: 131 Bytes
  • Size of remote file: 403 kB

Git LFS Details

  • SHA256: 50029937620e440d8aeff932167d79365e4d767bd0a407a5007d9ffe9e1e073b
  • Pointer size: 131 Bytes
  • Size of remote file: 397 kB
assets/gallery/metrics.json CHANGED
@@ -9,5 +9,6 @@
9
  "rho": 0.7571224581862139,
10
  "mae_n": 5.371393132649007
11
  },
 
12
  "pipeline": "lut_v2"
13
  }
 
9
  "rho": 0.7571224581862139,
10
  "mae_n": 5.371393132649007
11
  },
12
+ "feats_geometry": "marker-inpainted depth/mesh; force from stages()",
13
  "pipeline": "lut_v2"
14
  }
assets/mnist_examples.png ADDED

Git LFS Details

  • SHA256: 91da7009b1a1ef031b4c575c257776ada1cfc6169057ea1bfbb27798272b5f07
  • Pointer size: 132 Bytes
  • Size of remote file: 1.35 MB
assets/results_cnc.png CHANGED

Git LFS Details

  • SHA256: c72d1709d31c27f4c1d87145216e45a41c0b7717401dc0b22f94297b1aac6ade
  • Pointer size: 131 Bytes
  • Size of remote file: 171 kB

Git LFS Details

  • SHA256: bb0e7ec8c02f9cb4f0f19ed86ba9d95c25a116e1f79840b5c0e98dfce7d09e4c
  • Pointer size: 131 Bytes
  • Size of remote file: 172 kB
assets/results_feats.png CHANGED

Git LFS Details

  • SHA256: ee41be55a9f2a2a3b899d9406162a79e8b09a78fa810cff87d3130243f7c1998
  • Pointer size: 131 Bytes
  • Size of remote file: 134 kB

Git LFS Details

  • SHA256: 8944ebfc6031a3b2f95a9f61e23b948621c8502dd43d8b7bccc0f4a5f58be90f
  • Pointer size: 131 Bytes
  • Size of remote file: 135 kB
assets/results_glowtact.png CHANGED

Git LFS Details

  • SHA256: 6529477dfc630958faeec8a7ba590a623ca8ce5952e5711f973536ce29d84d7f
  • Pointer size: 131 Bytes
  • Size of remote file: 170 kB

Git LFS Details

  • SHA256: 64a8c72d1db99cd6912b9e73f53b0555cedf09c183100902b7e3e5660111b5a2
  • Pointer size: 131 Bytes
  • Size of remote file: 169 kB
debug_pipeline.html CHANGED
@@ -31,7 +31,8 @@ th,td{padding:7px 10px;border-bottom:1px solid var(--line);text-align:right}
31
  th{color:var(--dim);font-family:'IBM Plex Mono',monospace;font-size:.72rem;
32
  text-transform:uppercase;letter-spacing:.08em}
33
  td:first-child,th:first-child{text-align:left}
34
- img{max-width:100%;border-radius:8px;border:1px solid var(--line);display:block;margin:12px auto}
 
35
  a{color:#4fd8e0}a:hover{color:var(--force)}
36
  .pill{display:inline-block;border:1px solid var(--line);border-radius:999px;
37
  padding:3px 12px;font-size:.75rem;font-family:'IBM Plex Mono',monospace;
@@ -77,6 +78,22 @@ relied on.</p>
77
  z-supervised gain field. Dot-type control: FEATS presses are always centered,
78
  so the raw pipeline already gives &rho;=0.72 there, vs 0.46 raw on GlowTact —
79
  the gain field and scope, not markers, are the binding factors.</p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
80
  <h2>GlowTact — markerless, the LUT's home sensor (control)</h2><div class='srow'>
81
  <img src='assets/debug/glowtact_01.png' loading='lazy'>
82
  <img src='assets/debug/glowtact_04.png' loading='lazy'>
@@ -93,7 +110,7 @@ the gain field and scope, not markers, are the binding factors.</p>
93
  <img src='assets/debug/cnc_14.png' loading='lazy'>
94
  <img src='assets/debug/cnc_17.png' loading='lazy'>
95
  </div>
96
- <h2>FEATS — dot/marker type; difference imaging cancels static markers</h2><div class='srow'>
97
  <img src='assets/debug/feats_01.png' loading='lazy'>
98
  <img src='assets/debug/feats_04.png' loading='lazy'>
99
  <img src='assets/debug/feats_08.png' loading='lazy'>
 
31
  th{color:var(--dim);font-family:'IBM Plex Mono',monospace;font-size:.72rem;
32
  text-transform:uppercase;letter-spacing:.08em}
33
  td:first-child,th:first-child{text-align:left}
34
+ img,video{max-width:100%;height:auto;border-radius:8px;border:1px solid var(--line);
35
+ display:block;margin:12px auto}
36
  a{color:#4fd8e0}a:hover{color:var(--force)}
37
  .pill{display:inline-block;border:1px solid var(--line);border-radius:999px;
38
  padding:3px 12px;font-size:.75rem;font-family:'IBM Plex Mono',monospace;
 
78
  z-supervised gain field. Dot-type control: FEATS presses are always centered,
79
  so the raw pipeline already gives &rho;=0.72 there, vs 0.46 raw on GlowTact —
80
  the gain field and scope, not markers, are the binding factors.</p>
81
+ <h2>The same failure mode, on four datasets</h2>
82
+ <p>The cnc story above is not a cnc story. Exact per-pixel ground truth on
83
+ Tactile MNIST shows the identical effect from the other side: 420/420 touches
84
+ have a contact that runs off the pad, and a control that moves one sphere cap
85
+ from mid-pad to the edge collapses its reconstructed peak <b>1.39 &rarr; 0.30
86
+ mm</b> against a 0.90 mm truth. Sparsh shows it a third time — its clipped
87
+ frames carry the <i>highest</i> median force and still score worse. The cause
88
+ is structural: the Poisson solve's zero boundary cannot represent a surface
89
+ that leaves the frame, so a clipped contact is integrated as if it ended at the
90
+ image edge. <b>Contact visibility, not force range or gel type, is this
91
+ pipeline's biggest external failure mode.</b></p>
92
+ <p>The same ground truth also puts a number on the mask you see in column 4:
93
+ against the true contact region it scores IoU <b>0.614</b> with recall
94
+ <b>0.917</b> and over-segmentation <b>0.531</b> — it finds nearly all of the
95
+ contact, then adds half as much again in halo. That is why a halo pedestal has
96
+ to be subtracted before any 3D view.</p>
97
  <h2>GlowTact — markerless, the LUT's home sensor (control)</h2><div class='srow'>
98
  <img src='assets/debug/glowtact_01.png' loading='lazy'>
99
  <img src='assets/debug/glowtact_04.png' loading='lazy'>
 
110
  <img src='assets/debug/cnc_14.png' loading='lazy'>
111
  <img src='assets/debug/cnc_17.png' loading='lazy'>
112
  </div>
113
+ <h2>FEATS — dot/marker gel; these panels are the FORCE path, so the dots are still in them (marker inpainting ships for geometry only)</h2><div class='srow'>
114
  <img src='assets/debug/feats_01.png' loading='lazy'>
115
  <img src='assets/debug/feats_04.png' loading='lazy'>
116
  <img src='assets/debug/feats_08.png' loading='lazy'>
debug_pipeline_zh.html CHANGED
@@ -31,7 +31,8 @@ th,td{padding:7px 10px;border-bottom:1px solid var(--line);text-align:right}
31
  th{color:var(--dim);font-family:'IBM Plex Mono',monospace;font-size:.72rem;
32
  text-transform:uppercase;letter-spacing:.08em}
33
  td:first-child,th:first-child{text-align:left}
34
- img{max-width:100%;border-radius:8px;border:1px solid var(--line);display:block;margin:12px auto}
 
35
  a{color:#4fd8e0}a:hover{color:var(--force)}
36
  .pill{display:inline-block;border:1px solid var(--line);border-radius:999px;
37
  padding:3px 12px;font-size:.75rem;font-family:'IBM Plex Mono',monospace;
@@ -64,6 +65,9 @@ border-radius:6px;margin:4px 0}</style></head><body><div class="wrap">
64
  <td><b>0.94 0.94 0.93 0.91 0.95 0.95</b></td><td><b>0.941</b></td>
65
  <td><b>0.26 N</b></td></tr></table>
66
  <p class="note">每探头半半分拟合,7 种子,LUT 管线 + z 监督增益场。dot 类型对照:FEATS 按压恒居中,原始管线即得 &rho;=0.72,而 GlowTact 原始条件仅 0.46——起决定作用的是增益场与作用域,而非 marker。</p>
 
 
 
67
  <h2>GlowTact — 无 marker,LUT 标定所用传感器(对照)</h2><div class='srow'>
68
  <img src='assets/debug/glowtact_01.png' loading='lazy'>
69
  <img src='assets/debug/glowtact_04.png' loading='lazy'>
@@ -80,7 +84,7 @@ border-radius:6px;margin:4px 0}</style></head><body><div class="wrap">
80
  <img src='assets/debug/cnc_14.png' loading='lazy'>
81
  <img src='assets/debug/cnc_17.png' loading='lazy'>
82
  </div>
83
- <h2>FEATS — dot/marker 类型差分成像抵消静态 marker</h2><div class='srow'>
84
  <img src='assets/debug/feats_01.png' loading='lazy'>
85
  <img src='assets/debug/feats_04.png' loading='lazy'>
86
  <img src='assets/debug/feats_08.png' loading='lazy'>
 
31
  th{color:var(--dim);font-family:'IBM Plex Mono',monospace;font-size:.72rem;
32
  text-transform:uppercase;letter-spacing:.08em}
33
  td:first-child,th:first-child{text-align:left}
34
+ img,video{max-width:100%;height:auto;border-radius:8px;border:1px solid var(--line);
35
+ display:block;margin:12px auto}
36
  a{color:#4fd8e0}a:hover{color:var(--force)}
37
  .pill{display:inline-block;border:1px solid var(--line);border-radius:999px;
38
  padding:3px 12px;font-size:.75rem;font-family:'IBM Plex Mono',monospace;
 
65
  <td><b>0.94 0.94 0.93 0.91 0.95 0.95</b></td><td><b>0.941</b></td>
66
  <td><b>0.26 N</b></td></tr></table>
67
  <p class="note">每探头半半分拟合,7 种子,LUT 管线 + z 监督增益场。dot 类型对照:FEATS 按压恒居中,原始管线即得 &rho;=0.72,而 GlowTact 原始条件仅 0.46——起决定作用的是增益场与作用域,而非 marker。</p>
68
+ <h2>同一个失效模式,出现在四个数据集上</h2>
69
+ <p>上面这个 cnc 的故事其实不只属于 cnc。Tactile MNIST 上精确的逐像素真值从另一侧展示了完全相同的效应:420/420 次触碰的接触都跑出了 pad,而把单个球冠从 pad 中心移到边缘的对照,使重建峰值从 <b>1.39 塌到 0.30 mm</b>(真值 0.90 mm)。Sparsh 是第三次印证——它被裁切的帧力中位数<i>最高</i>,得分却更差。原因是结构性的:Poisson 求解的零边界无法表示跑出画面的曲面,所以被裁切的接触会被当成在图像边缘就结束了。<b>接触可见性——而不是力程或 gel 类型——才是这条管线最大的外部失效模式。</b></p>
70
+ <p>同一份真值也给第 4 列那个掩码定了量:相对真实接触区域,它的 IoU 为 <b>0.614</b>,召回 <b>0.917</b>,过分割 <b>0.531</b>——它几乎找全了接触,然后又多加了半个接触面积的光晕。这正是任何 3D 视图之前都必须先减掉光晕基座的原因。</p>
71
  <h2>GlowTact — 无 marker,LUT 标定所用传感器(对照)</h2><div class='srow'>
72
  <img src='assets/debug/glowtact_01.png' loading='lazy'>
73
  <img src='assets/debug/glowtact_04.png' loading='lazy'>
 
84
  <img src='assets/debug/cnc_14.png' loading='lazy'>
85
  <img src='assets/debug/cnc_17.png' loading='lazy'>
86
  </div>
87
+ <h2>FEATS — dot/marker gel这里的面板是力的路径,所以点仍然在图里(marker 修补只用于几何)</h2><div class='srow'>
88
  <img src='assets/debug/feats_01.png' loading='lazy'>
89
  <img src='assets/debug/feats_04.png' loading='lazy'>
90
  <img src='assets/debug/feats_08.png' loading='lazy'>
index.html CHANGED
@@ -63,7 +63,8 @@ tactile-estimated normal force, and DexForce-style force-informed action targets
63
  <span class="pill">dataset: yxma/React</span>
64
  <a class="pill" href="method.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ method design</a>
65
  <a class="pill" href="results.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ results matrix</a>
66
- <a class="pill" href="actions.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ action processing deep-dive</a></div>
 
67
  </div></header>
68
  <div class="wrap">
69
 
@@ -98,22 +99,61 @@ the optimization journey, and the step-by-step reconstruction debug are on the
98
  <img src="assets/depth_validation_panel.png" alt="raw | diff | depth">
99
  <p class="footnote">Strongest motherboard presses: raw | difference |
100
  LUT-reconstructed depth. More examples (20 panels, 10 clips) in the
101
- <a href="gallery.html">gallery</a>.</p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
102
  </div>
103
 
104
  <h2>Validated against four force-labeled datasets</h2>
105
  <div class="card method">
106
  <table>
107
- <tr><th>dataset (gel)</th><th>ours</th><th>FEATS U-net</th><th>FeelAnyForce</th></tr>
108
- <tr><td>FEATS val (marker)</td><td>0.77</td><td><b>0.96</b> in-domain</td><td>0.43</td></tr>
109
- <tr><td>FoTa cnc_Mini (markerless)</td><td>0.94 (in view)</td><td>0.07</td><td><b>0.83</b></td></tr>
110
- <tr><td>GlowTact (markerless)</td><td><b>0.98</b></td><td>0.04</td><td>0.90</td></tr>
111
- <tr><td>Sparsh / Meta (markerless, 10 pads)</td><td><b>0.97</b> *</td><td colspan="2">not run — no published predictions</td></tr>
112
- <tr><td>React (no GT — agreement)</td><td colspan="3">physics vs FeelAnyForce ρ = 0.91; both read ≈0 N off-contact</td></tr>
 
 
 
 
 
 
 
113
  </table>
114
  <p>Each network dominates its own gel domain and collapses
115
- outside it; the physics pipeline (0.74–0.99) is the only one that works
116
- everywhere. Predicted-vs-ground-truth scatters, per dataset, on the
 
 
 
117
  <a href="results.html"><b>results page</b></a>.</p>
118
  <p class="footnote" style="margin-bottom:0">* Sparsh is a <b>foreign sensor</b>.
119
  Our GlowTact table reaches 0.878 there; rebuilding the table from Sparsh's own
 
63
  <span class="pill">dataset: yxma/React</span>
64
  <a class="pill" href="method.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ method design</a>
65
  <a class="pill" href="results.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ results matrix</a>
66
+ <a class="pill" href="actions.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ action processing deep-dive</a>
67
+ <a class="pill" href="recon_workbench.html" style="text-decoration:none;color:#ffd9a0;border:1px solid rgba(255,255,255,.3)">→ 3D reconstruction workbench</a></div>
68
  </div></header>
69
  <div class="wrap">
70
 
 
99
  <img src="assets/depth_validation_panel.png" alt="raw | diff | depth">
100
  <p class="footnote">Strongest motherboard presses: raw | difference |
101
  LUT-reconstructed depth. More examples (20 panels, 10 clips) in the
102
+ <a href="gallery.html">gallery</a>; every stage of the reconstruction, with
103
+ its knobs, in the <a href="recon_workbench.html">3D workbench</a>. On marker
104
+ gels the depth and 3D products get one extra step — the dots are inpainted out
105
+ of reference and frame before differencing, which removes the dimple lattice
106
+ from the geometry but is deliberately <b>not</b> applied to the force
107
+ features, where it costs 0.04 ρ.</p>
108
+ </div>
109
+
110
+ <h2>How accurate is the geometry itself?</h2>
111
+ <div class="card method">
112
+ <p style="margin-top:0">Measured against exact per-pixel ground truth for the
113
+ first time — ray-cast mesh depth on 420 Tactile MNIST touches, non-spherical
114
+ geometry a sphere calibration cannot self-validate. The answer is a range set
115
+ by press depth, not a single number:</p>
116
+ <table>
117
+ <tr><th>press depth</th><th>0.30 mm</th><th>0.60 mm</th><th>1.00 mm</th>
118
+ <th>1.50 mm</th><th>2.25 mm</th></tr>
119
+ <tr><td>depth MAE</td><td><b>11 µm</b></td><td>35 µm</td><td>68 µm</td>
120
+ <td>127 µm</td><td>281 µm</td></tr>
121
+ <tr><td>peak recovered</td><td>1.00</td><td>0.97</td><td>0.77</td><td>0.68</td>
122
+ <td>0.55</td></tr>
123
+ </table>
124
+ <p>At 0.3 mm, with no per-frame alignment and no fitted indentation scale, we
125
+ are below every published 3D Cal figure; by 2.25 mm we recover barely half the
126
+ peak. <b>The working range is shallow contact, and no accuracy number should
127
+ be quoted without its press depth.</b> The same run cut our own over-doming
128
+ claim to +7–12%, demoted unobserved LUT bins to a minor factor, and confirmed
129
+ that the valid mask is halo-dominated (IoU 0.614, over-segmentation 0.531) —
130
+ details and the Taxim caveat on the
131
+ <a href="results.html">results page</a>.</p>
132
  </div>
133
 
134
  <h2>Validated against four force-labeled datasets</h2>
135
  <div class="card method">
136
  <table>
137
+ <tr><th>dataset (gel)</th><th>ours</th><th>shuffle control</th>
138
+ <th>FEATS U-net</th><th>FeelAnyForce</th></tr>
139
+ <tr><td>FEATS val (marker)</td><td>0.77</td>
140
+ <td>-0.00</td><td><b>0.96</b> in-domain</td><td>0.43</td></tr>
141
+ <tr><td>FoTa cnc_Mini (markerless)</td><td>0.95 (in view)</td>
142
+ <td>+0.06</td><td>0.07</td><td><b>0.83</b></td></tr>
143
+ <tr><td>GlowTact (markerless)</td><td><b>0.99</b></td>
144
+ <td>+0.17</td><td>0.04</td><td>0.90</td></tr>
145
+ <tr><td>Sparsh / Meta (markerless, 10 pads)</td>
146
+ <td><b>0.97</b> *</td>
147
+ <td>+0.26</td>
148
+ <td colspan="2">not run — no published predictions</td></tr>
149
+ <tr><td>React (no GT — agreement)</td><td colspan="4">physics vs FeelAnyForce ρ = 0.91; both read ≈0 N off-contact</td></tr>
150
  </table>
151
  <p>Each network dominates its own gel domain and collapses
152
+ outside it; the physics pipeline (0.77–0.99)
153
+ is the only one that works everywhere. Every row is a per-group half/half fit
154
+ with isotonic calibration, 5 seeds, reported beside the same protocol with the
155
+ force labels shuffled <i>within</i> each group. Predicted-vs-ground-truth
156
+ scatters, per dataset, on the
157
  <a href="results.html"><b>results page</b></a>.</p>
158
  <p class="footnote" style="margin-bottom:0">* Sparsh is a <b>foreign sensor</b>.
159
  Our GlowTact table reaches 0.878 there; rebuilding the table from Sparsh's own
method.html CHANGED
@@ -59,6 +59,7 @@ th,td{white-space:normal;min-width:110px}}
59
  A physics pipeline with exactly one fitted number.</p>
60
  <a class="pill" href="results.html">↖ results matrix</a>
61
  <a class="pill" href="index.html">overview</a>
 
62
  <a class="pill" href="actions.html">action transform</a>
63
  <a class="pill" href="method_zh.html">中文</a>
64
  </header>
@@ -112,14 +113,89 @@ supervised network reaches 2.1 N on the same range). A new sensor needs one
112
 
113
  <h2>Validation (LUT-v2 pipeline)</h2>
114
  <div class="card">
115
- <p style="margin-top:0">The LUT-v2 pipeline is validated on <b>three force-labeled
116
- datasets</b> (FEATS, FoTa cnc_Mini, GlowTact) and cross-checked against two
117
- neural estimators on identical frames — every predicted-vs-ground-truth
118
  scatter, per dataset, lives on the
119
  <a href="results.html"><b>results page</b></a>. Short version: physics
120
- 0.74-0.99 everywhere (GlowTact 0.98, cnc_Mini in-view 0.94, FEATS 0.77);
121
- each network 0.90+ in its own gel domain and collapsing
122
- outside it.</p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
123
  </div>
124
 
125
  <h2>Per-dataset calibration</h2>
 
59
  A physics pipeline with exactly one fitted number.</p>
60
  <a class="pill" href="results.html">↖ results matrix</a>
61
  <a class="pill" href="index.html">overview</a>
62
+ <a class="pill" href="recon_workbench.html">3D workbench</a>
63
  <a class="pill" href="actions.html">action transform</a>
64
  <a class="pill" href="method_zh.html">中文</a>
65
  </header>
 
113
 
114
  <h2>Validation (LUT-v2 pipeline)</h2>
115
  <div class="card">
116
+ <p style="margin-top:0">The LUT-v2 pipeline is validated on <b>four force-labeled
117
+ datasets</b> (FEATS, FoTa cnc_Mini, GlowTact, Sparsh) and cross-checked against
118
+ two neural estimators on identical frames — every predicted-vs-ground-truth
119
  scatter, per dataset, lives on the
120
  <a href="results.html"><b>results page</b></a>. Short version: physics
121
+ 0.77-0.99 everywhere (GlowTact 0.99, cnc_Mini in-view 0.95, Sparsh in-view
122
+ 0.97, FEATS 0.77), each reported beside a within-group label-shuffle control
123
+ that lands between &minus;0.00 and 0.26; each network 0.90+ in its own gel
124
+ domain and collapsing outside it.</p>
125
+ </div>
126
+
127
+ <h2>How accurate is the depth, in micrometres?</h2>
128
+ <div class="card">
129
+ <p style="margin-top:0">Every &rho; above scores <b>force</b>. The depth
130
+ underneath used to be checked only against our own analytic sphere cap. It is
131
+ now measured against exact per-pixel ground truth — ray-cast mesh depth on 420
132
+ Tactile MNIST touches, non-spherical geometry a sphere calibration cannot
133
+ self-validate — and the answer is a <b>range set by press depth</b>, not a
134
+ number:</p>
135
+ <table>
136
+ <tr><th>press depth</th><th>0.30 mm</th><th>0.60 mm</th><th>1.00 mm</th>
137
+ <th>1.50 mm</th><th>2.25 mm</th></tr>
138
+ <tr><td>MAE</td><td class="ok"><b>11 µm</b></td><td class="ok">35 µm</td>
139
+ <td>68 µm</td><td>127 µm</td><td class="warn">281 µm</td></tr>
140
+ <tr><td>peak recovered</td><td>1.00</td><td>0.97</td><td>0.77</td><td>0.68</td>
141
+ <td class="warn">0.55</td></tr>
142
+ </table>
143
+ <p>At 0.3 mm, with no per-frame alignment and no fitted indentation scale, the
144
+ Type-2 error is 96.5 µm — below all three of 3D Cal's published figures
145
+ (152.8&thinsp;/&thinsp;171.6&thinsp;/&thinsp;290.0 µm), which are reported
146
+ <i>with</i> both. By 2.25 mm we recover barely half the peak. <b>The working
147
+ range of this reconstruction is shallow contact, and no accuracy number here
148
+ should be quoted without its press depth.</b></p>
149
+ <p>The same run corrected two of our own claims and confirmed a third:
150
+ flat-top over-doming is <b>+7% to +12%</b>, not the +23&ndash;42% we implied
151
+ (compliant gel wraps a flat edge, so the true centre/rim ratio is
152
+ 1.33&ndash;1.40, not 1.0); unobserved LUT bins are a <b>minor</b> factor
153
+ (correlation 0.10&ndash;0.29 with error); and the <code>|dI|&gt;8</code> valid
154
+ mask really is halo-dominated (IoU 0.614, recall 0.917, over-segmentation
155
+ 0.531). The photometric table is <i>not</i> the weak link — its gradients sit
156
+ 24.4&deg; from the true ones off-domain versus 26.1&deg; on its own sensor.
157
+ Full tables, controls and the Taxim caveat on the
158
+ <a href="results.html"><b>results page</b></a>.</p>
159
+ </div>
160
+
161
+ <h2>Marker gels: one extra step, and what it is not for</h2>
162
+ <div class="card">
163
+ <p style="margin-top:0">FEATS is our only dotted gel. The dots occlude the gel,
164
+ so the photometric table has no valid colour underneath them and Poisson
165
+ integration bakes a dimple lattice into the depth map — visible as pockmarks
166
+ across the whole reconstructed surface. The fix is the one
167
+ <a href="https://arxiv.org/abs/2106.08851">GelSight Wedge</a>
168
+ (Wang/She/Dong/Adelson, ICRA 2021) gives in its Fig. 10 for marker holes —
169
+ detect the dots, then <i>fill</i> them rather than mask them — moved from
170
+ gradient space into image space: inpaint the dots out of <b>both</b> the
171
+ reference and the frame (cv2 Telea) and difference the inpainted pair.</p>
172
+ <table>
173
+ <tr><th>variant</th><th>dimple power at the 31.9 px pitch</th><th>force &rho;</th></tr>
174
+ <tr><td>baseline, no marker handling</td><td>1.523</td><td class="ok">0.7747</td></tr>
175
+ <tr><td><b>inpaint the image, ref + frame</b> — shipped, depth only</td>
176
+ <td class="ok"><b>0.890</b> (&times;0.65)</td><td>0.7371</td></tr>
177
+ <tr><td>inpaint the gradient field instead</td><td>0.924</td><td>0.7408</td></tr>
178
+ <tr><td>zero the gradient inside the holes</td><td class="warn">1.251</td><td>0.7620</td></tr>
179
+ </table>
180
+ <p>Two things worth reading off that table. <b>Filling beats masking</b>:
181
+ setting g&nbsp;:=&nbsp;0 on the holes barely helps (&times;0.93) because a zero
182
+ patch puts a dipole layer on every hole boundary, at exactly the lattice
183
+ frequency it is supposed to remove. And <b>the geometry column and the force
184
+ column disagree</b> — the variant that halves the dimple power is the one that
185
+ loses the most &rho;. So the step ships for <b>depth and 3D only</b>
186
+ (<code>marker_removal.stages_depth</code>, which wraps rather than edits the
187
+ force path); the force features on FEATS still come from the untouched
188
+ pipeline. Everything else on this site is markerless, where the detector finds
189
+ zero dots and the step is a bit-exact no-op.</p>
190
+ <p class="footnote">Controls, because &ldquo;inpainting helps&rdquo; is an easy
191
+ thing to fool yourself about: inpainting the same <i>area</i> of randomly
192
+ placed fake markers gives &rho; 0.7697 — no gain, so this is not smoothing.
193
+ The detector finds 63/63 dots with 0 rejects on the FEATS reference and stays
194
+ at 63 for every threshold from 3 to 16 grey levels, but exactly 0 blobs on the
195
+ GlowTact and cnc references, so it is marker-specific. It is also not a
196
+ complete fix: the dots shear with the gel (median 1.7 px, &gt;8 px on 8% of
197
+ frames) and a static reference mask leaves the displaced ones in place. Full
198
+ study: <code>python -m force_recovery.marker_study all</code>.</p>
199
  </div>
200
 
201
  <h2>Per-dataset calibration</h2>
method_zh.html CHANGED
@@ -59,6 +59,7 @@ th,td{white-space:normal;min-width:110px}}
59
  一条只有一个拟合参数的物理管线。</p>
60
  <a class="pill" href="results_zh.html">↖ 评测结果</a>
61
  <a class="pill" href="index.html">总览(英文)</a>
 
62
  <a class="pill" href="actions_zh.html">动作变换</a>
63
  <a class="pill" href="method.html">English</a>
64
  </header>
@@ -111,7 +112,67 @@ supervised network reaches 2.1 N on the same range). A new sensor needs one
111
 
112
  <h2>验证(LUT-v2 管线)</h2>
113
  <div class="card">
114
- <p style="margin-top:0">LUT-v2 管线在<b>个带力标注的数据集</b>(FEATS、FoTa cnc_Mini、GlowTact)上验证,并与两个神经网络估计器同帧对比——每个数据集的预测 vs 真值散点图都在<a href="results_zh.html"><b>评测结果页</b></a>。一句话版:物理方法在所有域 0.74–0.99(GlowTact 0.98、cnc_Mini 视野内 0.94、FEATS 0.77);每个网络在自己的 gel 域 0.90+,出域即塌。</p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
115
  </div>
116
 
117
  <h2>逐数据集标定</h2>
 
59
  一条只有一个拟合参数的物理管线。</p>
60
  <a class="pill" href="results_zh.html">↖ 评测结果</a>
61
  <a class="pill" href="index.html">总览(英文)</a>
62
+ <a class="pill" href="recon_workbench.html">3D 工作台</a>
63
  <a class="pill" href="actions_zh.html">动作变换</a>
64
  <a class="pill" href="method.html">English</a>
65
  </header>
 
112
 
113
  <h2>验证(LUT-v2 管线)</h2>
114
  <div class="card">
115
+ <p style="margin-top:0">LUT-v2 管线在<b>个带力标注的数据集</b>(FEATS、FoTa cnc_Mini、GlowTact、Sparsh)上验证,并与两个神经网络估计器同帧对比——每个数据集的预测 vs 真值散点图都在<a href="results_zh.html"><b>评测结果页</b></a>。一句话版:物理方法在所有域 0.77–0.99(GlowTact 0.99、cnc_Mini 视野内 0.95Sparsh 视野内 0.97、FEATS 0.77),每个数字旁边都给出组内标签打乱对照(落在 &minus;0.00 到 0.26 之间);每个网络在自己的 gel 域 0.90+,出域即塌。</p>
116
+ </div>
117
+
118
+ <h2>深度到底准到多少微米?</h2>
119
+ <div class="card">
120
+ <p style="margin-top:0">上面所有 &rho; 评的都是<b>力</b>。底层的深度此前只和我们自己的
121
+ 解析球冠比较过。现在它已经对着精确的逐像素真值测量——420 次 Tactile MNIST 触碰的
122
+ 网格光线投射深度,而且是球面标定无法自证的非球面几何——答案是一个
123
+ <b>由压入深度决定的区间</b>,而不是一个数字:</p>
124
+ <table>
125
+ <tr><th>压入深度</th><th>0.30 mm</th><th>0.60 mm</th><th>1.00 mm</th>
126
+ <th>1.50 mm</th><th>2.25 mm</th></tr>
127
+ <tr><td>MAE</td><td class="ok"><b>11 µm</b></td><td class="ok">35 µm</td>
128
+ <td>68 µm</td><td>127 µm</td><td class="warn">281 µm</td></tr>
129
+ <tr><td>峰值还原比</td><td>1.00</td><td>0.97</td><td>0.77</td><td>0.68</td>
130
+ <td class="warn">0.55</td></tr>
131
+ </table>
132
+ <p>在 0.3 mm,且没有逐帧配准、也没有拟合压入尺度的情况下,Type-2 误差为 96.5 µm——
133
+ 低于 3D Cal 公布的全部三个数字(152.8&thinsp;/&thinsp;171.6&thinsp;/&thinsp;290.0 µm),
134
+ 而他们的结果是<i>带</i>这两项才得到的。到 2.25 mm 我们只还原了大约一半峰值。
135
+ <b>这套重建的适用区间是浅接触,本页任何精度数字都必须连同其压入深度一起引用。</b></p>
136
+ <p>同一次实验纠正了我们自己的两个说法,并确认了第三个:平头过度穹顶化是
137
+ <b>+7% 到 +12%</b>,而不是我们暗示的 +23&ndash;42%(gel 柔顺、会包裹平边,
138
+ 所以真实的中心/边缘比是 1.33&ndash;1.40,而不是 1.0);未观测的 LUT bin 是
139
+ <b>次要</b>因素(与误差的相关性只有 0.10&ndash;0.29);而 <code>|dI|&gt;8</code>
140
+ 有效掩码确实被光晕主导(IoU 0.614、召回 0.917、过分割 0.531)。
141
+ 光度查找表<i>不是</i>薄弱环节——它的梯度在域外与真值相差 24.4&deg;,
142
+ 而在它自己的传感器上是 26.1&deg;。完整表格、对照与 Taxim 相关的边界说明见
143
+ <a href="results_zh.html"><b>评测结果页</b></a>。</p>
144
+ </div>
145
+
146
+ <h2>有 marker 的 gel:多一步,以及它不是为了什么</h2>
147
+ <div class="card">
148
+ <p style="margin-top:0">FEATS 是我们唯一一块带点的 gel。点会遮挡 gel,
149
+ 所以光度查找表在点下没有有效颜色,Poisson 积分会把一层点阵凹坑烤进深度图——
150
+ 在整个重建表面上表现为麻点。修法就是
151
+ <a href="https://arxiv.org/abs/2106.08851">GelSight Wedge</a>
152
+ (Wang/She/Dong/Adelson,ICRA 2021)图 10 针对 marker 孔洞给出的做法——
153
+ 检测这些点,然后<i>填充</i>而不是屏蔽——只不过从梯度域搬到了图像域:
154
+ 把点从<b>参考帧和当前帧</b>里一起修补掉(cv2 Telea),再对修补后的两幅图做差分。</p>
155
+ <table>
156
+ <tr><th>方案</th><th>31.9 px 点距频率上的凹坑功率</th><th>力 &rho;</th></tr>
157
+ <tr><td>基线,不处理 marker</td><td>1.523</td><td class="ok">0.7747</td></tr>
158
+ <tr><td><b>图像域修补,参考帧 + 当前帧</b>——已上线,仅用于深度</td>
159
+ <td class="ok"><b>0.890</b>(&times;0.65)</td><td>0.7371</td></tr>
160
+ <tr><td>改为修补梯度场</td><td>0.924</td><td>0.7408</td></tr>
161
+ <tr><td>把孔洞内的梯度置零</td><td class="warn">1.251</td><td>0.7620</td></tr>
162
+ </table>
163
+ <p>这张表有两点值得读出来。<b>填充胜过屏蔽</b>:把孔洞里的 g 置零几乎没用(&times;0.93),
164
+ 因为零补丁会在每个孔洞边界上放一层偶极子,频率恰好就是它本该去掉的那个点阵频率。
165
+ ��及<b>几何列和力列是相反的</b>——把凹坑功率减半的那个方案,恰恰是 &rho; 掉得最多的。
166
+ 所以这一步只用于<b>深度与 3D</b>(<code>marker_removal.stages_depth</code>,
167
+ 它是包一层而不是改动力的路径);FEATS 的力特征仍然来自未改动的管线。
168
+ 本站其余数据都是无 marker 的,检测器在那里找不到点,这一步是逐比特的空操作。</p>
169
+ <p class="footnote">对照实验,因为&ldquo;修补有用&rdquo;是很容易自欺的结论:
170
+ 把同样<i>面积</i>的随机假 marker 修补掉得到 &rho; 0.7697——没有增益,所以这不是平滑效应。
171
+ 检测器在 FEATS 参考帧上得到 63/63 个点、0 个误检,阈值从 3 到 16 灰阶都保持 63;
172
+ 而在 GlowTact 与 cnc 参考帧上恰好是 0 个,所以它是 marker 专用的。
173
+ 它也不是完全的修复:点会随 gel 剪切移动(中位 1.7 px,8% 的帧超过 8 px),
174
+ 静态参考掩码会漏掉位移后的点。完整研究:
175
+ <code>python -m force_recovery.marker_study all</code>。</p>
176
  </div>
177
 
178
  <h2>逐数据集标定</h2>
recon_workbench.html CHANGED
@@ -9,7 +9,8 @@ img{width:100%;background:#fff;border-radius:5px;margin:5px 0}
9
  table{border-collapse:collapse;margin:10px 0;font-size:13px}
10
  th,td{border:1px solid #2c3648;padding:5px 10px;text-align:left}
11
  th{color:#ffd9a0;font-weight:600}code{background:#1b2334;padding:1px 5px;border-radius:3px}
12
- .note{color:#8e99ab;font-size:13px}</style></head><body><div class="wrap">
 
13
  <h1>3D reconstruction workbench — every stage, 20 samples</h1>
14
  <p>Not a results figure: this is the full chain with every tunable quantity
15
  exposed, so a defect can be attributed to a stage. Columns:
@@ -19,34 +20,75 @@ exposed, so a defect can be attributed to a stage. Columns:
19
  <b>11</b> depth · <b>12</b> depth halo-removed · <b>13</b> radial profile vs
20
  analytic cap · <b>14</b> Open3D mesh.
21
  Reproduce: <code>xvfb-run -a -s "-screen 0 1400x1000x24" python -m
22
- force_recovery.recon_study glowtact</code></p>
23
- <h2>Three defects the numbers localise</h2>
24
- <p><b>1. Up to 22% of contact pixels land in LUT bins never observed during
25
- calibration</b> (column 6, dark = value invented by nearest-neighbour fill).
26
- It grows with depth: star 2% 22% from shallow to deep. Steep rim colours at
27
- large indentation are outside the calibrated set, so their gradient is made up.
28
- <b>2. Flat-topped indenters reconstruct as domes, not plateaus.</b> The
29
- centre/rim height ratio should be ≈1.0 for a flat punch; measured 1.23–1.42 for
30
- star / triangle / quad / B. The sphere's 1.43–1.46 is correct — for a sphere a
31
- dome IS the right answer, which is why the same number means different things
32
- per row. The interior of a flat contact carries no colour change, so Poisson
33
- can only raise it by integrating inward from the rim.
34
- <b>3. The valid mask is halo-dominated</b> (column 5): for star/triangle/quad
35
- it is a round blob while column 3 shows the shape crisply. The mask threshold
36
- <code>|dI| &gt; 8</code> admits the diffuse halo, and the halo pedestal then
37
- has to be removed downstream (column 12) rather than never being integrated.</p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
  <h2>Per-sample diagnostics</h2>
39
- <table><tr><th>sample</th><th>peak [mm]</th><th>unobserved LUT</th>
40
- <th>centre/rim</th><th>grad angle<br><span class="note">chance 90°</span></th>
41
- <th>profile RMS</th></tr><tr><td>round z=2.12mm F=4.4N</td><td>1.60</td><td>0%</td><td>1.46</td><td>6.8°</td><td>69 µm</td></tr><tr><td>round z=3.21mm F=9.4N</td><td>2.55</td><td>4%</td><td>1.43</td><td>5.8°</td><td>35 µm</td></tr><tr><td>round z=3.80mm F=14.8N</td><td>2.94</td><td>1%</td><td>1.44</td><td>5.1°</td><td>105 µm</td></tr><tr><td>star z=1.73mm F=3.0N</td><td>1.61</td><td>2%</td><td>1.44</td><td></td><td></td></tr><tr><td>star z=2.34mm F=8.0N</td><td>2.07</td><td>10%</td><td>1.33</td><td></td><td></td></tr><tr><td>star z=2.96mm F=11.0N</td><td>2.03</td><td>14%</td><td>1.36</td><td></td><td></td></tr><tr><td>star z=3.53mm F=16.0N</td><td>2.30</td><td>22%</td><td>1.35</td><td></td><td></td></tr><tr><td>triangle z=2.28mm F=4.9N</td><td>1.40</td><td>5%</td><td>1.42</td><td></td><td></td></tr><tr><td>triangle z=3.11mm F=9.7N</td><td>2.05</td><td>9%</td><td>1.31</td><td></td><td></td></tr><tr><td>triangle z=3.75mm F=14.7N</td><td>2.66</td><td>10%</td><td>1.32</td><td></td><td></td></tr><tr><td>quad z=2.16mm F=3.6N</td><td>1.19</td><td>4%</td><td>1.32</td><td></td><td></td></tr><tr><td>quad z=2.94mm F=9.2N</td><td>1.75</td><td>9%</td><td>1.24</td><td></td><td></td></tr><tr><td>quad z=3.52mm F=14.2N</td><td>2.32</td><td>10%</td><td>1.23</td><td></td><td></td></tr><tr><td>quad_small z=1.51mm F=3.9N</td><td>1.43</td><td>7%</td><td>1.32</td><td></td><td></td></tr><tr><td>quad_small z=2.55mm F=9.2N</td><td>2.51</td><td>6%</td><td>1.35</td><td></td><td></td></tr><tr><td>quad_small z=3.47mm F=13.7N</td><td>2.93</td><td>8%</td><td>1.35</td><td></td><td></td></tr><tr><td>quad_small z=3.97mm F=17.4N</td><td>2.66</td><td>12%</td><td>1.42</td><td></td><td></td></tr><tr><td>B z=2.91mm F=4.9N</td><td>1.96</td><td>7%</td><td>1.39</td><td></td><td></td></tr><tr><td>B z=4.04mm F=10.6N</td><td>2.83</td><td>6%</td><td>1.39</td><td></td><td></td></tr><tr><td>B z=4.48mm F=13.7N</td><td>2.87</td><td>5%</td><td>1.40</td><td></td><td></td></tr></table>
 
42
  <p class="note">Grad angle and profile RMS are only defined where the indenter
43
  is a sphere of known radius (the <code>round</code> family): the LUT gradient
44
- sits 5.1–6.8° from the analytic sphere gradient and the radial profile matches
45
- the cap to 35–105 µm, so on this sensor the table and the solver are both
46
- sound the defects above are about coverage, mask and flat-top geometry, not
47
- about the photometric map being wrong.</p>
48
  <footer style="margin-top:30px;color:#6f7a8c;font-size:13px">
49
  React force recovery · <a href="index.html" style="color:#ffd9a0">overview</a> ·
50
  <a href="method.html" style="color:#ffd9a0">method</a> ·
51
  <a href="results.html" style="color:#ffd9a0">results</a></footer>
52
- <img src='assets/recon/glowtact_00.png' loading='lazy'><br><img src='assets/recon/glowtact_01.png' loading='lazy'><br><img src='assets/recon/glowtact_02.png' loading='lazy'><br><img src='assets/recon/glowtact_03.png' loading='lazy'><br><img src='assets/recon/glowtact_04.png' loading='lazy'><br><img src='assets/recon/glowtact_05.png' loading='lazy'><br><img src='assets/recon/glowtact_06.png' loading='lazy'><br><img src='assets/recon/glowtact_07.png' loading='lazy'><br><img src='assets/recon/glowtact_08.png' loading='lazy'><br><img src='assets/recon/glowtact_09.png' loading='lazy'><br><img src='assets/recon/glowtact_10.png' loading='lazy'><br><img src='assets/recon/glowtact_11.png' loading='lazy'><br><img src='assets/recon/glowtact_12.png' loading='lazy'><br><img src='assets/recon/glowtact_13.png' loading='lazy'><br><img src='assets/recon/glowtact_14.png' loading='lazy'><br><img src='assets/recon/glowtact_15.png' loading='lazy'><br><img src='assets/recon/glowtact_16.png' loading='lazy'><br><img src='assets/recon/glowtact_17.png' loading='lazy'><br><img src='assets/recon/glowtact_18.png' loading='lazy'><br><img src='assets/recon/glowtact_19.png' loading='lazy'><br></div></body></html>
 
 
9
  table{border-collapse:collapse;margin:10px 0;font-size:13px}
10
  th,td{border:1px solid #2c3648;padding:5px 10px;text-align:left}
11
  th{color:#ffd9a0;font-weight:600}code{background:#1b2334;padding:1px 5px;border-radius:3px}
12
+ .note{color:#8e99ab;font-size:13px}
13
+ .good{color:#7be0a0}.bad{color:#ff9a8a}</style></head><body><div class="wrap">
14
  <h1>3D reconstruction workbench — every stage, 20 samples</h1>
15
  <p>Not a results figure: this is the full chain with every tunable quantity
16
  exposed, so a defect can be attributed to a stage. Columns:
 
20
  <b>11</b> depth · <b>12</b> depth halo-removed · <b>13</b> radial profile vs
21
  analytic cap · <b>14</b> Open3D mesh.
22
  Reproduce: <code>xvfb-run -a -s "-screen 0 1400x1000x24" python -m
23
+ force_recovery.recon_study glowtact</code>, then
24
+ <code>python -m force_recovery.recon_study page</code>.
25
+ The source here (GlowTact) is markerless, so the adopted marker step
26
+ (<code>marker_removal.stages_depth</code>) is a bit-exact no-op on these
27
+ frames; on a marker gel it inpaints the dots out of the reference and the
28
+ frame before column 3.</p>
29
+
30
+ <h2>What external ground truth did to this page</h2>
31
+ <p>Everything below used to be scored against our own expectations. It is now
32
+ scored against exact per-pixel ground truth ray-cast mesh depth on 420
33
+ Tactile MNIST touches (<code>mnist_validation</code>) and two of the three
34
+ defects we published moved:</p>
35
+ <table><tr><th>claim</th><th>status after per-pixel GT</th></tr>
36
+ <tr><td>Flat-topped indenters over-dome badly (centre/rim 1.23–1.42 where
37
+ 1.0 is correct)</td>
38
+ <td class="bad"><b>retracted and re-measured.</b> Compliant gel wraps a flat
39
+ edge, so c/r &gt; 1 is <i>expected</i>: the true depth map of a pressed digit
40
+ has c/r <b>1.400</b> (gel surface 1.334) and we reconstruct 1.539–1.562
41
+ (<b>+10–12%</b>); on an enclosed plateau control whose truth is exactly 1.000
42
+ we measure <b>1.069</b> (<b>+7%</b>). Most of the 1.23–1.42 in the table
43
+ below is real curvature, not artefact.</td></tr>
44
+ <tr><td>Up to 22% of contact pixels land in unobserved LUT bins — a top
45
+ defect</td>
46
+ <td class="bad"><b>demoted to a minor factor.</b> On the GT set 13.8% / 16.9%
47
+ of contact pixels are unobserved, and their correlation with per-touch
48
+ Type-2 error is only <b>0.098 / 0.294</b>. Still worth showing (column 6),
49
+ no longer worth blaming.</td></tr>
50
+ <tr><td>The valid mask is halo-dominated</td>
51
+ <td class="good"><b>confirmed and quantified.</b> Against the true contact
52
+ region the <code>|dI| &gt; 8</code> mask scores IoU <b>0.614</b>, recall
53
+ <b>0.917</b>, over-segmentation <b>0.531</b> — it finds nearly all of the
54
+ contact and then adds half as much again in halo.</td></tr>
55
+ <tr><td>(new) the photometric table is the weak link</td>
56
+ <td class="good"><b>ruled out.</b> LUT gradient direction vs true gel
57
+ gradient is <b>24.4°</b> on the GT renders against <b>26.1°</b> for this
58
+ table on its own real sensor — the same statistic, no worse off-domain.</td>
59
+ </tr></table>
60
+
61
+ <h2>The real headline: accuracy is a function of press depth</h2>
62
+ <p>Same digit meshes, re-rendered at five penetrations, no per-frame fitting
63
+ of any kind (the photometric table is calibrated once on that sensor's own
64
+ sphere presses, the recipe we already use per sensor):</p>
65
+ <table><tr><th>press depth [mm]</th><th>0.30</th><th>0.60</th><th>1.00</th>
66
+ <th>1.50</th><th>2.25 (what the dataset ships)</th></tr>
67
+ <tr><td>MAE [µm]</td><td class="good"><b>11.2</b></td><td class="good">35.0</td>
68
+ <td>67.8</td><td>127.4</td><td class="bad">281.1</td></tr>
69
+ <tr><td>Type-2 error [µm]</td><td class="good"><b>96.5</b></td><td>186.3</td>
70
+ <td>308.6</td><td>514.6</td><td class="bad">961.8</td></tr>
71
+ <tr><td>peak ours / GT</td><td>1.00</td><td>0.97</td><td>0.77</td><td>0.68</td>
72
+ <td class="bad">0.55</td></tr></table>
73
+ <p>So the honest public claim is a <b>range</b>, not a number: at ≤0.6 mm this
74
+ reconstruction is accurate on non-spherical ground truth with zero fitting; by
75
+ 2.25 mm it recovers barely half the peak. No accuracy figure on this site
76
+ should be quoted without the press depth it was measured at.</p>
77
+
78
  <h2>Per-sample diagnostics</h2>
79
+ <p class="note">centre/rim is reported <b>descriptively</b> — read it against
80
+ the GT ratio for that geometry (1.33–1.40 for a compliant press), never
81
+ against 1.0.</p>
82
+ <table><tr><th>sample</th><th>peak [mm]</th><th>unobserved LUT</th><th>centre/rim<br><span class='note'>GT for a compliant press: 1.33–1.40</span></th><th>grad angle<br><span class='note'>chance 90°</span></th><th>profile RMS</th></tr><tr><td>round z=2.12mm F=4.4N</td><td>1.60</td><td>0%</td><td>1.46</td><td>6.8</td><td>69</td></tr><tr><td>round z=3.21mm F=9.4N</td><td>2.55</td><td>4%</td><td>1.43</td><td>5.8</td><td>35</td></tr><tr><td>round z=3.80mm F=14.8N</td><td>2.94</td><td>1%</td><td>1.44</td><td>5.1</td><td>105</td></tr><tr><td>star z=1.73mm F=3.0N</td><td>1.61</td><td>2%</td><td>1.44</td><td></td><td></td></tr><tr><td>star z=2.34mm F=8.0N</td><td>2.07</td><td>10%</td><td>1.33</td><td></td><td></td></tr><tr><td>star z=2.96mm F=11.0N</td><td>2.03</td><td>14%</td><td>1.36</td><td></td><td></td></tr><tr><td>star z=3.53mm F=16.0N</td><td>2.30</td><td>22%</td><td>1.35</td><td></td><td></td></tr><tr><td>triangle z=2.28mm F=4.9N</td><td>1.40</td><td>5%</td><td>1.42</td><td></td><td></td></tr><tr><td>triangle z=3.11mm F=9.7N</td><td>2.05</td><td>9%</td><td>1.31</td><td></td><td></td></tr><tr><td>triangle z=3.75mm F=14.7N</td><td>2.66</td><td>10%</td><td>1.32</td><td></td><td></td></tr><tr><td>quad z=2.16mm F=3.6N</td><td>1.19</td><td>4%</td><td>1.32</td><td></td><td></td></tr><tr><td>quad z=2.94mm F=9.2N</td><td>1.75</td><td>9%</td><td>1.24</td><td></td><td></td></tr><tr><td>quad z=3.52mm F=14.2N</td><td>2.32</td><td>10%</td><td>1.23</td><td></td><td></td></tr><tr><td>quad_small z=1.51mm F=3.9N</td><td>1.43</td><td>7%</td><td>1.32</td><td></td><td></td></tr><tr><td>quad_small z=2.55mm F=9.2N</td><td>2.51</td><td>6%</td><td>1.35</td><td></td><td></td></tr><tr><td>quad_small z=3.47mm F=13.7N</td><td>2.93</td><td>8%</td><td>1.35</td><td></td><td></td></tr><tr><td>quad_small z=3.97mm F=17.4N</td><td>2.66</td><td>12%</td><td>1.42</td><td></td><td></td></tr><tr><td>B z=2.91mm F=4.9N</td><td>1.96</td><td>7%</td><td>1.39</td><td></td><td></td></tr><tr><td>B z=4.04mm F=10.6N</td><td>2.83</td><td>6%</td><td>1.39</td><td></td><td></td></tr><tr><td>B z=4.48mm F=13.7N</td><td>2.87</td><td>5%</td><td>1.40</td><td></td><td></td></tr></table>
83
  <p class="note">Grad angle and profile RMS are only defined where the indenter
84
  is a sphere of known radius (the <code>round</code> family): the LUT gradient
85
+ sits 5.1–6.8° from the analytic sphere gradient and the radial profile
86
+ matches the cap to 35–105 µm, so on this sensor the table and the solver
87
+ are both sound. Every one of these frames is a deep press (1.5–4.5 mm)
88
+ i.e. the regime the sweep above shows is our worst.</p>
89
  <footer style="margin-top:30px;color:#6f7a8c;font-size:13px">
90
  React force recovery · <a href="index.html" style="color:#ffd9a0">overview</a> ·
91
  <a href="method.html" style="color:#ffd9a0">method</a> ·
92
  <a href="results.html" style="color:#ffd9a0">results</a></footer>
93
+ <img src='assets/recon/glowtact_00.png' loading='lazy'><br><img src='assets/recon/glowtact_01.png' loading='lazy'><br><img src='assets/recon/glowtact_02.png' loading='lazy'><br><img src='assets/recon/glowtact_03.png' loading='lazy'><br><img src='assets/recon/glowtact_04.png' loading='lazy'><br><img src='assets/recon/glowtact_05.png' loading='lazy'><br><img src='assets/recon/glowtact_06.png' loading='lazy'><br><img src='assets/recon/glowtact_07.png' loading='lazy'><br><img src='assets/recon/glowtact_08.png' loading='lazy'><br><img src='assets/recon/glowtact_09.png' loading='lazy'><br><img src='assets/recon/glowtact_10.png' loading='lazy'><br><img src='assets/recon/glowtact_11.png' loading='lazy'><br><img src='assets/recon/glowtact_12.png' loading='lazy'><br><img src='assets/recon/glowtact_13.png' loading='lazy'><br><img src='assets/recon/glowtact_14.png' loading='lazy'><br><img src='assets/recon/glowtact_15.png' loading='lazy'><br><img src='assets/recon/glowtact_16.png' loading='lazy'><br><img src='assets/recon/glowtact_17.png' loading='lazy'><br><img src='assets/recon/glowtact_18.png' loading='lazy'><br><img src='assets/recon/glowtact_19.png' loading='lazy'><br>
94
+ </div></body></html>
results.html CHANGED
@@ -57,11 +57,12 @@ th,td{white-space:normal;min-width:110px}}
57
 
58
  <header>
59
  <div class="kicker">React force recovery · results</div>
60
- <h1>Three estimators × three ground-truth datasets</h1>
61
  <p class="sub">Every method evaluated on every dataset with force-sensor labels,
62
  predicted vs ground truth per dataset — so dataset quality is controlled within
63
  each row, and differences between panels are differences between methods.</p>
64
  <a class="pill" href="method.html">↖ how the method is designed</a>
 
65
  <a class="pill" href="index.html">overview</a>
66
  <a class="pill" href="gallery.html">gallery</a>
67
  <a class="pill" href="results_zh.html">中文</a>
@@ -74,15 +75,43 @@ each row, and differences between panels are differences between methods.</p>
74
  <th>FeelAnyForce<br>(trained: markerless)</th></tr>
75
  <tr><td>FEATS val (marker)</td><td>0.77</td>
76
  <td class="best">0.96 · in-domain</td><td class="dead">0.43</td></tr>
77
- <tr><td>FoTa cnc_Mini (markerless)</td><td>0.94 (in view)</td>
78
  <td class="dead">0.07</td><td class="best">0.83</td></tr>
79
- <tr><td>GlowTact (markerless)</td><td>0.98</td>
80
  <td class="dead">0.04</td><td class="best">0.90</td></tr>
 
 
 
81
  </table>
82
  <p>The pattern is the finding: <b>each network dominates its own gel domain and
83
  collapses outside it</b> (FEATS 0.96 → 0.04–0.07; FeelAnyForce 0.90 → 0.43),
84
  while <b>the physics pipeline is the only estimator that works everywhere</b>
85
- (0.74–0.99) — it never sees training data, so it has no domain to leave.</p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86
  </div>
87
 
88
  <h2>FEATS dataset — marker-dot gel</h2>
@@ -90,14 +119,40 @@ while <b>the physics pipeline is the only estimator that works everywhere</b>
90
  <p class="footnote">In-domain, the FEATS U-net is excellent (ρ=0.96) — the
91
  negative results elsewhere are domain effects, not a weak model. FeelAnyForce,
92
  markerless-trained, degrades on the dotted gel (0.43): the same knife cuts both
93
- ways.</p></div>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
94
 
95
  <h2>FoTa cnc_Mini — markerless gel</h2>
96
  <div class="card"><img src="assets/results_cnc.png" alt="cnc_Mini panels">
97
  <p class="footnote">Hard conditions: only 4 contact-free frames, 62% of presses
98
  near the pad border — the press grid is larger than the field of view (see
99
  <a href="debug_pipeline.html">pipeline debug</a>). Strictly in view our ρ is
100
- 0.94 (MAE 0.26 N); FeelAnyForce reaches 0.92 non-edge.</p></div>
101
 
102
  <h2>GlowTact — markerless gel, cleaned</h2>
103
  <div class="card"><img src="assets/results_glowtact.png" alt="GlowTact panels">
@@ -144,6 +199,85 @@ frame indices than forces (silently drifting the labels, &rho;&asymp;0 until
144
  paired within each trajectory), sharp/batch_2 is stored BGR while the other
145
  nine are RGB, and flat/batch_2 ships only 3 of 4 image files.</p></div>
146
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
147
  <h2>React — no ground truth, so: do independent methods agree?</h2>
148
  <div class="card"><img src="assets/results_react.png" alt="React agreement">
149
  <p class="footnote">On the dataset we actually care about, the two surviving
@@ -156,15 +290,24 @@ contact-free. Neither can copy the other's mistakes.</p></div>
156
  Mini): FeelAnyForce as the primary labeller, the physics pipeline as an
157
  independent audit, disagreement rows flagged. <b>For any new gel or sensor</b>
158
  where no trained model matches the domain: the physics pipeline is the only
159
- option that works out of the box — and its FEATS-dataset score (0.74) shows
160
- what it does on a domain nobody tuned it for.</p></div>
 
 
 
 
 
 
 
 
161
 
162
  <p><a href="debug_pipeline.html"><b>Pipeline debug page</b></a>: raw
163
  image &rarr; force step by step on all three datasets, and the cnc
164
- field-of-view ablation (in-view &rho;=0.94).</p>
165
 
166
  <footer>React force recovery · <a href="index.html">overview</a> ·
167
  <a href="method.html">method design</a> · <a href="gallery.html">gallery</a> ·
168
  data: FEATS (2411.03315) · FoTa/T3 (2406.13640) · GlowTact
169
- (dacongming666/GlowTact_Datasets) · FeelAnyForce (2410.02048)</footer>
 
170
  </div></body></html>
 
57
 
58
  <header>
59
  <div class="kicker">React force recovery · results</div>
60
+ <h1>Three estimators × four ground-truth datasets</h1>
61
  <p class="sub">Every method evaluated on every dataset with force-sensor labels,
62
  predicted vs ground truth per dataset — so dataset quality is controlled within
63
  each row, and differences between panels are differences between methods.</p>
64
  <a class="pill" href="method.html">↖ how the method is designed</a>
65
+ <a class="pill" href="recon_workbench.html">3D workbench</a>
66
  <a class="pill" href="index.html">overview</a>
67
  <a class="pill" href="gallery.html">gallery</a>
68
  <a class="pill" href="results_zh.html">中文</a>
 
75
  <th>FeelAnyForce<br>(trained: markerless)</th></tr>
76
  <tr><td>FEATS val (marker)</td><td>0.77</td>
77
  <td class="best">0.96 · in-domain</td><td class="dead">0.43</td></tr>
78
+ <tr><td>FoTa cnc_Mini (markerless)</td><td>0.95 (in view)</td>
79
  <td class="dead">0.07</td><td class="best">0.83</td></tr>
80
+ <tr><td>GlowTact (markerless)</td><td>0.99</td>
81
  <td class="dead">0.04</td><td class="best">0.90</td></tr>
82
+ <tr><td>Sparsh / Meta (markerless, 10 gel pads)</td>
83
+ <td>0.97 (in view, self-calibrated table)</td>
84
+ <td colspan="2">not run &mdash; no published predictions</td></tr>
85
  </table>
86
  <p>The pattern is the finding: <b>each network dominates its own gel domain and
87
  collapses outside it</b> (FEATS 0.96 → 0.04–0.07; FeelAnyForce 0.90 → 0.43),
88
  while <b>the physics pipeline is the only estimator that works everywhere</b>
89
+ (0.77&ndash;0.99) — it never sees training data, so it has no domain to leave.</p>
90
+ </div>
91
+
92
+ <h2>Every number beside the control that could have killed it</h2>
93
+ <div class="card">
94
+ <p style="margin-top:0">Re-run end to end after marker inpainting was added to
95
+ the depth pipeline. One protocol on all four datasets: within each group
96
+ (indenter family / probe / capture group / gel pad) half the frames fit a
97
+ 5-feature least-squares model calibrated by isotonic regression, the other half
98
+ are scored; 5 seeds, median reported. The control column repeats the identical
99
+ protocol with the force labels <b>permuted within each group</b> — group
100
+ structure and force distribution untouched, only the frame-to-force pairing
101
+ destroyed.</p>
102
+ <table class="matrix"><tr><th>dataset</th><th>n (eval)</th><th>&rho;</th><th>&rho; across seeds</th><th>MAE [N]</th><th>within-group shuffle</th></tr><tr><td>GlowTact (markerless, 0-20 N)</td><td>201</td><td class='best'><b>0.986</b></td><td>0.981&ndash;0.987</td><td>0.525</td><td>+0.171</td></tr><tr><td>FoTa cnc_Mini (markerless, in view)</td><td>337</td><td class='best'><b>0.946</b></td><td>0.929&ndash;0.949</td><td>0.252</td><td>+0.056</td></tr><tr><td>FEATS (marker gel)</td><td>186</td><td class='best'><b>0.775</b></td><td>0.713&ndash;0.787</td><td>5.025</td><td>-0.003</td></tr><tr><td>Sparsh / Meta (markerless, Sparsh LUT, in view)</td><td>1667</td><td class='best'><b>0.968</b></td><td>0.967&ndash;0.971</td><td>0.042</td><td>+0.264</td></tr><tr><td>FEATS (marker gel, dots inpainted — rejected for force)</td><td>186</td><td class='dead'><b>0.737</b></td><td>0.682&ndash;0.816</td><td>4.934</td><td>-0.010</td></tr></table>
103
+ <p class="footnote">Nothing moved. That is the expected result and it is worth
104
+ being explicit about: marker inpainting was adopted for <b>geometry only</b>,
105
+ so the force path was deliberately left byte-identical, and a spot check that
106
+ recomputes cached features from the raw frames confirms it
107
+ (max |cached − fresh| = 0 over 40 cnc frames). The last row is the same FEATS
108
+ frames and the same splits with the marker-inpainted features fed to the force
109
+ model instead — it <b>loses</b> 0.037 ρ, which is why it did not ship there.</p>
110
+ <p class="footnote">A within-group shuffle is the right control here, not a
111
+ global one: on FeelAnyForce the pooled ρ survived <i>global</i> shuffling at
112
+ 0.442 vs 0.455, which is how we caught that its frame join had never been
113
+ demonstrated. Reading the table: cnc and GlowTact sit ~0.9 above their
114
+ controls; FEATS sits 0.78 above a control that is flat at 0.00.</p>
115
  </div>
116
 
117
  <h2>FEATS dataset — marker-dot gel</h2>
 
119
  <p class="footnote">In-domain, the FEATS U-net is excellent (ρ=0.96) — the
120
  negative results elsewhere are domain effects, not a weak model. FeelAnyForce,
121
  markerless-trained, degrades on the dotted gel (0.43): the same knife cuts both
122
+ ways.</p>
123
+ <h3>Removing the marker dots: a geometry win, not a force win</h3>
124
+ <img src="assets/feats_marker_removal.png" alt="FEATS before and after marker inpainting">
125
+ <p class="footnote">The dots occlude the gel, so the photometric table has no
126
+ valid colour under them and Poisson integrates a dimple lattice into the depth
127
+ map — visible as pockmarks all over the <b>3D mesh BEFORE</b> panels. Detecting
128
+ the dots on the reference and inpainting them out of <b>both</b> the reference
129
+ and the frame before differencing (cv2 Telea, the image-space cousin of
130
+ GelSight Wedge's Fig. 10 hole interpolation) removes it: lattice power at the
131
+ 31.9 px marker pitch drops <b>1.523 → 0.890</b> (×0.65), lower on <b>91%</b> of
132
+ 120 frames above 1 N, Wilcoxon p = 2.6e-19. The detector is marker-specific,
133
+ not a blob finder — 63/63 dots and 0 rejects on this reference, stable for
134
+ every threshold from 3 to 16 grey levels, and exactly <b>0</b> blobs on the
135
+ markerless GlowTact and cnc references, where the step is a bit-exact no-op.</p>
136
+ <p class="footnote">Two things it does <b>not</b> do. It does not help force:
137
+ on identical frames and splits ρ goes 0.7747 → 0.7371 and every paired median
138
+ delta is negative, so the force features still come from the untouched
139
+ pipeline. And it does not catch every dot — the dots shear with the gel
140
+ (median 1.7 px, &gt;8 px on 8% of frames), so a static reference mask leaves
141
+ the displaced ones behind, which is what the residual dots in column 3 are.
142
+ Controls: inpainting the same <i>area</i> of randomly placed fake markers gives
143
+ 0.7697, i.e. no gain, so the small changes are not "inpainting = smoothing".
144
+ What actually caps FEATS is the reference, not the dots — on the 20 lightest
145
+ presses |dI| already sits at 11 grey levels off-dot and 88% of off-dot pixels
146
+ pass the |dI|&gt;8 valid test, so the mask is nearly the whole frame and the
147
+ features integrate reference mismatch; a per-indenter light-press reference
148
+ made it worse still (0.7747 → 0.7261).</p></div>
149
 
150
  <h2>FoTa cnc_Mini — markerless gel</h2>
151
  <div class="card"><img src="assets/results_cnc.png" alt="cnc_Mini panels">
152
  <p class="footnote">Hard conditions: only 4 contact-free frames, 62% of presses
153
  near the pad border — the press grid is larger than the field of view (see
154
  <a href="debug_pipeline.html">pipeline debug</a>). Strictly in view our ρ is
155
+ 0.95 (MAE 0.25 N); FeelAnyForce reaches 0.92 non-edge.</p></div>
156
 
157
  <h2>GlowTact — markerless gel, cleaned</h2>
158
  <div class="card"><img src="assets/results_glowtact.png" alt="GlowTact panels">
 
199
  paired within each trajectory), sharp/batch_2 is stored BGR while the other
200
  nine are RGB, and flat/batch_2 ships only 3 of 4 image files.</p></div>
201
 
202
+ <h2>First external per-pixel ground truth — and it moves the headline</h2>
203
+ <div class="card">
204
+ <p style="margin-top:0">Everything above scores <b>force</b>. Until now the
205
+ <b>depth</b> underneath was only ever checked against our own analytic sphere
206
+ cap, with the cap's amplitude anchored on the reconstruction itself. Tactile
207
+ MNIST supplies what was missing: exact per-pixel depth, ray-cast from the
208
+ 3D-printed digit meshes, on <b>non-spherical</b> geometry a sphere calibration
209
+ cannot self-validate — 420 touches, 106 objects. The pose bookkeeping is
210
+ verified end to end rather than assumed (re-rendering the ground-truth height
211
+ map reproduces the shipped image to ~2/255 grey levels).</p>
212
+ <img src="assets/mnist_examples.png" alt="reconstruction vs exact mesh ground truth">
213
+ <p><b>The finding is a range, not a number: accuracy is a steep function of
214
+ press depth.</b> Same digit meshes, re-rendered at five penetrations, no
215
+ per-frame alignment and no fitted indentation scale:</p>
216
+ <table class="matrix">
217
+ <tr><th>press depth [mm]</th><th>0.30</th><th>0.60</th><th>1.00</th><th>1.50</th>
218
+ <th>2.25 — what the dataset ships</th></tr>
219
+ <tr><td>MAE [µm]</td><td class="best"><b>11.2</b></td><td class="best">35.0</td>
220
+ <td>67.8</td><td>127.4</td><td class="dead">281.1</td></tr>
221
+ <tr><td>Type-2 error [µm]</td><td class="best"><b>96.5</b></td><td>186.3</td>
222
+ <td>308.6</td><td>514.6</td><td class="dead">961.8</td></tr>
223
+ <tr><td>peak recovered (ours / GT)</td><td>1.00</td><td>0.97</td><td>0.77</td>
224
+ <td>0.68</td><td class="dead">0.55</td></tr>
225
+ </table>
226
+ <p class="footnote">At 0.3 mm the Type-2 error is <b>96.5 µm</b>, below every
227
+ one of 3D Cal's three published figures (152.8 / 171.6 / 290.0 µm) — and
228
+ theirs are reported <i>with</i> a 2D cross-correlation alignment and a fitted
229
+ indentation scale, ours with neither. At 0.6 mm (186.3 µm) we sit inside their
230
+ range. At the 2.25 mm press this dataset actually ships we are an order of
231
+ magnitude worse and recover barely half the peak. <b>So: no accuracy number on
232
+ this site should be quoted without the press depth it was measured at, and the
233
+ working range of this reconstruction is shallow contact.</b></p>
234
+ <p><b>What it corrects, what it confirms.</b></p>
235
+ <table class="matrix">
236
+ <tr><th>earlier claim</th><th>after per-pixel ground truth</th></tr>
237
+ <tr><td>Flat-topped indenters over-dome badly — centre/rim 1.23&ndash;1.42
238
+ where 1.0 would be correct</td>
239
+ <td class="dead"><b>retracted.</b> The premise was wrong: compliant gel wraps
240
+ around a flat edge, so c/r &gt; 1 is <i>expected</i>. GT says the true depth
241
+ map of a pressed digit has c/r <b>1.400</b> (the gel surface itself 1.334)
242
+ while we reconstruct 1.539&ndash;1.562 &mdash; <b>+10&ndash;12%</b>; on an
243
+ enclosed plateau control whose truth is exactly 1.000 we measure <b>1.069</b>,
244
+ <b>+7%</b>. The over-doming is real and small, not the +23&ndash;42% we
245
+ implied.</td></tr>
246
+ <tr><td>Up to 22% of contact pixels land in unobserved LUT bins &mdash; a
247
+ leading defect</td>
248
+ <td class="dead"><b>demoted.</b> 13.8% / 16.9% on the GT set, correlating with
249
+ per-touch Type-2 error at only <b>0.098 / 0.294</b>.</td></tr>
250
+ <tr><td>The <code>|dI| &gt; 8</code> valid mask is halo-dominated</td>
251
+ <td class="best"><b>confirmed and quantified.</b> Against the true contact
252
+ region: IoU <b>0.614</b>, recall <b>0.917</b>, over-segmentation <b>0.531</b>
253
+ &mdash; it finds nearly all the contact and then adds half as much again in
254
+ halo.</td></tr>
255
+ <tr><td>The photometric table might be the weak link off-domain</td>
256
+ <td class="best"><b>ruled out.</b> LUT gradient vs true gel gradient is
257
+ <b>24.4&deg;</b> on these renders against <b>26.1&deg;</b> for the same table
258
+ on its own real sensor, and refitting the table on in-domain sphere renders
259
+ does not improve the digits (316.9 vs 273.4 µm).</td></tr>
260
+ </table>
261
+ <p><b>One failure mode, now seen on four datasets.</b> 420/420 of these touches
262
+ have a contact that runs off the pad, and a control that moves a single sphere
263
+ cap from mid-pad to the edge collapses its peak <b>1.39 &rarr; 0.30 mm</b>
264
+ against a 0.90 mm truth (Type-2 292 &rarr; 450 µm). That is the same effect as
265
+ cnc_Mini's press grid being larger than the field of view (&rho; 0.11 &rarr;
266
+ 0.94 in the field-of-view ablation) and Sparsh's clipped-disc frames
267
+ scoring worse despite carrying the highest median force. <b>Contact
268
+ visibility, not force range or gel type, is this pipeline's single biggest
269
+ external failure mode</b> &mdash; the Poisson solve's zero boundary cannot
270
+ represent a surface that leaves the frame.</p>
271
+ <p class="footnote"><b>Honest caveat.</b> These images are Taxim renders, not
272
+ real GelSight frames, so this validates the <i>geometry solver</i> &mdash;
273
+ table, mask, integration &mdash; rather than the sensor model, and the domain
274
+ check above is what licenses reading it that way. Taxim's gel also follows the
275
+ object geometry to within ~38 µm, so it barely models real gel
276
+ non-conformance: the true numbers on a physical sensor at the same press depth
277
+ will be worse, not better. Reproduce with
278
+ <code>python -m force_recovery.mnist_validation stage1|controls|sweep</code>.</p>
279
+ </div>
280
+
281
  <h2>React — no ground truth, so: do independent methods agree?</h2>
282
  <div class="card"><img src="assets/results_react.png" alt="React agreement">
283
  <p class="footnote">On the dataset we actually care about, the two surviving
 
290
  Mini): FeelAnyForce as the primary labeller, the physics pipeline as an
291
  independent audit, disagreement rows flagged. <b>For any new gel or sensor</b>
292
  where no trained model matches the domain: the physics pipeline is the only
293
+ option that works out of the box — and its FEATS-dataset score (0.77) shows
294
+ what it does on a domain nobody tuned it for.</p>
295
+ <p><b>And state the scope with it.</b> Those ρ are
296
+ <i>force</i> scores, per group. The <i>depth</i> underneath is now measured
297
+ against exact per-pixel ground truth and is a strong function of press depth —
298
+ 11 µm MAE at 0.3 mm, 281 µm at 2.25 mm — so the honest one-line summary is:
299
+ <b>accurate shallow-contact geometry, monotone force within a calibrated
300
+ group, and neither claim survives a contact that leaves the field of view.</b>
301
+ Marker gels get one extra step (dots inpainted before differencing) which
302
+ buys geometry and not force.</p></div>
303
 
304
  <p><a href="debug_pipeline.html"><b>Pipeline debug page</b></a>: raw
305
  image &rarr; force step by step on all three datasets, and the cnc
306
+ field-of-view ablation (in-view &rho;=0.95).</p>
307
 
308
  <footer>React force recovery · <a href="index.html">overview</a> ·
309
  <a href="method.html">method design</a> · <a href="gallery.html">gallery</a> ·
310
  data: FEATS (2411.03315) · FoTa/T3 (2406.13640) · GlowTact
311
+ (dacongming666/GlowTact_Datasets) · FeelAnyForce (2410.02048) ·
312
+ Tactile MNIST (TUDa-RL) · Sparsh (facebook/gelsight-force-estimation)</footer>
313
  </div></body></html>
results_zh.html CHANGED
@@ -57,10 +57,11 @@ th,td{white-space:normal;min-width:110px}}
57
 
58
  <header>
59
  <div class="kicker">React 力恢复 · 评测结果</div>
60
- <h1>三个估计器 × 个真值数据集</h1>
61
  <p class="sub">每个方法在每个有力传感器标注的数据集上评测,逐数据集画预测 vs 真值——
62
  数据集质量在行内被控制,面板之间的差异就是方法之间的差异。</p>
63
  <a class="pill" href="method_zh.html">↖ 方法设计</a>
 
64
  <a class="pill" href="index.html">总览(英文)</a>
65
  <a class="pill" href="gallery.html">图库</a>
66
  <a class="pill" href="results.html">English</a>
@@ -73,25 +74,62 @@ th,td{white-space:normal;min-width:110px}}
73
  <th>FeelAnyForce<br>(训练域:无点)</th></tr>
74
  <tr><td>FEATS val(有点)</td><td>0.77</td>
75
  <td class="best">0.96 · 本域</td><td class="dead">0.43</td></tr>
76
- <tr><td>FoTa cnc_Mini(无点)</td><td>0.94 (视野内)</td>
77
  <td class="dead">0.07</td><td class="best">0.83</td></tr>
78
- <tr><td>GlowTact(无点)</td><td>0.98</td>
79
  <td class="dead">0.04</td><td class="best">0.90</td></tr>
 
 
 
80
  </table>
81
  <p>规律本身就是发现:<b>每个网络只统治自己的 gel 域,出域即塌</b>
82
  (FEATS 0.96 → 0.04–0.07;FeelAnyForce 0.90 → 0.43);
83
- 而<b>物理管线是唯一在所有域都工作的估计器</b>(0.74–0.99)——
84
  它没见过训练数据,所以也没有可以离开的域。</p>
85
  </div>
86
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
87
  <h2>FEATS 数据集——有点 gel</h2>
88
  <div class="card"><img src="assets/results_feats.png" alt="FEATS dataset panels">
89
  <p class="footnote">本域内 FEATS U-net 非常出色(ρ=0.96)——它在别处的失败是域效应,
90
- 不是模型弱。FeelAnyForce(无点训练)在有点 gel 上退化到 0.43:这把刀两边都割。</p></div>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
91
 
92
  <h2>FoTa cnc_Mini——无点 gel</h2>
93
  <div class="card"><img src="assets/results_cnc.png" alt="cnc_Mini panels">
94
- <p class="footnote">条件苛刻:只有 4 帧无接触,62% 的按压贴边。我们严格视野内 ρ=0.94(MAE 0.26 N,见管线调试页);
95
  FeelAnyForce 非边缘达 0.92。</p></div>
96
 
97
  <h2>GlowTact——无点 gel,清洗过</h2>
@@ -127,6 +165,64 @@ FeelAnyForce 非边缘达 0.92。</p></div>
127
  逐轨迹配对前 &rho;&asymp;0)、sharp/batch_2 以 BGR 存储而其余九个是 RGB、
128
  flat/batch_2 只有 3 个图像文件。</p></div>
129
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
130
  <h2>React——没有真值,那就问:独立方法是否一致?</h2>
131
  <div class="card"><img src="assets/results_react.png" alt="React agreement">
132
  <p class="footnote">在我们真正关心的数据集上,两个幸存的估计器——物理(零训练)与
@@ -136,12 +232,18 @@ FeelAnyForce(20 万帧)——一致性 ρ=0.91,且物理管线判无接触的每
136
  <h2>结论</h2>
137
  <div class="card"><p style="margin-top:0"><b>给 React 打标签</b>(无点 Mini):FeelAnyForce 主标,
138
  物理管线独立审计,分歧行打标。<b>换任何新 gel 或新传感器</b>、没有训练模型匹配该域时:
139
- 物理管线是唯一开箱即用的选项——它在 FEATS 数据集上的 0.74 就是"无人为它调过的域"上的表现。</p></div>
 
 
 
 
 
140
 
141
- <p><a href="debug_pipeline.html"><b>管线调试页</b></a>:三个数据集上从原始图像到力的逐步展示,以及 cnc 视野消融(视野内 &rho;=0.94)。</p>
142
 
143
  <footer>React 力恢复 · <a href="index.html">总览</a> ·
144
  <a href="method_zh.html">方法设计</a> · <a href="gallery.html">图库</a> ·
145
  数据:FEATS (2411.03315) · FoTa/T3 (2406.13640) · GlowTact
146
- (dacongming666/GlowTact_Datasets) · FeelAnyForce (2410.02048)</footer>
 
147
  </div></body></html>
 
57
 
58
  <header>
59
  <div class="kicker">React 力恢复 · 评测结果</div>
60
+ <h1>三个估计器 × 个真值数据集</h1>
61
  <p class="sub">每个方法在每个有力传感器标注的数据集上评测,逐数据集画预测 vs 真值——
62
  数据集质量在行内被控制,面板之间的差异就是方法之间的差异。</p>
63
  <a class="pill" href="method_zh.html">↖ 方法设计</a>
64
+ <a class="pill" href="recon_workbench.html">3D 工作台</a>
65
  <a class="pill" href="index.html">总览(英文)</a>
66
  <a class="pill" href="gallery.html">图库</a>
67
  <a class="pill" href="results.html">English</a>
 
74
  <th>FeelAnyForce<br>(训练域:无点)</th></tr>
75
  <tr><td>FEATS val(有点)</td><td>0.77</td>
76
  <td class="best">0.96 · 本域</td><td class="dead">0.43</td></tr>
77
+ <tr><td>FoTa cnc_Mini(无点)</td><td>0.95 (视野内)</td>
78
  <td class="dead">0.07</td><td class="best">0.83</td></tr>
79
+ <tr><td>GlowTact(无点)</td><td>0.99</td>
80
  <td class="dead">0.04</td><td class="best">0.90</td></tr>
81
+ <tr><td>Sparsh / Meta(无点,10 块 gel pad)</td>
82
+ <td>0.97(视野内,自标定查找表)</td>
83
+ <td colspan="2">未运行 &mdash; 没有公开的预测结果</td></tr>
84
  </table>
85
  <p>规律本身就是发现:<b>每个网络只统治自己的 gel 域,出域即塌</b>
86
  (FEATS 0.96 → 0.04–0.07;FeelAnyForce 0.90 → 0.43);
87
+ 而<b>物理管线是唯一在所有域都工作的估计器</b>(0.77&ndash;0.99)——
88
  它没见过训练数据,所以也没有可以离开的域。</p>
89
  </div>
90
 
91
+ <h2>每个数字都配上足以否定它的对照</h2>
92
+ <div class="card">
93
+ <p style="margin-top:0">在深度管线加入 marker 修补之后端到端重跑。四个数据集使用同一套协议:
94
+ 在每个组内(压头族 / 探头 / 采集组 / gel pad)一半帧拟合 5 特征最小二乘模型并用保序回归标定,
95
+ 另一半用于评分;5 个随机种子,报告中位数。对照列用完全相同的协议,但把力标签
96
+ <b>在组内打乱</b>——组结构和力的分布都不变,只破坏帧与力的配对。</p>
97
+ <table class="matrix"><tr><th>dataset</th><th>n (eval)</th><th>&rho;</th><th>&rho; across seeds</th><th>MAE [N]</th><th>within-group shuffle</th></tr><tr><td>GlowTact (markerless, 0-20 N)</td><td>201</td><td class='best'><b>0.986</b></td><td>0.981&ndash;0.987</td><td>0.525</td><td>+0.171</td></tr><tr><td>FoTa cnc_Mini (markerless, in view)</td><td>337</td><td class='best'><b>0.946</b></td><td>0.929&ndash;0.949</td><td>0.252</td><td>+0.056</td></tr><tr><td>FEATS (marker gel)</td><td>186</td><td class='best'><b>0.775</b></td><td>0.713&ndash;0.787</td><td>5.025</td><td>-0.003</td></tr><tr><td>Sparsh / Meta (markerless, Sparsh LUT, in view)</td><td>1667</td><td class='best'><b>0.968</b></td><td>0.967&ndash;0.971</td><td>0.042</td><td>+0.264</td></tr><tr><td>FEATS (marker gel, dots inpainted — rejected for force)</td><td>186</td><td class='dead'><b>0.737</b></td><td>0.682&ndash;0.816</td><td>4.934</td><td>-0.010</td></tr></table>
98
+ <p class="footnote">没有任何数字变动。这正是预期结果,也值得明确说出来:marker 修补只被采纳用于
99
+ <b>几何</b>,力的路径被刻意保持逐字节不变;从原始帧重算缓存特征的抽查也证实了这一点
100
+ (40 帧 cnc 上 max |缓存 − 重算| = 0)。最后一行是同样的 FEATS 帧、同样的划分,
101
+ 只是把 marker 修补后的特征喂给力模型——ρ <b>下降</b> 0.037,这就是它没有进入力路径的原因。</p>
102
+ <p class="footnote">这里正确的对照是组内打乱,而不是全局打乱:在 FeelAnyForce 上,
103
+ 混合后的 ρ 在<i>全局</i>打乱下仍有 0.442(对比 0.455),我们正是这样发现它的帧对齐从未被验证。
104
+ 读表方式:cnc 与 GlowTact 高出各自对照约 0.9;FEATS 高出 0.78,而它的对照平在 0.00。</p>
105
+ </div>
106
+
107
  <h2>FEATS 数据集——有点 gel</h2>
108
  <div class="card"><img src="assets/results_feats.png" alt="FEATS dataset panels">
109
  <p class="footnote">本域内 FEATS U-net 非常出色(ρ=0.96)——它在别处的失败是域效应,
110
+ 不是模型弱。FeelAnyForce(无点训练)在有点 gel 上退化到 0.43:这把刀两边都割。</p>
111
+ <h3>去掉 marker 点:几何上的胜利,不是力上的</h3>
112
+ <img src="assets/feats_marker_removal.png" alt="FEATS marker 修补前后对比">
113
+ <p class="footnote">点会遮挡 gel,因此光度查找表在点下没有有效颜色,Poisson 积分会把一层
114
+ 点阵凹坑积进深度图——在 <b>3D mesh BEFORE</b> 面板上表现为满屏麻点。在参考帧上检测这些点,
115
+ 并在做差分<b>之前</b>把它们从参考帧<b>和</b>当前帧里一起���补掉(cv2 Telea,相当于 GelSight
116
+ Wedge 图 10 中孔洞插值的图像域版本),凹坑就消失了:31.9 px 点距对应频率上的功率从
117
+ <b>1.523 降到 0.890</b>(×0.65),在 1 N 以上的 120 帧中有 <b>91%</b> 变低,
118
+ Wilcoxon p = 2.6e-19。检测器是 marker 专用的,不是通用斑点检测器——在该参考帧上 63/63 个点、
119
+ 0 个误检,阈值从 3 到 16 灰阶都稳定;而在无 marker 的 GlowTact 与 cnc 参考帧上恰好检出
120
+ <b>0</b> 个,此时该步骤是逐比特的空操作。</p>
121
+ <p class="footnote">它<b>做不到</b>两件事。第一,它对力没有帮助:在相同帧、相同划分下
122
+ ρ 从 0.7747 降到 0.7371,且每个配对中位差都是负的,所以力特征仍然来自未改动的管线。
123
+ 第二,它抓不全所有点——点会随 gel 剪切移动(中位 1.7 px,8% 的帧超过 8 px),
124
+ 静态参考掩码追不上位移的点,第 3 列里残留的点就是它们。对照实验:把同样<i>面积</i>的
125
+ 随机假 marker 修补掉得到 0.7697,即没有增益,说明这些微小变化并不是"修补=平滑=更好"。
126
+ 真正卡住 FEATS 的是参考帧而非点——在最轻的 20 次按压上,点外区域的 |dI| 已经有 11 个灰阶,
127
+ 88% 的点外像素通过 |dI|&gt;8 的有效性判据,掩码几乎覆盖整幅画面,特征积分的是参考帧失配;
128
+ 换成逐压头的轻压参考帧反而更差(0.7747 → 0.7261)。</p></div>
129
 
130
  <h2>FoTa cnc_Mini——无点 gel</h2>
131
  <div class="card"><img src="assets/results_cnc.png" alt="cnc_Mini panels">
132
+ <p class="footnote">条件苛刻:只有 4 帧无接触,62% 的按压贴边。我们严格视野内 ρ=0.95(MAE 0.25 N,见管线调试页);
133
  FeelAnyForce 非边缘达 0.92。</p></div>
134
 
135
  <h2>GlowTact——无点 gel,清洗过</h2>
 
165
  逐轨迹配对前 &rho;&asymp;0)、sharp/batch_2 以 BGR 存储而其余九个是 RGB、
166
  flat/batch_2 只有 3 个图像文件。</p></div>
167
 
168
+ <h2>第一份外部逐像素真值——它改变了主结论</h2>
169
+ <div class="card">
170
+ <p style="margin-top:0">上面评的都是<b>力</b>。在此之前,底层的<b>深度</b>只被拿来和我们自己的
171
+ 解析球冠比较,而球冠的幅值还是从重建本身锚定的。Tactile MNIST 补上了缺失的一环:
172
+ 由 3D 打印数字网格光线投射得到的精确逐像素深度,且是球面标定无法自证的<b>非球面</b>几何——
173
+ 420 次触碰、106 个物体。位姿对应关系是端到端验证过的而非假设的
174
+ (用真值高度图重渲染可复现出厂图像,误差约 2/255 灰阶)。</p>
175
+ <img src="assets/mnist_examples.png" alt="重建与精确网格真值对比">
176
+ <p><b>结论是一个区间,不是一个数字:精度是压入深度的陡峭函数。</b>
177
+ 同样的数字网格,在五个压入深度上重渲染,没有逐帧配准,也没有拟合压入尺度:</p>
178
+ <table class="matrix">
179
+ <tr><th>压入深度 [mm]</th><th>0.30</th><th>0.60</th><th>1.00</th><th>1.50</th>
180
+ <th>2.25 — 数据集实际发布的</th></tr>
181
+ <tr><td>MAE [µm]</td><td class="best"><b>11.2</b></td><td class="best">35.0</td>
182
+ <td>67.8</td><td>127.4</td><td class="dead">281.1</td></tr>
183
+ <tr><td>Type-2 误差 [µm]</td><td class="best"><b>96.5</b></td><td>186.3</td>
184
+ <td>308.6</td><td>514.6</td><td class="dead">961.8</td></tr>
185
+ <tr><td>峰值还原比(我们 / 真值)</td><td>1.00</td><td>0.97</td><td>0.77</td>
186
+ <td>0.68</td><td class="dead">0.55</td></tr>
187
+ </table>
188
+ <p class="footnote">在 0.3 mm 处 Type-2 误差为 <b>96.5 µm</b>,低于 3D Cal 公布的全部三个数字
189
+ (152.8 / 171.6 / 290.0 µm)——而且他们的结果是<i>带</i>二维互相关配准和拟合压入尺度得到的,
190
+ 我们两者都没有。在 0.6 mm(186.3 µm)我们落在他们的区间内。而在该数据集实际使用的 2.25 mm
191
+ 压入下,我们差了一个数量级,峰值只还原了一半左右。<b>所以:本站任何精度数字都必须连同
192
+ 它所对应的压入深度一起引用,而这套重建的适用区间是浅接触。</b></p>
193
+ <p><b>它纠正了什么,又确认了什么。</b></p>
194
+ <table class="matrix">
195
+ <tr><th>此前的说法</th><th>逐像素真值之后</th></tr>
196
+ <tr><td>平头压头过度穹顶化——中心/边缘比 1.23&ndash;1.42,而 1.0 才是正确值</td>
197
+ <td class="dead"><b>撤回。</b>前提本身就是错的:gel 是柔顺的,会包裹平边,所以
198
+ c/r &gt; 1 是<i>应当出现</i>的。真值显示被压数字的真实深度图 c/r 为 <b>1.400</b>
199
+ (gel 表面本身是 1.334),而我们重建出 1.539&ndash;1.562——<b>+10&ndash;12%</b>;
200
+ 在真值恰为 1.000 的封闭平台对照上我们测得 <b>1.069</b>,即 <b>+7%</b>。
201
+ 过度穹顶化真实存在但很小,不是我们暗示的 +23&ndash;42%。</td></tr>
202
+ <tr><td>多达 22% 的接触像素落在未观测过的 LUT bin 里——一个��要缺陷</td>
203
+ <td class="dead"><b>降级。</b>在真值集上为 13.8% / 16.9%,
204
+ 与逐次触碰 Type-2 误差的相关性仅 <b>0.098 / 0.294</b>。</td></tr>
205
+ <tr><td><code>|dI| &gt; 8</code> 有效掩码被光晕主导</td>
206
+ <td class="best"><b>确认并量化。</b>相对真实接触区域:IoU <b>0.614</b>,
207
+ 召回 <b>0.917</b>,过分割 <b>0.531</b>——它几乎找全了接触,然后又多加了半个接触面积的光晕。</td></tr>
208
+ <tr><td>出域时光度查找表可能是薄弱环节</td>
209
+ <td class="best"><b>排除。</b>LUT 梯度与真实 gel 梯度的夹角在这些渲染上是 <b>24.4&deg;</b>,
210
+ 而同一张表在它自己的真实传感器上是 <b>26.1&deg;</b>;用域内球压重渲染重新拟合查找表
211
+ 也没有改善数字(316.9 vs 273.4 µm)。</td></tr>
212
+ </table>
213
+ <p><b>同一个失效模式,现在已在四个数据集上看到。</b>这里 420/420 次触碰的接触都跑出了 pad,
214
+ 而一个把单个球冠从 pad 中心移到边缘的对照,使峰值从 <b>1.39 塌到 0.30 mm</b>
215
+ (真值 0.90 mm,Type-2 292 &rarr; 450 µm)。这与 cnc_Mini 的按压网格大于视野
216
+ (视野消融实验里 &rho; 0.11 &rarr; 0.94)、以及 Sparsh 中被裁切的接触圆盘虽然力中位数最高
217
+ 却得分更差,是同一回事。<b>接触可见性——而不是力程或 gel 类型——才是这条管线最大的外部失效模式</b>:
218
+ Poisson 求解的零边界无法表示一个跑出画面的曲面。</p>
219
+ <p class="footnote"><b>诚实的边界。</b>这些图像是 Taxim 渲染,不是真实 GelSight 帧,
220
+ 所以它验证的是<i>几何求解器</i>(查找表、掩码、积分),而不是传感器模型;
221
+ 上面的域检查正是允许这样解读的依据。另外 Taxim 的 gel 与物体几何的偏差只有约 38 µm,
222
+ 几乎没有建模真实 gel 的非贴合性:在同样压深下,真实传感器上的数字只会更差,不会更好。
223
+ 复现:<code>python -m force_recovery.mnist_validation stage1|controls|sweep</code>。</p>
224
+ </div>
225
+
226
  <h2>React——没有真值,那就问:独立方法是否一致?</h2>
227
  <div class="card"><img src="assets/results_react.png" alt="React agreement">
228
  <p class="footnote">在我们真正关心的数据集上,两个幸存的估计器——物理(零训练)与
 
232
  <h2>结论</h2>
233
  <div class="card"><p style="margin-top:0"><b>给 React 打标签</b>(无点 Mini):FeelAnyForce 主标,
234
  物理管线独立审计,分歧行打标。<b>换任何新 gel 或新传感器</b>、没有训练模型匹配该域时:
235
+ 物理管线是唯一开箱即用的选项——它在 FEATS 数据集上的 0.77 就是"无人为它调过的域"上的表现。</p>
236
+ <p><b>并且要把适用范围一起说清楚。</b>上面那些 ρ 是逐组的<i>力</i>得分。
237
+ 底层的<i>深度</i>现在已经用精确逐像素真值测过,它是压入深度的强函数——
238
+ 0.3 mm 处 MAE 11 µm,2.25 mm 处 281 µm——所以诚实的一句话总结是:
239
+ <b>浅接触几何精确、标定组内力单调,而一旦接触跑出视野,两条结论都不成立。</b>
240
+ 有 marker 的 gel 多一步(差分前把点修补掉),它买到的是几何,不是力。</p></div>
241
 
242
+ <p><a href="debug_pipeline.html"><b>管线调试页</b></a>:三个数据集上从原始图像到力的逐步展示,以及 cnc 视野消融(视野内 &rho;=0.95)。</p>
243
 
244
  <footer>React 力恢复 · <a href="index.html">总览</a> ·
245
  <a href="method_zh.html">方法设计</a> · <a href="gallery.html">图库</a> ·
246
  数据:FEATS (2411.03315) · FoTa/T3 (2406.13640) · GlowTact
247
+ (dacongming666/GlowTact_Datasets) · FeelAnyForce (2410.02048) ·
248
+ Tactile MNIST (TUDa-RL) · Sparsh (facebook/gelsight-force-estimation)</footer>
249
  </div></body></html>