File size: 10,304 Bytes
3cbbecf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fa83a91
 
3cbbecf
 
 
 
 
 
 
 
 
 
553edff
 
 
 
3cbbecf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f7d1c58
3cbbecf
512969b
3cbbecf
f7d1c58
3cbbecf
 
 
 
 
512969b
3cbbecf
 
 
 
 
 
 
 
 
 
 
 
512969b
 
 
3cbbecf
 
 
 
fa83a91
 
 
 
 
3cbbecf
999caf9
 
55f7a18
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
999caf9
3cbbecf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ac68efd
 
 
 
3cbbecf
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
<!DOCTYPE html>
<html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width,initial-scale=1">
<title>Results β€” Force Estimation on GelSight Mini</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link href="https://fonts.googleapis.com/css2?family=Fraunces:opsz,wght@9..144,600;9..144,700&family=IBM+Plex+Mono:wght@400;500&family=IBM+Plex+Sans:wght@400;600&family=Noto+Serif+SC:wght@600;700&family=Noto+Sans+SC:wght@400;500&display=swap" rel="stylesheet">
<style>
:root{--paper:#0d1526;--ink:#dce7f5;--dim:#7c8db0;--grid:#1a2540;
--force:#ffb347;--target:#ff7847;--ok:#7be0a0;--card:#111b31;--line:#24345;}
*{box-sizing:border-box}
body{margin:0;background:var(--paper);color:var(--ink);
font-family:'IBM Plex Sans','Noto Sans SC',sans-serif;line-height:1.65;font-size:16px;
background-image:linear-gradient(var(--grid) 1px,transparent 1px),
linear-gradient(90deg,var(--grid) 1px,transparent 1px);background-size:44px 44px}
.wrap{max-width:880px;margin:0 auto;padding:0 24px 80px}
h1,h2{font-family:'Fraunces','Noto Serif SC',serif}
h1{font-size:2.1rem;font-weight:700;line-height:1.2;margin:0 0 10px}
h2{font-size:1.25rem;font-weight:600;margin:46px 0 8px;color:var(--force)}
.kicker{font-family:'IBM Plex Mono',monospace;font-size:.76rem;letter-spacing:.16em;
text-transform:uppercase;color:var(--force);margin-bottom:14px}
header{padding:60px 0 26px;border-bottom:1px dashed var(--line)}
.sub{color:var(--dim);max-width:720px}
.card{background:var(--card);border:1px solid var(--line);border-radius:10px;
padding:18px 22px;margin:14px 0}
.step{display:grid;grid-template-columns:34px 200px 1fr;gap:12px;
padding:10px 0;border-bottom:1px solid var(--line);align-items:baseline}
.step:last-child{border-bottom:none}
.step .n{font-family:'Fraunces',serif;font-size:1.25rem;color:var(--target)}
.step .t{font-family:'IBM Plex Mono',monospace;font-size:.85rem}
.step .d{color:var(--dim);font-size:.9rem}
table{border-collapse:collapse;width:100%;font-size:.88rem;margin:10px 0;
font-variant-numeric:tabular-nums}
th,td{padding:7px 10px;border-bottom:1px solid var(--line);text-align:right}
th{color:var(--dim);font-family:'IBM Plex Mono',monospace;font-size:.72rem;
text-transform:uppercase;letter-spacing:.08em}
td:first-child,th:first-child{text-align:left}
img,video{max-width:100%;height:auto;border-radius:8px;border:1px solid var(--line);
display:block;margin:12px auto}
a{color:#4fd8e0}a:hover{color:var(--force)}
.pill{display:inline-block;border:1px solid var(--line);border-radius:999px;
padding:3px 12px;font-size:.75rem;font-family:'IBM Plex Mono',monospace;
color:var(--dim);margin-right:8px;margin-top:12px;text-decoration:none}
code{font-family:'IBM Plex Mono',monospace;background:#0a1120;padding:1px 5px;
border-radius:4px;font-size:.85em}
.ok{color:var(--ok)}.warn{color:#e0b34f}
footer{margin-top:60px;padding-top:18px;border-top:1px dashed var(--line);
color:var(--dim);font-size:.8rem}
@media(max-width:640px){.step{grid-template-columns:28px 1fr;grid-auto-flow:dense}
.step .d{grid-column:2}
h1{font-size:1.45rem}
table{display:block;overflow-x:auto;white-space:nowrap;-webkit-overflow-scrolling:touch}
th,td{white-space:normal;min-width:110px}}

.matrix td.best{color:var(--ok);font-weight:700}
.matrix td.dead{color:#c0392b}
</style></head><body><div class="wrap">

<header>
<div class="kicker">React force recovery Β· results</div>
<h1>Three estimators Γ— three ground-truth datasets</h1>
<p class="sub">Every method evaluated on every dataset with force-sensor labels,
predicted vs ground truth per dataset β€” so dataset quality is controlled within
each row, and differences between panels are differences between methods.</p>
<a class="pill" href="method.html">β†– how the method is designed</a>
<a class="pill" href="index.html">overview</a>
<a class="pill" href="gallery.html">gallery</a>
<a class="pill" href="results_zh.html">δΈ­ζ–‡</a>
</header>

<h2>The matrix (Spearman ρ, predicted vs F/T ground truth)</h2>
<div class="card"><table class="matrix">
<tr><th>dataset (gel type)</th><th>Ours β€” physics,<br>0 training frames</th>
<th>FEATS U-net<br>(trained: marker gel)</th>
<th>FeelAnyForce<br>(trained: markerless)</th></tr>
<tr><td>FEATS val (marker)</td><td>0.77</td>
<td class="best">0.96 Β· in-domain</td><td class="dead">0.43</td></tr>
<tr><td>FoTa cnc_Mini (markerless)</td><td>0.94 (in view)</td>
<td class="dead">0.07</td><td class="best">0.83</td></tr>
<tr><td>GlowTact (markerless)</td><td>0.98</td>
<td class="dead">0.04</td><td class="best">0.90</td></tr>
</table>
<p>The pattern is the finding: <b>each network dominates its own gel domain and
collapses outside it</b> (FEATS 0.96 β†’ 0.04–0.07; FeelAnyForce 0.90 β†’ 0.43),
while <b>the physics pipeline is the only estimator that works everywhere</b>
(0.74–0.99) β€” it never sees training data, so it has no domain to leave.</p>
</div>

<h2>FEATS dataset β€” marker-dot gel</h2>
<div class="card"><img src="assets/results_feats.png" alt="FEATS dataset panels">
<p class="footnote">In-domain, the FEATS U-net is excellent (ρ=0.96) β€” the
negative results elsewhere are domain effects, not a weak model. FeelAnyForce,
markerless-trained, degrades on the dotted gel (0.43): the same knife cuts both
ways.</p></div>

<h2>FoTa cnc_Mini β€” markerless gel</h2>
<div class="card"><img src="assets/results_cnc.png" alt="cnc_Mini panels">
<p class="footnote">Hard conditions: only 4 contact-free frames, 62% of presses
near the pad border β€” the press grid is larger than the field of view (see
<a href="debug_pipeline.html">pipeline debug</a>). Strictly in view our ρ is
0.94 (MAE 0.26 N); FeelAnyForce reaches 0.92 non-edge.</p></div>

<h2>GlowTact β€” markerless gel, cleaned</h2>
<div class="card"><img src="assets/results_glowtact.png" alt="GlowTact panels">
<p class="footnote">Friendliest ground truth (centred presses, 10 free frames,
0–20 N). Per-indenter, calibrated within each family under the physical scope
(contact fully in view, gel not bottomed out), ρ is 0.975–0.992 across all six
indenters with MAE ≀ 0.73 N. Caveat: that calibration is refit per family, so
it measures rank agreement within a group, not a transferable absolute-newton
scale β€” see <a href="method.html">per-dataset calibration</a>.</p></div>

<h2>Sparsh (Meta) β€” a fourth dataset, and where the method breaks</h2>
<div class="card"><img src="assets/results_sparsh.png" alt="Sparsh results">
<p class="footnote">10 gel pads (6 sphere, 2 flat, 2 sharp), force in newtons.
Our GlowTact table applied to this foreign sensor reaches &rho;=0.878 on
in-view frames. Rebuilding the table from <b>Sparsh's own sphere presses</b>
β€” 708 frames, radius fitted at R=2.44 mm from a&sup2;=d(2R&minus;d) β€” takes it
to <b>&rho;=0.968, MAE 0.042 N</b>, against a labels-shuffled-within-pad
control of 0.23. Fitting on one pad and applying it <b>unchanged</b> to another
costs nothing measurable (0.96&ndash;0.98 everywhere): one table, six gel pads.</p>
<img src="assets/sparsh_dome.png" alt="dome before and after">
<p class="footnote">Why the table matters more than the fit: with the wrong
sensor's table a sphere press integrates to a <b>bilobed shape with a central
dip</b>; with the self-calibrated table it is a single dome matching the
analytic spherical cap (residual RMS 0.179 &rarr; <b>0.0545 mm</b>). Measured
before any integration, the LUT gradient sits <b>93.3&deg;</b> from the analytic
sphere gradient β€” chance is 90&deg; β€” and self-calibration brings it to
<b>4.5&deg;</b> (within 30&deg;: 15% &rarr; 99%). Reverse control: the Sparsh
table fails on GlowTact frames too, so this is a per-sensor property, not a
bad table.</p>
<p class="footnote"><b>What this costs and where it still fails.</b> The price
is one set of sphere presses with logged depth on the target sensor: this is
<b>calibrate once per sensor</b>, not zero-shot. Rank order transfers across
indenter shapes but <i>absolute newtons do not</i> β€” a sphere-fitted model
applied to a flat punch degrades to MAE 0.37&ndash;0.40 N, and it got
<i>worse</i> with the correct table, because true geometry widens the real
feature-scale gap between a sphere and a punch. The sharp indenter is
unsupported (in-view &rho; 0.58). Shear stays out of reach by construction β€”
the top shear decile keeps 1.6&times; the residual whichever table is used.
Frames are restricted to a visible contact disc: 36% of presses show none and
11% are clipped, and the clipped subset carries the <i>highest</i> median force
yet scores worse, so this is visibility, not force-range filtering. Three
dataset defects had to be fixed first: flat/sharp trajectories ship 5 more
frame indices than forces (silently drifting the labels, &rho;&asymp;0 until
paired within each trajectory), sharp/batch_2 is stored BGR while the other
nine are RGB, and flat/batch_2 ships only 3 of 4 image files.</p></div>

<h2>React β€” no ground truth, so: do independent methods agree?</h2>
<div class="card"><img src="assets/results_react.png" alt="React agreement">
<p class="footnote">On the dataset we actually care about, the two surviving
estimators β€” physics (zero training) and FeelAnyForce (200K frames) β€” agree at
ρ=0.91, and FeelAnyForce reads β‰ˆ0 N on every frame the physics pipeline calls
contact-free. Neither can copy the other's mistakes.</p></div>

<h2>Takeaway</h2>
<div class="card"><p style="margin-top:0"><b>For labelling React</b> (markerless
Mini): FeelAnyForce as the primary labeller, the physics pipeline as an
independent audit, disagreement rows flagged. <b>For any new gel or sensor</b>
where no trained model matches the domain: the physics pipeline is the only
option that works out of the box β€” and its FEATS-dataset score (0.74) shows
what it does on a domain nobody tuned it for.</p></div>

<p><a href="debug_pipeline.html"><b>Pipeline debug page</b></a>: raw
image &rarr; force step by step on all three datasets, and the cnc
field-of-view ablation (in-view &rho;=0.94).</p>

<footer>React force recovery Β· <a href="index.html">overview</a> Β·
<a href="method.html">method design</a> Β· <a href="gallery.html">gallery</a> Β·
data: FEATS (2411.03315) Β· FoTa/T3 (2406.13640) Β· GlowTact
(dacongming666/GlowTact_Datasets) Β· FeelAnyForce (2410.02048)</footer>
</div></body></html>