WolfDavid commited on
Commit
7a26764
·
1 Parent(s): daee26b

fix(01-05): pose the arms out of the authored T-pose and assert it

Browse files

A VRM's rest pose is a T-pose and nothing in the stage ever moved the arms,
so the deployed avatar stood with both arms straight out while every AVTR-01
check passed (docs/evidence/2026-09-05-deployed-tpose.png).

vrm-stage.js now rotates each upper arm 1.2 rad about Z toward the body and
the forearm a further 0.2 rad, once, on the normalized bones at mount. The
idle life only writes spine, hips and head, so the two never compete.

The pose is measured every frame on the RAW skeleton after render and
published as debug.armDown.{left,right}: the downward component of the
world-space shoulder->elbow direction (-1 down, 0 T-pose). Measured -0.932
both sides, which is sin(1.2) exactly.

Tests at all three layers assert armDown < -0.7 (ARM_DOWN_MAX): the
standalone stage, both Gradio transports through the facade (so the field
is proven to cross the iframe boundary), and the deployed Space. Each waits
for the first rendered frame via breathValue rather than a fixed sleep,
because the first frame compiles every MToon shader and was measured over
1.5 s on the headless fleet. Mutation-tested: with ARM_REST_DROP = 0 the
standalone test fails with armDown=0.000.

.planning/phases/01-voice-avatar-loop-skeleton/01-VALIDATION.md CHANGED
@@ -93,9 +93,12 @@ the planner MUST bind each row to a task and the executor MUST fill in Status.
93
  ## Observable Signals
94
 
95
  The canvas is opaque — **never screenshot-diff the VRM.** The avatar must expose
96
- `window.Avatar.__debug = { ready, mountCount, vrmMetaTitle, threeInstanceCount, currentVisemes: {aa,ih,ou,ee,oh}, blinkValue, clockOffset }`
97
  so every visual assertion becomes a numeric one.
98
 
 
 
 
99
  - **Audio:** `speech-start` / `speech-end` custom events, `AudioContext.state`, `AudioBuffer.duration`. Duration ratios are the reliable proxy for "it actually spoke slower".
100
  - **Module identity:** `performance.getEntriesByType('resource').filter(e => e.name.includes('three.mjs')).length === 1`
101
  - **Latency:** `performance.mark`/`measure` names emitted per stage, scraped by the harness.
 
93
  ## Observable Signals
94
 
95
  The canvas is opaque — **never screenshot-diff the VRM.** The avatar must expose
96
+ `window.Avatar.__debug = { ready, mountCount, vrmMetaTitle, threeInstanceCount, currentVisemes: {aa,ih,ou,ee,oh}, visemePeaks, blinkValue, blinkCount, breathValue, armDown: {left,right}, clockOffset }`
97
  so every visual assertion becomes a numeric one.
98
 
99
+ - **Rest pose:** `armDown.{left,right}` is the downward component of the world-space shoulder→elbow direction on the raw skeleton (-1 down, 0 T-pose). Assert `< -0.7` (`ARM_DOWN_MAX`). Added 2026-09-05 after the deployed avatar passed every AVTR-01 check while standing in a T-pose.
100
+ - **Reading `__debug`:** it is a snapshot refreshed only by `getDebug()`. Always `await window.Avatar.getDebug()` per sample; polling `__debug` directly reads frozen values.
101
+
102
  - **Audio:** `speech-start` / `speech-end` custom events, `AudioContext.state`, `AudioBuffer.duration`. Duration ratios are the reliable proxy for "it actually spoke slower".
103
  - **Module identity:** `performance.getEntriesByType('resource').filter(e => e.name.includes('three.mjs')).length === 1`
104
  - **Latency:** `performance.mark`/`measure` names emitted per stage, scraped by the harness.
avatar/vrm-stage.js CHANGED
@@ -27,6 +27,16 @@ const SWAY_HIPS = 0.02; // rad
27
  const THINK_TILT = 0.08; // rad
28
  const LISTEN_LEAN = 0.05; // rad
29
 
 
 
 
 
 
 
 
 
 
 
30
  /**
31
  * Copy only structured-cloneable scalars out of a VRM meta block.
32
  * `vrm.meta.thumbnailImage` is an HTMLImageElement, which cannot cross a frame
@@ -135,6 +145,11 @@ export async function mountStage(canvasEl, vrmUrl, emit = () => {}) {
135
  blinkValue: 0,
136
  blinkCount: 0,
137
  breathValue: 0,
 
 
 
 
 
138
  clockOffset: 0,
139
  thinking: false,
140
  listening: false,
@@ -152,6 +167,42 @@ export async function mountStage(canvasEl, vrmUrl, emit = () => {}) {
152
  cameraY: camera.position.y,
153
  };
154
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
155
  const clock = new THREE.Clock();
156
  let elapsed = 0;
157
  let nextBlinkAt = BLINK_MIN + Math.random() * BLINK_SPREAD;
@@ -216,6 +267,7 @@ export async function mountStage(canvasEl, vrmUrl, emit = () => {}) {
216
  if (onTick) onTick(dt);
217
  vrm.update(dt); // MUST run after expression values are set, every frame
218
  renderer.render(scene, camera);
 
219
  });
220
 
221
  const ro = new ResizeObserver(() => {
@@ -231,7 +283,11 @@ export async function mountStage(canvasEl, vrmUrl, emit = () => {}) {
231
  scene,
232
  camera,
233
  renderer,
234
- getDebug: () => ({ ...debug, currentVisemes: { ...debug.currentVisemes } }),
 
 
 
 
235
  setExpressionWeights(weights) {
236
  const em = vrm.expressionManager;
237
  for (const v of VISEME_NAMES) {
 
27
  const THINK_TILT = 0.08; // rad
28
  const LISTEN_LEAN = 0.05; // rad
29
 
30
+ // Rest pose. A VRM's authored rest pose IS a T-pose - both arms straight out along X -
31
+ // and nothing about the idle life above moves the arms, so without this block the
32
+ // avatar stands like a scarecrow on a page whose every other test is green. The upper
33
+ // arm is rotated about Z toward the body (1.2 rad leaves it ~21 deg out from vertical,
34
+ // a relaxed A-pose rather than pinned to the hip) and the forearm follows a little
35
+ // further in. Applied ONCE to the normalized bones at mount; the idle life only ever
36
+ // writes spine, hips and head, so the two never compete for a bone.
37
+ const ARM_REST_DROP = 1.2; // rad about Z, sign per side
38
+ const FOREARM_REST_DROP = 0.2; // rad about Z, sign per side
39
+
40
  /**
41
  * Copy only structured-cloneable scalars out of a VRM meta block.
42
  * `vrm.meta.thumbnailImage` is an HTMLImageElement, which cannot cross a frame
 
145
  blinkValue: 0,
146
  blinkCount: 0,
147
  breathValue: 0,
148
+ // Downward component of each upper arm's world-space direction, read from the
149
+ // RAW skeleton the renderer skins (not the normalized rig the pose is written to):
150
+ // -1 is straight down, 0 is the T-pose, +1 is straight up. Tests assert on this so
151
+ // a regression to the T-pose fails a number instead of needing an eyeball.
152
+ armDown: { left: 0, right: 0 },
153
  clockOffset: 0,
154
  thinking: false,
155
  listening: false,
 
167
  cameraY: camera.position.y,
168
  };
169
 
170
+ // Bring the arms down out of the authored T-pose. Normalized bone axes are world-
171
+ // aligned at rest under both VRM 0.0 and 1.0, so the same signed Z rotation lowers
172
+ // the left arm (+X) and the right arm (-X) symmetrically.
173
+ for (const [side, sign] of [
174
+ ['left', -1],
175
+ ['right', 1],
176
+ ]) {
177
+ const upper = bone(`${side}UpperArm`);
178
+ const lower = bone(`${side}LowerArm`);
179
+ if (upper) upper.rotation.z = sign * ARM_REST_DROP;
180
+ if (lower) lower.rotation.z = sign * FOREARM_REST_DROP;
181
+ }
182
+
183
+ // Measure the pose the way a viewer sees it: the direction from shoulder to elbow in
184
+ // world space, on the raw bones that drive the mesh. Read after render, when the
185
+ // world matrices are current.
186
+ const rawBone = (name) =>
187
+ vrm.humanoid?.getRawBoneNode?.(name) ?? vrm.humanoid?.getNormalizedBoneNode?.(name) ?? null;
188
+ const armPairs = {
189
+ left: [rawBone('leftUpperArm'), rawBone('leftLowerArm')],
190
+ right: [rawBone('rightUpperArm'), rawBone('rightLowerArm')],
191
+ };
192
+ const shoulderPos = new THREE.Vector3();
193
+ const elbowPos = new THREE.Vector3();
194
+ function measureArms() {
195
+ for (const side of ['left', 'right']) {
196
+ const [upper, lower] = armPairs[side];
197
+ if (!upper || !lower) continue;
198
+ upper.getWorldPosition(shoulderPos);
199
+ lower.getWorldPosition(elbowPos);
200
+ const dir = elbowPos.sub(shoulderPos);
201
+ const len = dir.length();
202
+ debug.armDown[side] = len > 0 ? dir.y / len : 0;
203
+ }
204
+ }
205
+
206
  const clock = new THREE.Clock();
207
  let elapsed = 0;
208
  let nextBlinkAt = BLINK_MIN + Math.random() * BLINK_SPREAD;
 
267
  if (onTick) onTick(dt);
268
  vrm.update(dt); // MUST run after expression values are set, every frame
269
  renderer.render(scene, camera);
270
+ measureArms();
271
  });
272
 
273
  const ro = new ResizeObserver(() => {
 
283
  scene,
284
  camera,
285
  renderer,
286
+ getDebug: () => ({
287
+ ...debug,
288
+ currentVisemes: { ...debug.currentVisemes },
289
+ armDown: { ...debug.armDown },
290
+ }),
291
  setExpressionWeights(weights) {
292
  const em = vrm.expressionManager;
293
  for (const v of VISEME_NAMES) {
docs/evidence/2026-09-05-local-rest-pose.png ADDED
tests/e2e/test_avatar_loop.py CHANGED
@@ -20,6 +20,8 @@ import time
20
  import pytest
21
  import requests
22
 
 
 
23
  # Applied at module scope AND per test. The module-level mark is the one that matters -
24
  # it cannot be forgotten on a test plan 01-09 adds later - while the per-test decorators
25
  # are what this plan's acceptance check counts. Re-applying the same mark is a no-op.
@@ -85,6 +87,16 @@ async (ms) => {
85
  }
86
  """
87
 
 
 
 
 
 
 
 
 
 
 
88
  # Resource timing for the VRM itself. Recorded rather than asserted: plan 01-10 reuses
89
  # it in docs/LATENCY.md, and a slow CDN is not a reason to fail the spike.
90
  VRM_TIMING = """
@@ -237,6 +249,30 @@ def test_idle_life(page, space_url, warm_space):
237
  )
238
 
239
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
240
  @pytest.mark.deployed
241
  def test_no_remount(page, space_url, warm_space):
242
  """AVTR-01, and the direct answer to RESEARCH's Open Question 3.
 
20
  import pytest
21
  import requests
22
 
23
+ from tests.e2e.test_stage_standalone import ARM_DOWN_MAX, FIRST_FRAME_TIMEOUT_MS
24
+
25
  # Applied at module scope AND per test. The module-level mark is the one that matters -
26
  # it cannot be forgotten on a test plan 01-09 adds later - while the per-test decorators
27
  # are what this plan's acceptance check counts. Re-applying the same mark is a no-op.
 
87
  }
88
  """
89
 
90
+ # ready fires before the first frame renders and the first frame compiles every shader;
91
+ # the pose is measured post-render, so wait for a tick to have happened. breathValue is
92
+ # written every tick and is exactly 0 only before the first one.
93
+ FIRST_FRAME = """
94
+ async () => {
95
+ const d = window.Avatar ? await window.Avatar.getDebug() : null;
96
+ return !!d && d.breathValue !== 0;
97
+ }
98
+ """
99
+
100
  # Resource timing for the VRM itself. Recorded rather than asserted: plan 01-10 reuses
101
  # it in docs/LATENCY.md, and a slow CDN is not a reason to fail the spike.
102
  VRM_TIMING = """
 
249
  )
250
 
251
 
252
+ @pytest.mark.deployed
253
+ def test_arms_at_sides(page, space_url, warm_space):
254
+ """AVTR-01. The deployed avatar stands with its arms down, as a number.
255
+
256
+ Revision 41ee90e passed test_vrm_ready and test_idle_life while standing in a full
257
+ T-pose (docs/evidence/2026-09-05-deployed-tpose.png): threeInstanceCount was 1, the
258
+ VRM had its meta, it blinked and breathed - and nothing had ever posed the arms. This
259
+ is the assertion that would have caught it. armDown is the downward component of the
260
+ world-space shoulder->elbow direction on the raw skeleton: 0 is the T-pose.
261
+ """
262
+ _open_ready(page, space_url)
263
+ page.wait_for_function(FIRST_FRAME, timeout=FIRST_FRAME_TIMEOUT_MS)
264
+ debug = read_debug(page)
265
+ arms = debug.get("armDown") if debug else None
266
+ print(f"[deployed] armDown={arms}")
267
+
268
+ assert arms is not None, "armDown is missing from the deployed debug surface"
269
+ for side in ("left", "right"):
270
+ assert arms[side] < ARM_DOWN_MAX, (
271
+ f"{side} arm reads armDown={arms[side]:.3f} on the deployed page; the avatar is "
272
+ "standing in a T-pose again (0 is the T-pose, -1 straight down)"
273
+ )
274
+
275
+
276
  @pytest.mark.deployed
277
  def test_no_remount(page, space_url, warm_space):
278
  """AVTR-01, and the direct answer to RESEARCH's Open Question 3.
tests/e2e/test_facade_parity.py CHANGED
@@ -13,6 +13,7 @@ from __future__ import annotations
13
 
14
  import pytest
15
 
 
16
  from tests.test_transport_seam import avatar_surface
17
 
18
  pytestmark = pytest.mark.slow
@@ -21,6 +22,16 @@ TRANSPORTS = ("inline", "iframe")
21
  AVATAR_READY = "() => !!window.Avatar && !!window.Avatar.__debug && window.Avatar.__debug.ready"
22
  BOOT_TIMEOUT_MS = 120_000
23
 
 
 
 
 
 
 
 
 
 
 
24
  # Deliberately double-quoted and built from data: the surface list must come from
25
  # facade.js, never from a literal in this file.
26
  #
@@ -85,6 +96,10 @@ def live_avatars(browser, gradio_apps):
85
  "deferred": {name: page.evaluate(CALL_DEFERRED, name) for name in DEFERRED_METHODS},
86
  "wired": {name: page.evaluate(PROBE_WIRED, name) for name in WIRED_METHODS},
87
  }
 
 
 
 
88
  finally:
89
  page.close()
90
  return captured
@@ -116,6 +131,23 @@ def test_transports_expose_identical_debug_keys(live_avatars):
116
  )
117
 
118
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
119
  def test_deferred_methods_fail_loudly_not_silently(live_avatars):
120
  """The remaining stubs are placeholders, not accidental no-ops.
121
 
 
13
 
14
  import pytest
15
 
16
+ from tests.e2e.test_stage_standalone import ARM_DOWN_MAX, FIRST_FRAME_TIMEOUT_MS
17
  from tests.test_transport_seam import avatar_surface
18
 
19
  pytestmark = pytest.mark.slow
 
22
  AVATAR_READY = "() => !!window.Avatar && !!window.Avatar.__debug && window.Avatar.__debug.ready"
23
  BOOT_TIMEOUT_MS = 120_000
24
 
25
+ # The pose is measured after a render, and ready fires before the first one. Same
26
+ # frame-rendered signal as the standalone suite, read through the facade because the
27
+ # iframe transport's stage is in another document.
28
+ FIRST_FRAME = """
29
+ async () => {
30
+ const d = window.Avatar ? await window.Avatar.getDebug() : null;
31
+ return !!d && d.breathValue !== 0;
32
+ }
33
+ """
34
+
35
  # Deliberately double-quoted and built from data: the surface list must come from
36
  # facade.js, never from a literal in this file.
37
  #
 
96
  "deferred": {name: page.evaluate(CALL_DEFERRED, name) for name in DEFERRED_METHODS},
97
  "wired": {name: page.evaluate(PROBE_WIRED, name) for name in WIRED_METHODS},
98
  }
99
+ page.wait_for_function(FIRST_FRAME, timeout=FIRST_FRAME_TIMEOUT_MS)
100
+ captured[transport]["arm_down"] = page.evaluate(
101
+ "async () => (await window.Avatar.getDebug()).armDown"
102
+ )
103
  finally:
104
  page.close()
105
  return captured
 
131
  )
132
 
133
 
134
+ def test_arms_rest_at_sides_under_both_transports(live_avatars):
135
+ """The rest pose is measured on the stage, but a learner sees it through a transport.
136
+
137
+ Under the iframe transport the number has to survive a postMessage round trip, which
138
+ is exactly the kind of field that gets added to the inline path and forgotten on the
139
+ fallback. Same threshold as the standalone and deployed suites.
140
+ """
141
+ for transport in TRANSPORTS:
142
+ arms = live_avatars[transport]["arm_down"]
143
+ assert arms is not None, f"{transport}: armDown never reached window.Avatar.getDebug()"
144
+ for side in ("left", "right"):
145
+ assert arms[side] < ARM_DOWN_MAX, (
146
+ f"{transport}: {side} arm reads armDown={arms[side]:.3f}; "
147
+ "0 is the T-pose, -1 straight down"
148
+ )
149
+
150
+
151
  def test_deferred_methods_fail_loudly_not_silently(live_avatars):
152
  """The remaining stubs are placeholders, not accidental no-ops.
153
 
tests/e2e/test_stage_standalone.py CHANGED
@@ -17,6 +17,20 @@ pytestmark = pytest.mark.slow
17
  READY_TIMEOUT_MS = 30_000
18
  STAGE_READY = "() => !!window.__stageDebug && window.__stageDebug.ready === true"
19
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
  # Sampling happens INSIDE the page at animation-frame rate rather than by polling from
21
  # Python. A blink is a 120 ms ramp on a 1.8-5.8 s schedule, so its peak occupies well
22
  # under 2% of the wall clock: a 250 ms poll would miss it on most runs and the test
@@ -115,6 +129,35 @@ def test_idle_life_values_change(page, static_server):
115
  )
116
 
117
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
118
  def test_demo_timeline_drives_visemes(page, static_server):
119
  """A generated timeline against the real ずんだもん WAV, clocked off the AudioContext.
120
 
 
17
  READY_TIMEOUT_MS = 30_000
18
  STAGE_READY = "() => !!window.__stageDebug && window.__stageDebug.ready === true"
19
 
20
+ # armDown is the downward component of the shoulder->elbow direction in world space:
21
+ # -1 straight down, 0 the authored T-pose, +1 straight up. The rest pose puts the upper
22
+ # arm 1.2 rad from horizontal (sin 1.2 = 0.93 down), so -0.7 - an arm within 45 deg of
23
+ # vertical - separates "posed" from "T-pose" by a wide margin while leaving room for the
24
+ # idle sway. Shared verbatim with the deployed suite.
25
+ ARM_DOWN_MAX = -0.7
26
+
27
+ # `ready` is announced before the first frame renders, and the first frame compiles every
28
+ # MToon shader - measured at over 1.5 s on the headless fleet - so the debug snapshot can
29
+ # still hold its boot-time zeros well after ready. breathValue is written on every tick and
30
+ # is only exactly 0 before the first one, which makes it the frame-rendered signal.
31
+ FIRST_FRAME = "() => !!window.__stageDebug && window.__stageDebug.breathValue !== 0"
32
+ FIRST_FRAME_TIMEOUT_MS = 30_000
33
+
34
  # Sampling happens INSIDE the page at animation-frame rate rather than by polling from
35
  # Python. A blink is a 120 ms ramp on a 1.8-5.8 s schedule, so its peak occupies well
36
  # under 2% of the wall clock: a 250 ms poll would miss it on most runs and the test
 
129
  )
130
 
131
 
132
+ def test_arms_rest_at_sides(page, static_server):
133
+ """AVTR-01's other half: the avatar stands like a person, not a T-posed statue.
134
+
135
+ Every previous assertion in this file passed while the deployed avatar held both arms
136
+ straight out - a VRM's authored rest pose IS a T-pose and nothing ever moved the arms
137
+ (docs/evidence/2026-09-05-deployed-tpose.png). The number is read from the RAW bones
138
+ the mesh is skinned from, after a render, so it measures what the viewer sees.
139
+ """
140
+ _open_stage(page, static_server)
141
+ page.wait_for_function(FIRST_FRAME, timeout=FIRST_FRAME_TIMEOUT_MS)
142
+ arms = page.evaluate("() => window.__stageDebug.armDown")
143
+ print(f"[standalone] armDown={arms}")
144
+
145
+ assert arms is not None, "armDown is missing from __stageDebug; the stage never measured"
146
+ for side in ("left", "right"):
147
+ assert arms[side] < ARM_DOWN_MAX, (
148
+ f"{side} arm reads armDown={arms[side]:.3f} (0 is the T-pose, -1 straight down); "
149
+ "the rest pose in vrm-stage.js is not reaching the rendered skeleton"
150
+ )
151
+
152
+ # The idle life must not undo it: two seconds of breathing and sway later, still down.
153
+ page.wait_for_timeout(2_000)
154
+ later = page.evaluate("() => window.__stageDebug.armDown")
155
+ for side in ("left", "right"):
156
+ assert later[side] < ARM_DOWN_MAX, (
157
+ f"{side} arm drifted to armDown={later[side]:.3f} after 2 s of idle life"
158
+ )
159
+
160
+
161
  def test_demo_timeline_drives_visemes(page, static_server):
162
  """A generated timeline against the real ずんだもん WAV, clocked off the AudioContext.
163