File size: 11,520 Bytes
da46df0
 
199a6c3
487b8cf
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82717b7
 
 
 
 
199a6c3
 
 
 
 
 
 
 
 
487b8cf
 
 
 
 
 
cf68e57
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
487b8cf
 
cf68e57
487b8cf
82717b7
da46df0
 
 
 
 
 
 
 
199a6c3
da46df0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
199a6c3
 
da46df0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
# Resume here — Japanese Learning Avatar

> ## ✅ 2026-09-05 (latest) — 01-05 CLOSED: spike CONFIRMED, `AVATAR_TRANSPORT=inline`, Space at `487b8cf`
>
> Commit `7a26764` poses the arms out of the authored T-pose in `avatar/vrm-stage.js`
> (upper arm 1.2 rad about Z toward the body, forearm a further 0.2 rad, applied once on the
> normalized bones at mount) and publishes `armDown.{left,right}` — the downward component
> of the world-space shoulder→elbow direction on the RAW skeleton, read after every render.
> Measured **-0.932 both sides** (= sin 1.2). Local screenshot: `docs/evidence/2026-09-05-local-rest-pose.png`
> (before: `docs/evidence/2026-09-05-deployed-tpose.png`).
>
> **Tests now assert `armDown < -0.7` at all three layers:** `test_stage_standalone.py::test_arms_rest_at_sides`,
> `test_facade_parity.py::test_arms_rest_at_sides_under_both_transports` (proves the field
> crosses the iframe postMessage boundary), and the deployed
> `test_avatar_loop.py::test_arms_at_sides`. Mutation-tested: `ARM_REST_DROP = 0` fails the
> standalone test with `armDown=0.000`. Each waits for the first rendered frame via
> `breathValue !== 0` — `ready` fires *before* the first frame and that frame compiles every
> MToon shader (>1.5 s measured headless), so a fixed sleep read boot-time zeros.
>
> **Verified:** `ruff` clean; 82 quick-loop; full `pytest tests/e2e/` **17 passed** as a whole
> (5 deployed skipped without `--space-url`).
>
> **Watched, not fixed:** two later full-suite runs each dropped one to three
> `test_asr_standalone.py` tests (`speech_ja.wav` read as rms 0.00057 → `rms-floor`;
> `test_tier_wasm_fallback`). The ASR harness never imports `vrm-stage.js`, and the file passes
> 7/7 alone. Signature matches Chromium's `--use-file-for-fake-audio-capture=…%noloop`
> finishing the WAV before the page's `getUserMedia` attaches when the fleet boots slowly.
> Pre-existing; not caused by this change.
>
> **Deployed and verified.** Pushed `41ee90e..487b8cf`; Space RUNNING in 106 s. Deployed suite
> `pytest tests/e2e/test_avatar_loop.py --space-url https://wolfdavid-japanese-learning-avatar.hf.space`
> **5 passed, twice**, incl. the new `test_arms_at_sides` (live `armDown` -0.932 both sides).
> Live screenshot: `docs/evidence/2026-09-05-deployed-rest-pose.png` — relaxed A-pose.
>
> **Verdict written:** `docs/SPIKE-VERDICT.md` — CONFIRMED, evidence table complete, including
> the huggingface.co Space-page embed read from inside the cross-origin iframe
> (`docs/evidence/2026-09-05-hub-embed.png`). `01-05-SUMMARY.md` closes the plan; **DPLY-01 is
> complete**. `HANDOFF.json` is gone — nothing is blocked any more.
>
> **Next:** `/gsd:execute-phase 1` resumes at **01-08**, the plan that makes the avatar speak.
> One human row is still open and does not block 01-08: **AVTR-01's mobile check** — load the
> Space on a real phone, note device / OS / browser / renders / smooth-or-slideshow, and put it in
> the verdict's *Mobile* line.
>
> **"It does not speak" is still expected.** `dispatchTurn()` throws until plan 01-08.
> **`window.Avatar.__debug` is still a stale snapshot** — always `await window.Avatar.getDebug()`.

> <details><summary>Earlier banner (2026-09-05, before the fix) — kept for the diagnosis</summary>
>
> ## ⚠ 2026-09-05 — READ THIS FIRST: the deployed avatar is in a T-pose
>
> The Space is **live and working** at
> <https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar> (revision `41ee90e`,
> `RUNNING` on `zero-a10g`) — but the avatar stands with **both arms straight out**.
> Screenshot evidence: `docs/evidence/2026-09-05-deployed-tpose.png`.
>
> **This is NOT the two-three.js-instances bug.** The screenshot shows correct materials,
> hair, face and toon shading, and the instruments report `threeInstanceCount = 1`,
> `vrmMetaTitle = 'VRM1_Constraint_Twist_Sample'`, `vrmSpecVersion = '1'`. A plain-glTF
> fallback would have no `vrm.meta` at all.
>
> **Root cause: nothing ever applies an idle POSE.** A VRM's authored rest pose *is* a
> T-pose. Phase 1's idle life only adds tiny offsets — spine 0.012 rad, hips 0.02 rad,
> head 0.05–0.08 rad, all under 1.5° — and **never touches the arms**. No code in this
> repo has ever posed them.
>
> **The fix** (detail in `.planning/HANDOFF.json` → `the_fix`): hand-pose
> `leftUpperArm`/`rightUpperArm` about Z by ~±1.2 rad at mount in `avatar/vrm-stage.js`,
> before the render loop starts. Zero new dependencies, ~10 lines, composes with the
> existing additive idle. The alternative is a `.vrma` clip via
> `@pixiv/three-vrm-animation@3.5.5`, which is better long-term but adds a dependency and
> an asset to license — better as a later polish plan.
>
> **Then add a deployed assertion that the arms are DOWN.** AVTR-01 passed every existing
> check while looking visibly wrong. That is the real lesson: do not mark AVTR-01 complete.
>
> **"It does not speak" is expected, not a bug.** `dispatchTurn()` throws
> `not wired yet - plan 01-08 implements it`. Plan **01-08** is the one that makes it talk.
>
> **Diagnostic trap — cost real time this session:** `window.Avatar.__debug` is a *stale
> snapshot*, refreshed only inside `getDebug()`. Polling it in a rAF loop returns frozen
> values forever and looks exactly like a dead render loop. Always
> `await window.Avatar.getDebug()` per sample; copy `IDLE_SAMPLER` from
> `tests/e2e/test_avatar_loop.py`.
>
> </details>

**Updated 2026-09-05 (T-pose fix). Originally written 2026-08-27.**
Working tree is clean. Everything is committed. The fix IS deployed: the Space runs revision `487b8cf`. Local master carries only docs commits on top.

To pick back up, run `/gsd:resume-work` in this directory. It reads `.planning/HANDOFF.json`
first, which carries the same state in machine-readable form.

---

## One-line status

Phase 01 is **7 of 10 plans complete and unblocked**: the Space is deployed, the spike is CONFIRMED, and 01-08 is next.

---

## Why nothing was pushed

This repo's **only** git remote is `space` — the public Hugging Face Space at
`WolfDavid/japanese-learning-avatar`. There is no GitHub remote. Branch `master` has no upstream.

So `git push` here does not mean "back up my work." It means **publish the app publicly and
trigger a Space rebuild** — which is precisely plan 01-05 Task 1, the gate that is waiting on you.

Pushing right now would also produce a **broken** public build. There is no `README.md` in this
repo; the Space's front-matter lives only on the HF side and was auto-generated wrong when the
Space was created:

| Front-matter key | Currently on HF | Must become | Why |
|---|---|---|---|
| `sdk_version` | `6.26.0` | `6.22.0` | must match `gradio==6.22.0` in `requirements.txt` |
| `python_version` | `'3.12'` | `3.12.12` | ZeroGPU provides only 3.10.13 and 3.12.12 |

Writing that corrected README is part of plan 01-05. Push after 01-05 runs, not before.

**42 commits are unpushed by design, not by omission.**

If what you actually want is an off-machine backup, that is a *different* action from deploying:
add a private GitHub remote and push there. That was not done unprompted because the standing
rule for this account is to ask before pushing nested repos.

---

## Where the work stands

| Plan | What it is | State |
|---|---|---|
| 01-01 | Hosting decision + Space creation + VRM sourcing | Complete |
| 01-02 | Toolchain, Git LFS, package skeleton, test scaffolding | Complete |
| 01-03 | Avatar stage, shared facade + turn loop, both transports | Complete (resumed from the crash) |
| 01-04 | VOICEVOX TTS, `requirements.txt`, ground-truth fixtures | Complete |
| 01-06 | Mora-to-viseme timeline builder | Complete |
| 01-07 | Push-to-talk mic gate + tiered browser ASR | Complete |
| 01-05 | Space manifest, deploy the spike, verdict | Complete — CONFIRMED, inline |
| **01-08** | **Turn-loop wiring — closes the round trip** | **NEXT** (`/gsd:execute-phase 1`) |
| 01-09 | Full deployed E2E suite + `LICENSES.md` | Blocked behind 01-08 |
| 01-10 | Vendor modules, latency harness (user-gated) | Blocked behind 01-09 |

The dependency chain is real, not a preference: `01-08 depends_on [01-05, 01-06, 01-07]`,
`01-09 depends_on [01-08]`, `01-10 depends_on [01-09]`. Nothing proceeds until 01-05 does.

### Verified green at handoff

- 82 quick-loop tests in ~12.8 s
- 15 end-to-end browser tests
- `ruff check .` and `ruff format --check .` both clean

**Run `pytest tests/e2e/` as a whole before trusting a green phase.** Per-file runs hid two real
defects this session — each plan's own verification passed while the full suite failed.

---

## What deploying would actually test

The risky part is already retired. three.js + `@pixiv/three-vrm` render inside a real
Gradio 6.22.0 `gr.HTML` component locally, with `threeInstanceCount === 1` and
`mountCount === 1`, and the iframe fallback boots identically under `AVATAR_TRANSPORT=iframe`.

What remains genuinely deployment-specific: CDN reach from `*.hf.space`, `set_static_paths`
behind the Space proxy, cold-start behaviour, and mobile.

Also note: **1 of 2 free ZeroGPU slots is already consumed** by this Space. The remaining slot is
the last free Gradio Space this account can create without PRO — budget it against the other
HF-profile projects. `cpu-basic` is not a fallback; it returns HTTP 402 for Gradio Spaces.

---

## Traps found the hard way — do not re-derive these

`01-RESEARCH.md` is **wrong** on these points. Copy from
`src/japanese_avatar/ui/avatar_component.py` and `avatar/asr.js` instead.

1. **`js_on_load` cannot use top-level `await`.** Gradio 6.22.0 compiles it into a plain
   non-async `Function`, so the documented snippet throws `SyntaxError` and `window.Avatar`
   never exists. Use an async IIFE with its own `.catch`.
2. **Custom props are `**kwargs`, not `props={...}`.** The dict form creates one prop literally
   named `props`, so `props.vrmUrl` comes back `undefined`.
3. **ASR `dtype:'q8'` cannot create an ONNX session on the WASM backend at all** — for every
   Whisper size tested. It works on WebGPU, so the prescribed default would ship green on a dev
   machine and dead on exactly the browsers the WASM tier exists to serve. **`q4` is the only
   quantisation working on both tiers.**
4. **The prescribed WebGPU→WASM fallback does not fall back.** A failed WebGPU init poisons the
   ONNX Runtime Web backend registry for the whole page. `avatar/asr.js` probes
   `requestAdapter()` first so the doomed call is never issued.

## Two spec corrections carried forward

- **Frame quantisation:** the correct form is `round(round(sec * 93.75) / speed)`, **not**
  `round(sec / speed * 93.75)`. Measured 24/24 correct vs 8/24. Already shipped in `visemes.py`.
- **The slow/long duration ratio is 1.3410852713178294, not exactly 1/0.75** — 4 frames out from
  what `01-VALIDATION.md` and plan 01-06 assert. The plan's stated truth was written from a wrong
  assumption. **Update `01-VALIDATION.md` when 01-10 touches it**, so the phase verifier does not
  read this as an unmet must-have.

## Standing warning about requirements

Phase 1 plan frontmatter **over-claims shared requirements**. AVTR-01, AVTR-02, VOIC-02/03/04/05,
DPLY-01 and DPLY-04 all appear in the `requirements:` field of plans that do not actually satisfy
them. Never blind-run `requirements mark-complete` — verify the acceptance test has really run.
Three executors this session correctly declined to mark requirements complete for this reason.