WolfDavid's picture
docs: hand off Phase 01 at the deploy gate
da46df0
|
Raw History Blame
6.05 kB

Resume here — Japanese Learning Avatar

Written 2026-08-27, stepping away from the machine. Working tree is clean. Everything is committed. Nothing is pushed, deliberately — see below.

To pick back up, run /gsd:resume-work in this directory. It reads .planning/HANDOFF.json first, which carries the same state in machine-readable form.


One-line status

Phase 01 is 6 of 10 plans complete and hard-blocked on a decision only you can make: whether to publicly deploy the Hugging Face Space.


Why nothing was pushed

This repo's only git remote is space — the public Hugging Face Space at WolfDavid/japanese-learning-avatar. There is no GitHub remote. Branch master has no upstream.

So git push here does not mean "back up my work." It means publish the app publicly and trigger a Space rebuild — which is precisely plan 01-05 Task 1, the gate that is waiting on you.

Pushing right now would also produce a broken public build. There is no README.md in this repo; the Space's front-matter lives only on the HF side and was auto-generated wrong when the Space was created:

Front-matter key Currently on HF Must become Why
sdk_version 6.26.0 6.22.0 must match gradio==6.22.0 in requirements.txt
python_version '3.12' 3.12.12 ZeroGPU provides only 3.10.13 and 3.12.12

Writing that corrected README is part of plan 01-05. Push after 01-05 runs, not before.

42 commits are unpushed by design, not by omission.

If what you actually want is an off-machine backup, that is a different action from deploying: add a private GitHub remote and push there. That was not done unprompted because the standing rule for this account is to ask before pushing nested repos.


Where the work stands

Plan What it is State
01-01 Hosting decision + Space creation + VRM sourcing Complete
01-02 Toolchain, Git LFS, package skeleton, test scaffolding Complete
01-03 Avatar stage, shared facade + turn loop, both transports Complete (resumed from the crash)
01-04 VOICEVOX TTS, requirements.txt, ground-truth fixtures Complete
01-06 Mora-to-viseme timeline builder Complete
01-07 Push-to-talk mic gate + tiered browser ASR Complete
01-05 Space manifest, deploy the spike, verdict BLOCKED — your call
01-08 Turn-loop wiring — closes the round trip Blocked behind 01-05
01-09 Full deployed E2E suite + LICENSES.md Blocked behind 01-08
01-10 Vendor modules, latency harness (user-gated) Blocked behind 01-09

The dependency chain is real, not a preference: 01-08 depends_on [01-05, 01-06, 01-07], 01-09 depends_on [01-08], 01-10 depends_on [01-09]. Nothing proceeds until 01-05 does.

Verified green at handoff

  • 82 quick-loop tests in ~12.8 s
  • 15 end-to-end browser tests
  • ruff check . and ruff format --check . both clean

Run pytest tests/e2e/ as a whole before trusting a green phase. Per-file runs hid two real defects this session — each plan's own verification passed while the full suite failed.


What deploying would actually test

The risky part is already retired. three.js + @pixiv/three-vrm render inside a real Gradio 6.22.0 gr.HTML component locally, with threeInstanceCount === 1 and mountCount === 1, and the iframe fallback boots identically under AVATAR_TRANSPORT=iframe.

What remains genuinely deployment-specific: CDN reach from *.hf.space, set_static_paths behind the Space proxy, cold-start behaviour, and mobile.

Also note: 1 of 2 free ZeroGPU slots is already consumed by this Space. The remaining slot is the last free Gradio Space this account can create without PRO — budget it against the other HF-profile projects. cpu-basic is not a fallback; it returns HTTP 402 for Gradio Spaces.


Traps found the hard way — do not re-derive these

01-RESEARCH.md is wrong on these points. Copy from src/japanese_avatar/ui/avatar_component.py and avatar/asr.js instead.

  1. js_on_load cannot use top-level await. Gradio 6.22.0 compiles it into a plain non-async Function, so the documented snippet throws SyntaxError and window.Avatar never exists. Use an async IIFE with its own .catch.
  2. Custom props are **kwargs, not props={...}. The dict form creates one prop literally named props, so props.vrmUrl comes back undefined.
  3. ASR dtype:'q8' cannot create an ONNX session on the WASM backend at all — for every Whisper size tested. It works on WebGPU, so the prescribed default would ship green on a dev machine and dead on exactly the browsers the WASM tier exists to serve. q4 is the only quantisation working on both tiers.
  4. The prescribed WebGPU→WASM fallback does not fall back. A failed WebGPU init poisons the ONNX Runtime Web backend registry for the whole page. avatar/asr.js probes requestAdapter() first so the doomed call is never issued.

Two spec corrections carried forward

  • Frame quantisation: the correct form is round(round(sec * 93.75) / speed), not round(sec / speed * 93.75). Measured 24/24 correct vs 8/24. Already shipped in visemes.py.
  • The slow/long duration ratio is 1.3410852713178294, not exactly 1/0.75 — 4 frames out from what 01-VALIDATION.md and plan 01-06 assert. The plan's stated truth was written from a wrong assumption. Update 01-VALIDATION.md when 01-10 touches it, so the phase verifier does not read this as an unmet must-have.

Standing warning about requirements

Phase 1 plan frontmatter over-claims shared requirements. AVTR-01, AVTR-02, VOIC-02/03/04/05, DPLY-01 and DPLY-04 all appear in the requirements: field of plans that do not actually satisfy them. Never blind-run requirements mark-complete — verify the acceptance test has really run. Three executors this session correctly declined to mark requirements complete for this reason.