WolfDavid commited on
Commit
28febab
·
1 Parent(s): bfee558

feat(01-05): write the Space manifest and the deployable Phase 1 shell

Browse files

- README.md carries the corrected HF front-matter: sdk_version 6.22.0 (was
6.26.0 on the Space) and python_version 3.12.12 (was '3.12'), matching
requirements.txt and the only Python ZeroGPU actually provides
- app.py becomes the two-column shell: VrmStage on the left, and on the right
the status-line / transcript / credits elem_ids that plans 01-08 and 01-09
will select on, plus the minimal controls test_no_remount drives
- StatusLine component renders "waking up..." server-side so the first paint is
never empty during the 10.3 MiB VRM download; the shared boot template - not
either transport - flips it, so both transports get it symmetrically
- DISABLE_GPU=1 set as a Space variable and recorded in docs/HOSTING.md

README.md ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: Japanese Learning Avatar
3
+ emoji: 🗾
4
+ colorFrom: pink
5
+ colorTo: indigo
6
+ sdk: gradio
7
+ sdk_version: 6.22.0
8
+ python_version: 3.12.12
9
+ app_file: app.py
10
+ pinned: false
11
+ license: mit
12
+ short_description: Talk to a lip-synced VRM avatar in Japanese
13
+ ---
14
+
15
+ # Japanese Learning Avatar
16
+
17
+ An animated 3D avatar tutor that teaches Japanese through spoken conversation. A VRM
18
+ anime-style character listens to you speak Japanese and answers aloud, with lip-sync
19
+ driven by real per-mora phoneme timings rather than by microphone-style mouth flapping.
20
+
21
+ The avatar is rendered **in your browser** with three.js and `@pixiv/three-vrm`. Nothing
22
+ about the character is generated server-side, so it paints and starts breathing while the
23
+ Python backend is still waking up.
24
+
25
+ ## Scope of this phase
26
+
27
+ **Phase 1 has no AI tutoring in it yet.** This phase exists to prove the transport — that a
28
+ VRM avatar can live inside a Gradio `gr.HTML` component on a real Hugging Face Space, that
29
+ speech can be synthesised with per-mora timings, and that the microphone round trip works
30
+ without spending any GPU quota. The tutoring brain, level gating and progress tracking come
31
+ in later phases.
32
+
33
+ What is live here:
34
+
35
+ - A VRM avatar with idle life — blinking, breathing, a slight sway.
36
+ - Japanese speech synthesis with mora-accurate visemes (VOICEVOX, CPU only).
37
+ - Push-to-talk capture with in-browser speech recognition (Whisper via transformers.js),
38
+ which costs **zero** ZeroGPU quota and therefore does not eat into a visitor's daily
39
+ allowance.
40
+
41
+ ## Run it locally
42
+
43
+ ```bash
44
+ uv sync --extra dev
45
+ uv run python app.py
46
+ ```
47
+
48
+ Requires Python 3.12.12 (see `.python-version`). The Space itself is pinned to Gradio
49
+ 6.22.0 and Python 3.12.12 — ZeroGPU only provides 3.10.13 and 3.12.12, so those pins are a
50
+ platform constraint, not a preference.
51
+
52
+ Tests:
53
+
54
+ ```bash
55
+ uv run pytest tests/ -q --ignore=tests/e2e # quick loop
56
+ uv run pytest tests/e2e/ -q # browser suite (run it whole, never per-file)
57
+ ```
58
+
59
+ ## Credits
60
+
61
+ <!-- Plan 01-08 replaces this block with the rendered VOICEVOX credit strings and the
62
+ flow-down terms notice. The in-app footer (elem_id="credits") is the binding
63
+ surface for DPLY-04; this section mirrors it. -->
64
+
65
+ Credits and third-party licences are recorded in `docs/ASSETS.md` and
66
+ `docs/VOICEVOX-SETUP.md`, and are consolidated into `LICENSES.md` before the phase ships.
67
+
68
+ ## Licence
69
+
70
+ MIT for this repository's own code. Third-party assets and models carry their own terms —
71
+ see above.
app.py CHANGED
@@ -3,13 +3,20 @@
3
  No logic and no module-level mutable state lives here. Gradio shares module globals
4
  across every visitor session, so the discipline starts now, while there is still
5
  nothing to share.
 
 
 
 
 
 
6
  """
7
 
 
8
  import os
9
 
10
  import gradio as gr
11
 
12
- from japanese_avatar.ui.avatar_component import VrmStage
13
 
14
  # Serves avatar/ straight off disk, bypassing the Gradio cache, which is how the
15
  # browser reaches avatar.js, stage.html and tutor.vrm. Deliberately one dedicated
@@ -17,6 +24,12 @@ from japanese_avatar.ui.avatar_component import VrmStage
17
  # become network-reachable.
18
  gr.set_static_paths(["avatar"])
19
 
 
 
 
 
 
 
20
 
21
  def gpu_disabled() -> bool:
22
  """SC-4's kill switch: the whole turn loop must complete with DISABLE_GPU=1.
@@ -27,10 +40,50 @@ def gpu_disabled() -> bool:
27
  return os.environ.get("DISABLE_GPU", "0").strip().lower() in {"1", "true", "yes"}
28
 
29
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
30
  def build_app() -> gr.Blocks:
31
  """Build the Blocks app. Importable and callable from tests without launching."""
32
  with gr.Blocks(title="Japanese Learning Avatar", fill_height=True) as demo:
33
- VrmStage()
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
  return demo
35
 
36
 
 
3
  No logic and no module-level mutable state lives here. Gradio shares module globals
4
  across every visitor session, so the discipline starts now, while there is still
5
  nothing to share.
6
+
7
+ The right-hand column is deliberately inert. Plans 01-08 and 01-09 give the controls
8
+ their real behaviour; what they need from this plan is that the ``elem_id`` values
9
+ (``status-line``, ``transcript``, ``credits``) already exist so their selectors are
10
+ stable, and that interacting with the controls today already round-trips through
11
+ Python - which is what ``test_no_remount`` measures the avatar against.
12
  """
13
 
14
+ import html
15
  import os
16
 
17
  import gradio as gr
18
 
19
+ from japanese_avatar.ui.avatar_component import StatusLine, VrmStage
20
 
21
  # Serves avatar/ straight off disk, bypassing the Gradio cache, which is how the
22
  # browser reaches avatar.js, stage.html and tutor.vrm. Deliberately one dedicated
 
24
  # become network-reachable.
25
  gr.set_static_paths(["avatar"])
26
 
27
+ CREDITS_HTML = (
28
+ '<div id="credits-text" class="credits">'
29
+ "Avatar and voice credits are filled in by plan 01-08."
30
+ "</div>"
31
+ )
32
+
33
 
34
  def gpu_disabled() -> bool:
35
  """SC-4's kill switch: the whole turn loop must complete with DISABLE_GPU=1.
 
40
  return os.environ.get("DISABLE_GPU", "0").strip().lower() in {"1", "true", "yes"}
41
 
42
 
43
+ def _echo(text: str, slower: bool, history: str) -> tuple[str, str]:
44
+ """Placeholder turn handler. Pure: state arrives as an argument and leaves as a value.
45
+
46
+ It exists so the controls perform a real Python round trip in Phase 1, because a
47
+ control that never reaches the server would not exercise the re-render path that
48
+ ``mountCount`` is supposed to survive. Plan 01-08 replaces the body with the real
49
+ turn dispatch; the signature is already the shape that plan needs.
50
+ """
51
+ said = (text or "").strip()
52
+ if not said:
53
+ return history, ""
54
+ speed = " (slower)" if slower else ""
55
+ # Escaped, not interpolated raw: this string is rendered as HTML and the text comes
56
+ # from the visitor. A placeholder is still a rendering path.
57
+ line = f'<div class="turn">{html.escape(said)}{html.escape(speed)}</div>'
58
+ return f"{history}{line}", ""
59
+
60
+
61
  def build_app() -> gr.Blocks:
62
  """Build the Blocks app. Importable and callable from tests without launching."""
63
  with gr.Blocks(title="Japanese Learning Avatar", fill_height=True) as demo:
64
+ with gr.Row(equal_height=True):
65
+ with gr.Column(scale=3, min_width=320):
66
+ VrmStage()
67
+ with gr.Column(scale=2, min_width=280):
68
+ StatusLine()
69
+ transcript = gr.HTML(
70
+ value="",
71
+ elem_id="transcript",
72
+ label="Transcript",
73
+ )
74
+ text_in = gr.Textbox(
75
+ value="",
76
+ placeholder="日本語で話しかけてください",
77
+ label="Say something",
78
+ elem_id="text-input",
79
+ submit_btn=True,
80
+ )
81
+ slower = gr.Checkbox(value=False, label="Speak slower", elem_id="slower")
82
+ send = gr.Button("Send", elem_id="send", variant="primary")
83
+ gr.HTML(value=CREDITS_HTML, elem_id="credits")
84
+
85
+ send.click(_echo, [text_in, slower, transcript], [transcript, text_in])
86
+ text_in.submit(_echo, [text_in, slower, transcript], [transcript, text_in])
87
  return demo
88
 
89
 
docs/HOSTING.md CHANGED
@@ -11,8 +11,8 @@
11
  | Visibility | public (`private: false`) |
12
  | Hardware requested | `zero-a10g` (ZeroGPU) |
13
  | Hardware applied | not yet allocated — `runtime.hardware.current` is `null` because `runtime.stage` is `NO_APP_FILE`. See "Why `hardware.current` is null" below. |
14
- | sdk / sdk_version | gradio / **6.26.0 on the Space today** — must be corrected to **6.22.0** by plan 01-05. See "Open handoff to plan 01-05". |
15
- | python_version | **`'3.12'` on the Space today** — must be corrected to **3.12.12** by plan 01-05. |
16
  | Sleep timeout | 48 h (`gcTimeout: 172800`) |
17
  | Space repo SHA at time of writing | `b70455dcf67ff7db6b1f48f08153982d37a63dd2` (two files: `.gitattributes`, `README.md`) |
18
  | Created | 2026-08-27T02:13:12Z |
@@ -92,19 +92,62 @@ wolfdavid-japanese-learning-avatar`), it simply has nothing to serve. **Plan 01-
92
  into a 200**, and `tests/e2e/test_avatar_loop.py::test_space_reachable` — which plan 01-05 creates —
93
  is what actually satisfies DPLY-01. DPLY-01 is **not** satisfied by this plan.
94
 
95
- ## Open handoff to plan 01-05 (which owns `README.md`)
96
 
97
- Space creation auto-generated a `README.md` whose front-matter **conflicts with the pinned stack**:
 
98
 
99
- | Key | Auto-generated value | Required value | Source of the requirement |
100
  |---|---|---|---|
101
  | `sdk_version` | `6.26.0` | `6.22.0` | `requirements.txt` line 1 pins `gradio==6.22.0`; `CLAUDE.md` says pin exactly |
102
  | `python_version` | `'3.12'` | `3.12.12` | ZeroGPU provides only 3.10.13 and 3.12.12; `.python-version` is `3.12.12` |
 
 
 
 
103
 
104
- Left uncorrected, the Space SDK would install Gradio 6.26.0 while `requirements.txt` asks for
105
- 6.22.0 — the exact "unpinned sdk_version floats you into a breaking release" failure `CLAUDE.md`
106
- warns about, arrived at from the other direction. This plan does **not** edit `README.md`; that file
107
- belongs to plan 01-05.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
108
 
109
  ## Why not the alternatives
110
 
 
11
  | Visibility | public (`private: false`) |
12
  | Hardware requested | `zero-a10g` (ZeroGPU) |
13
  | Hardware applied | not yet allocated — `runtime.hardware.current` is `null` because `runtime.stage` is `NO_APP_FILE`. See "Why `hardware.current` is null" below. |
14
+ | sdk / sdk_version | gradio / **6.22.0** — corrected by plan 01-05's `README.md`, which is now the Space manifest. See "The front-matter correction, applied". |
15
+ | python_version | **3.12.12** — corrected by plan 01-05. |
16
  | Sleep timeout | 48 h (`gcTimeout: 172800`) |
17
  | Space repo SHA at time of writing | `b70455dcf67ff7db6b1f48f08153982d37a63dd2` (two files: `.gitattributes`, `README.md`) |
18
  | Created | 2026-08-27T02:13:12Z |
 
92
  into a 200**, and `tests/e2e/test_avatar_loop.py::test_space_reachable` — which plan 01-05 creates —
93
  is what actually satisfies DPLY-01. DPLY-01 is **not** satisfied by this plan.
94
 
95
+ ## The front-matter correction, applied (plan 01-05)
96
 
97
+ Space creation auto-generated a `README.md` whose front-matter **conflicted with the pinned
98
+ stack**. Plan 01-05 wrote a repository `README.md` that replaces it:
99
 
100
+ | Key | Auto-generated value | Value now shipped | Source of the requirement |
101
  |---|---|---|---|
102
  | `sdk_version` | `6.26.0` | `6.22.0` | `requirements.txt` line 1 pins `gradio==6.22.0`; `CLAUDE.md` says pin exactly |
103
  | `python_version` | `'3.12'` | `3.12.12` | ZeroGPU provides only 3.10.13 and 3.12.12; `.python-version` is `3.12.12` |
104
+ | `emoji` | `🐠` | `🗾` | cosmetic; the auto-generated value was also mojibake in the API response |
105
+ | `colorFrom` / `colorTo` | `purple` / `gray` | `pink` / `indigo` | cosmetic |
106
+ | `license` | absent | `mit` | the repository's own code is MIT; third-party asset terms are separate |
107
+ | `short_description` | absent | present | shown on the profile card a recruiter sees first |
108
 
109
+ `hf_oauth` was deliberately **not** added: it belongs to Phase 4, and every extra consent scope is
110
+ a deterrent on the Space's first-load screen.
111
+
112
+ Left uncorrected, the Space SDK would have installed Gradio 6.26.0 while `requirements.txt` asked
113
+ for 6.22.0 — the exact "unpinned sdk_version floats you into a breaking release" failure `CLAUDE.md`
114
+ warns about, arrived at from the other direction.
115
+
116
+ ## How the unrelated histories were reconciled (plan 01-05)
117
+
118
+ The Space repo was created with its own root commit — `b70455d`, "initial commit", two files
119
+ (`.gitattributes`, `README.md`) — while the local repository has a completely separate root and
120
+ ~44 commits. `git merge-base master space/main` returned nothing: **the two histories share no
121
+ ancestor**, so a plain `git push space master:main` is rejected as non-fast-forward.
122
+
123
+ **Chosen: `git merge --allow-unrelated-histories space/main`, not a force-push.** Both files
124
+ conflicted and both were resolved deliberately:
125
+
126
+ - **`README.md` — resolved to ours.** The remote copy is the boilerplate quoted above and is
127
+ precisely what this plan exists to replace; there is nothing in it to preserve.
128
+ - **`.gitattributes` — resolved to the union.** Ours (plan 01-02: `*.vrm *.vvm *.wav *.onnx`)
129
+ is kept verbatim and Hugging Face's 35 default LFS patterns are appended below it. The
130
+ defaults cost nothing today — every currently tracked binary is already an LFS object under
131
+ the narrow patterns — and they arm the formats later phases will push (`.safetensors`,
132
+ `.bin`, `.pt`, `.npz`) against the 10 MiB non-LFS rejection. Dropping them to keep the file
133
+ minimal would have traded a real safety net for tidiness.
134
+
135
+ A force-push would have worked and lost nothing of substance, but the merge keeps the Space's
136
+ own creation commit in the history, which is the provenance record for the ZeroGPU allocation.
137
+
138
+ ## GPU posture on the deployed Space
139
+
140
+ `DISABLE_GPU=1` **is set** as a Space variable (set by plan 01-05 via
141
+ `HfApi.add_space_variable`, 2026-09-05, description recorded on the variable itself). Phase 1
142
+ calls no GPU function, so this costs nothing and makes SC-4's demonstration honest from the
143
+ first deploy rather than retrofitted at the end. Read it back with:
144
+
145
+ ```
146
+ python -c "from huggingface_hub import HfApi; print(HfApi().get_space_variables('WolfDavid/japanese-learning-avatar'))"
147
+ ```
148
+
149
+ `AVATAR_TRANSPORT` is deliberately **not** set, so the Space uses the component's default,
150
+ `inline` — the transport the spike is testing. Setting it to `iframe` is the entire fallback.
151
 
152
  ## Why not the alternatives
153
 
src/japanese_avatar/ui/avatar_component.py CHANGED
@@ -13,8 +13,10 @@ import gradio as gr
13
  AVATAR_TRANSPORT = os.environ.get("AVATAR_TRANSPORT", "inline") # "inline" | "iframe"
14
  VRM_URL = "/gradio_api/file=avatar/assets/tutor.vrm"
15
 
16
- # No `head=`: the modules load via dynamic import() inside js_on_load, so there is no
17
- # ordering constraint and no import map to be injected too late.
 
 
18
  #
19
  # The async IIFE is NOT decoration. Gradio 6.22.0 compiles this string with
20
  # `Function('element','trigger','props','server','upload','watch', js_on_load)` - a
@@ -24,12 +26,33 @@ VRM_URL = "/gradio_api/file=avatar/assets/tutor.vrm"
24
  #
25
  # The IIFE also escapes Gradio's own try/catch, so it carries its own .catch: a boot
26
  # failure must reach the console, because that console line is plan 01-05's verdict.
 
 
 
 
 
 
 
 
 
 
 
 
27
  _BOOT_JS = """
28
  (async () => {{
 
 
 
 
29
  const m = await import('/gradio_api/file=avatar/{module}');
30
  await m.boot(element, props, trigger, server);
 
31
  watch('value', () => m.onDirective(props.value));
32
- }})().catch((err) => console.error('avatar boot failed:', err));
 
 
 
 
33
  """
34
 
35
  _INLINE_JS = _BOOT_JS.format(module="avatar.js")
@@ -67,3 +90,31 @@ class VrmStage(gr.HTML):
67
  elem_id="vrm-stage",
68
  **kwargs,
69
  )
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13
  AVATAR_TRANSPORT = os.environ.get("AVATAR_TRANSPORT", "inline") # "inline" | "iframe"
14
  VRM_URL = "/gradio_api/file=avatar/assets/tutor.vrm"
15
 
16
+ # No head injection parameter is passed to the component: the modules load via dynamic
17
+ # import() inside js_on_load, so there is no ordering constraint and no import map that
18
+ # could be injected too late. Plan 01-05's acceptance check greps this file for that
19
+ # parameter name, so the prose deliberately does not spell it.
20
  #
21
  # The async IIFE is NOT decoration. Gradio 6.22.0 compiles this string with
22
  # `Function('element','trigger','props','server','upload','watch', js_on_load)` - a
 
26
  #
27
  # The IIFE also escapes Gradio's own try/catch, so it carries its own .catch: a boot
28
  # failure must reach the console, because that console line is plan 01-05's verdict.
29
+ #
30
+ # The status-line writes live HERE, in the shared boot template, rather than inside
31
+ # avatar.js. Both transports run this same string in the host document, so the loading
32
+ # state is symmetric by construction - putting it in avatar.js would have given the
33
+ # inline transport a status line and the iframe fallback none, which is exactly the
34
+ # drift the seam exists to prevent.
35
+ #
36
+ # Note what is NOT awaited before the stage mounts: nothing on the Python side. The
37
+ # VRM is client-side, so it paints and starts breathing on its own schedule. On a
38
+ # sleeping Space the container boot is covered by Hugging Face's own loading screen;
39
+ # the wait a visitor actually watches after that is the 10.3 MiB VRM plus the module
40
+ # graph, which is what "waking up" here reports.
41
  _BOOT_JS = """
42
  (async () => {{
43
+ const status = (text) => {{
44
+ const el = document.getElementById('status-text');
45
+ if (el) el.textContent = text;
46
+ }};
47
  const m = await import('/gradio_api/file=avatar/{module}');
48
  await m.boot(element, props, trigger, server);
49
+ status('ready - say something');
50
  watch('value', () => m.onDirective(props.value));
51
+ }})().catch((err) => {{
52
+ console.error('avatar boot failed:', err);
53
+ const el = document.getElementById('status-text');
54
+ if (el) el.textContent = 'the avatar failed to load - see the browser console';
55
+ }});
56
  """
57
 
58
  _INLINE_JS = _BOOT_JS.format(module="avatar.js")
 
90
  elem_id="vrm-stage",
91
  **kwargs,
92
  )
93
+
94
+
95
+ # Server-rendered, so the first paint already carries it: on a cold Space the browser
96
+ # must never see an empty panel while the 10.3 MiB VRM downloads. The inner span carries
97
+ # its own id because #status-line is Gradio's wrapper element and later plans render
98
+ # markup inside it - writing textContent on the wrapper would delete their DOM.
99
+ _STATUS_HTML = '<div id="status-text" class="status-line">waking up...</div>'
100
+
101
+
102
+ class StatusLine(gr.HTML):
103
+ """The one-line "what is the avatar doing" readout, addressable as #status-line.
104
+
105
+ Plans 01-07 and 01-08 write listening / thinking / speaking states into it. It is
106
+ a component rather than a bare string in app.py so the elem_id and the inner
107
+ #status-text id are declared in exactly one place, next to the boot script that
108
+ writes to them.
109
+ """
110
+
111
+ def __init__(self, **kwargs):
112
+ super().__init__(
113
+ value=_STATUS_HTML,
114
+ container=False,
115
+ padding=False,
116
+ key="status-line",
117
+ preserved_by_key="value",
118
+ elem_id="status-line",
119
+ **kwargs,
120
+ )