WolfDavid commited on
Commit
16c6052
Β·
1 Parent(s): 84bdf64

docs(01-04): complete Japanese voice plan

Browse files
.planning/REQUIREMENTS.md CHANGED
@@ -13,7 +13,7 @@
13
 
14
  ### Voice Loop
15
 
16
- - [ ] **VOIC-01**: Avatar speaks Japanese aloud via TTS that emits per-mora timings (CPU-path, zero GPU quota)
17
  - [ ] **VOIC-02**: Learner can speak Japanese via push-to-talk; ASR transcribes with visible transcript
18
  - [ ] **VOIC-03**: Learner can replay the avatar's last utterance and request it slower
19
  - [ ] **VOIC-04**: Learner can type instead of speak at any point (full text fallback)
@@ -102,7 +102,7 @@ Which phases cover which requirements. Updated during roadmap creation.
102
  | AVTR-01 | Phase 1 | Pending |
103
  | AVTR-02 | Phase 1 | Pending |
104
  | AVTR-03 | Phase 3 | Pending |
105
- | VOIC-01 | Phase 1 | Pending |
106
  | VOIC-02 | Phase 1 | Pending |
107
  | VOIC-03 | Phase 1 | Pending |
108
  | VOIC-04 | Phase 1 | Pending |
 
13
 
14
  ### Voice Loop
15
 
16
+ - [x] **VOIC-01**: Avatar speaks Japanese aloud via TTS that emits per-mora timings (CPU-path, zero GPU quota)
17
  - [ ] **VOIC-02**: Learner can speak Japanese via push-to-talk; ASR transcribes with visible transcript
18
  - [ ] **VOIC-03**: Learner can replay the avatar's last utterance and request it slower
19
  - [ ] **VOIC-04**: Learner can type instead of speak at any point (full text fallback)
 
102
  | AVTR-01 | Phase 1 | Pending |
103
  | AVTR-02 | Phase 1 | Pending |
104
  | AVTR-03 | Phase 3 | Pending |
105
+ | VOIC-01 | Phase 1 | Complete |
106
  | VOIC-02 | Phase 1 | Pending |
107
  | VOIC-03 | Phase 1 | Pending |
108
  | VOIC-04 | Phase 1 | Pending |
.planning/ROADMAP.md CHANGED
@@ -34,7 +34,7 @@ Decimal phases appear between their surrounding integers in numeric order.
34
  **Plans**: 10 plans in 7 waves
35
  - [x] 01-02-PLAN.md β€” Toolchain, LFS arming, package skeleton, test scaffolding (local-only) [wave 1]
36
  - [ ] 01-01-PLAN.md β€” Hosting decision + Space creation + VRM sourcing (user-gated) [wave 2]
37
- - [ ] 01-04-PLAN.md β€” VOICEVOX TTS, requirements.txt, ground-truth AudioQuery fixtures [wave 2]
38
  - [ ] 01-03-PLAN.md β€” Avatar stage, shared facade + turn loop, both transports [wave 3]
39
  - [ ] 01-05-PLAN.md β€” Space manifest, deploy the spike, first deployed E2E, verdict (user-gated) [wave 4]
40
  - [ ] 01-06-PLAN.md β€” Mora-to-viseme timeline builder, test-first [wave 4]
@@ -109,7 +109,7 @@ Phases execute in numeric order: 1 β†’ 2 β†’ 3 β†’ 4 β†’ 5 β†’ 6
109
 
110
  | Phase | Plans Complete | Status | Completed |
111
  |-------|----------------|--------|-----------|
112
- | 1. Voice + Avatar Loop Skeleton | 1/10 | In Progress | - |
113
  | 2. Japanese Language Core | 0/TBD | Not started | - |
114
  | 3. Tutoring Brain | 0/TBD | Not started | - |
115
  | 4. Accounts & Persistence | 0/TBD | Not started | - |
 
34
  **Plans**: 10 plans in 7 waves
35
  - [x] 01-02-PLAN.md β€” Toolchain, LFS arming, package skeleton, test scaffolding (local-only) [wave 1]
36
  - [ ] 01-01-PLAN.md β€” Hosting decision + Space creation + VRM sourcing (user-gated) [wave 2]
37
+ - [x] 01-04-PLAN.md β€” VOICEVOX TTS, requirements.txt, ground-truth AudioQuery fixtures [wave 2]
38
  - [ ] 01-03-PLAN.md β€” Avatar stage, shared facade + turn loop, both transports [wave 3]
39
  - [ ] 01-05-PLAN.md β€” Space manifest, deploy the spike, first deployed E2E, verdict (user-gated) [wave 4]
40
  - [ ] 01-06-PLAN.md β€” Mora-to-viseme timeline builder, test-first [wave 4]
 
109
 
110
  | Phase | Plans Complete | Status | Completed |
111
  |-------|----------------|--------|-----------|
112
+ | 1. Voice + Avatar Loop Skeleton | 2/10 | In Progress | - |
113
  | 2. Japanese Language Core | 0/TBD | Not started | - |
114
  | 3. Tutoring Brain | 0/TBD | Not started | - |
115
  | 4. Accounts & Persistence | 0/TBD | Not started | - |
.planning/STATE.md CHANGED
@@ -3,13 +3,13 @@ gsd_state_version: 1.0
3
  milestone: v1.0
4
  milestone_name: milestone
5
  status: unknown
6
- stopped_at: Completed 01-02-PLAN.md (local foundations). Plan 01-01 still gated on the human hosting decision.
7
- last_updated: "2026-08-27T00:02:37.063Z"
8
  progress:
9
  total_phases: 6
10
  completed_phases: 0
11
  total_plans: 10
12
- completed_plans: 1
13
  ---
14
 
15
  # Project State
@@ -24,36 +24,39 @@ See: .planning/PROJECT.md (updated 2026-08-08)
24
  ## Current Position
25
 
26
  Phase: 01 (voice-avatar-loop-skeleton) β€” EXECUTING
27
- Plan: 2 of 10 (1 complete)
28
 
29
- **Completed:** 01-02 (local foundations β€” pytest/ruff on Python 3.12.12, Git LFS armed, `--space-url` plumbing, audio fixtures).
30
- **Not started:** 01-01 β€” it is deliberately out of order. 01-02 is local-only and network-free precisely so it could run in parallel with 01-01's **human hosting checkpoint, which has NOT cleared**. Do not treat "Plan 2" as meaning 01-01 is done.
31
- **Next actionable:** 01-03 (avatar transport seam) is unblocked β€” its dependencies are the test loop and package skeleton, both of which now exist. 01-01, 01-04 and 01-05 remain gated on the hosting decision and the VRM/VOICEVOX asset choices.
 
 
32
 
33
  ## Performance Metrics
34
 
35
  **Velocity:**
36
 
37
- - Total plans completed: 1
38
- - Average duration: 21 min
39
- - Total execution time: 0.35 hours
40
 
41
  **By Phase:**
42
 
43
  | Phase | Plans | Total | Avg/Plan |
44
  |-------|-------|-------|----------|
45
- | 01 | 1 | 21min | 21min |
46
 
47
  **Per-plan detail:**
48
 
49
  | Plan | Duration | Tasks | Files |
50
  |------|----------|-------|-------|
51
  | Phase 01 P02 | 21min | 3 tasks | 16 files |
 
52
 
53
  **Recent Trend:**
54
 
55
- - Last 5 plans: 01-02 (21min)
56
- - Trend: β€” (single data point)
57
 
58
  *Updated after each plan completion*
59
 
@@ -71,6 +74,9 @@ Recent decisions affecting current work:
71
  - [Phase 01]: Ruff excludes *.md: ruff 0.16 formats Python blocks embedded in Markdown, which made the phase's own planning docs fail the lint gate. Lint governs shipped Python, not prose.
72
  - [Phase 01]: uv.lock is committed so a Space rebuild resolves the same dependency set that passed locally; requirements.txt (plan 01-04) remains the Space's contract.
73
  - [Phase 01]: Git LFS armed for .vrm/.vvm/.wav/.onnx in Wave 1, before any binary exists - retires the 10 MiB non-LFS push rejection ahead of plans 01-01 and 01-04.
 
 
 
74
 
75
  ### Pending Todos
76
 
@@ -81,13 +87,14 @@ None yet.
81
  ### Blockers/Concerns
82
 
83
  - **Hosting model unresolved (pre-Phase-1 decision).** STACK research found free Gradio Spaces on personal accounts now require PRO; the free path is ZeroGPU with per-visitor quota (~2 min/day anonymous). Confirm WolfDavid account eligibility (verified email, >30 days old, <2 existing ZeroGPU Spaces) during Phase 1 planning; if ineligible, PRO ($9/mo) becomes a decision before any build.
84
- - **Asset licensing gates a design branch.** VOICEVOX character terms need a human read (third-party web synthesis, not just generated audio) before TTS is built around it β€” the fallback loses free mora timings and changes the lip-sync design. VRM model licenses vary per author; self-authored VRoid Studio character removes the question.
85
  - **Per-turn GPU-second cost is unmeasured.** The entire free-tier UX budget depends on this number β€” measure in Phase 1/3, do not assume.
86
  - **REQUIREMENTS.md originally miscounted v1 as 31.** Actual count is 33; traceability and coverage corrected during roadmap creation.
87
  - **Phase 1 plan frontmatter over-claims shared requirements β€” do not blind-run `requirements mark-complete`.** DPLY-01 appears in the `requirements:` field of plans 01-01, 01-02, 01-05, 01-09 and 01-10, but it is only *satisfied* by 01-05 (`test_space_reachable`). Executing 01-02 marked it Complete; this was reverted, and DPLY-01 is correctly `Pending`. The same over-claiming applies to AVTR-01/02, VOIC-02/03/04/05 and DPLY-04 across plans 01-09 and 01-10. Verify a requirement's actual acceptance test has run before checking it off.
 
88
 
89
  ## Session Continuity
90
 
91
- Last session: 2026-08-27T00:02:37.058Z
92
- Stopped at: Completed 01-02-PLAN.md (local foundations). Plan 01-01 still gated on the human hosting decision.
93
  Resume file: None
 
3
  milestone: v1.0
4
  milestone_name: milestone
5
  status: unknown
6
+ stopped_at: Completed 01-04-PLAN.md (Japanese voice). VOIC-01 satisfied. Plan 01-01 still gated on the human hosting decision.
7
+ last_updated: "2026-08-27T00:44:59.201Z"
8
  progress:
9
  total_phases: 6
10
  completed_phases: 0
11
  total_plans: 10
12
+ completed_plans: 2
13
  ---
14
 
15
  # Project State
 
24
  ## Current Position
25
 
26
  Phase: 01 (voice-avatar-loop-skeleton) β€” EXECUTING
27
+ Plan: 3 of 10 (2 complete)
28
 
29
+ **Completed:** 01-02 (local foundations β€” pytest/ruff on Python 3.12.12, Git LFS armed, `--space-url` plumbing, audio fixtures) and 01-04 (Japanese voice β€” VOICEVOX assets via LFS, `synthesize()`, `requirements.txt`, ground-truth fixtures). **VOIC-01 is Complete.**
30
+ **Not started:** 01-01 β€” it is deliberately out of order. 01-02 and 01-04 are both local-only precisely so they could run in parallel with 01-01's **human hosting checkpoint, which has NOT cleared**. The plan counter is a count, not a plan ID: "Plan 3 of 10" does not mean 01-01 or 01-03 is done.
31
+ **Next actionable:** 01-03 (avatar transport seam) and 01-06 (viseme timeline) are both unblocked β€” 01-06's dependency was 01-04's fixtures, which now exist. 01-05 has its `requirements.txt` but still needs the hosting decision. 01-01 remains gated on the human checkpoint.
32
+
33
+ > **Before writing `visemes.py` in 01-06, read `docs/VOICEVOX-SETUP.md` Β§ "Frame quantisation".** 01-RESEARCH.md's formula is wrong for `speedScale != 1.0`; the corrected form and its evidence are recorded there and in the Blockers/Concerns section below.
34
 
35
  ## Performance Metrics
36
 
37
  **Velocity:**
38
 
39
+ - Total plans completed: 2
40
+ - Average duration: 27 min
41
+ - Total execution time: 0.9 hours
42
 
43
  **By Phase:**
44
 
45
  | Phase | Plans | Total | Avg/Plan |
46
  |-------|-------|-------|----------|
47
+ | 01 | 2 | 54min | 27min |
48
 
49
  **Per-plan detail:**
50
 
51
  | Plan | Duration | Tasks | Files |
52
  |------|----------|-------|-------|
53
  | Phase 01 P02 | 21min | 3 tasks | 16 files |
54
+ | Phase 01 P04 | 33min | 3 tasks | 26 files |
55
 
56
  **Recent Trend:**
57
 
58
+ - Last 5 plans: 01-02 (21min), 01-04 (33min)
59
+ - Trend: slower, but 01-04 downloaded ~190 MB of assets and ran an empirical investigation the plan did not budget for
60
 
61
  *Updated after each plan completion*
62
 
 
74
  - [Phase 01]: Ruff excludes *.md: ruff 0.16 formats Python blocks embedded in Markdown, which made the phase's own planning docs fail the lint gate. Lint governs shipped Python, not prose.
75
  - [Phase 01]: uv.lock is committed so a Space rebuild resolves the same dependency set that passed locally; requirements.txt (plan 01-04) remains the Space's contract.
76
  - [Phase 01]: Git LFS armed for .vrm/.vvm/.wav/.onnx in Wave 1, before any binary exists - retires the 10 MiB non-LFS push rejection ahead of plans 01-01 and 01-04.
77
+ - [Phase 01]: voicevox_vvm 0.17.0 (not 0.16.4) pairs with core 0.17.0; the vvm ships format v2 + StyleType::StreamingTalk, both introduced in core 0.17.0
78
+ - [Phase 01]: ONNX Runtime uses lazy first-use download (option c): no pip distribution exists and a Gradio-SDK Space has no build hook to run the official downloader
79
+ - [Phase 01]: AudioQuery serialised to a plain ENGINE-schema dict at the module boundary; voicevox_core 0.17.0 exposes no .json() method, but the type is a real dataclass
80
 
81
  ### Pending Todos
82
 
 
87
  ### Blockers/Concerns
88
 
89
  - **Hosting model unresolved (pre-Phase-1 decision).** STACK research found free Gradio Spaces on personal accounts now require PRO; the free path is ZeroGPU with per-visitor quota (~2 min/day anonymous). Confirm WolfDavid account eligibility (verified email, >30 days old, <2 existing ZeroGPU Spaces) during Phase 1 planning; if ineligible, PRO ($9/mo) becomes a decision before any build.
90
+ - **Asset licensing β€” VOICEVOX half RESOLVED (01-04), VRM half still open.** The VOICEVOX branch is closed: all three licence layers were read in full, third-party web synthesis is explicitly contemplated (software cl.3 attaches a flow-down obligation rather than prohibiting it), embedded `.vvm` redistribution is permitted (model cl.2), and the rights holder declines qualification questions as policy (SSS LLC 免責村項 2) β€” so there is no human step available or required. Compliance is a visible `VOICEVOX:γšγ‚“γ γ‚‚γ‚“` credit plus a terms notice wherever audio is obtainable; all facts recorded in `docs/VOICEVOX-SETUP.md`. **Still open:** VRM model licences vary per author; a self-authored VRoid Studio character removes that question (plan 01-01).
91
  - **Per-turn GPU-second cost is unmeasured.** The entire free-tier UX budget depends on this number β€” measure in Phase 1/3, do not assume.
92
  - **REQUIREMENTS.md originally miscounted v1 as 31.** Actual count is 33; traceability and coverage corrected during roadmap creation.
93
  - **Phase 1 plan frontmatter over-claims shared requirements β€” do not blind-run `requirements mark-complete`.** DPLY-01 appears in the `requirements:` field of plans 01-01, 01-02, 01-05, 01-09 and 01-10, but it is only *satisfied* by 01-05 (`test_space_reachable`). Executing 01-02 marked it Complete; this was reverted, and DPLY-01 is correctly `Pending`. The same over-claiming applies to AVTR-01/02, VOIC-02/03/04/05 and DPLY-04 across plans 01-09 and 01-10. Verify a requirement's actual acceptance test has run before checking it off.
94
+ - CORRECTION for plan 01-06 (not a blocker, a spec fix): 01-RESEARCH.md's frame quantisation is wrong for speedScale != 1.0. Correct form is round(round(length*93.75)/speed), NOT round(length/speed*93.75) - verified 24/24 vs 8/24 across 6 speeds x 4 sentences, worst error 5 frames (53ms). Also: pauseLength/pauseLengthScale do not exist in voicevox_core 0.17.0, and the slow/long duration ratio is 1.341085, not exactly 1/0.75. See docs/VOICEVOX-SETUP.md section 'Frame quantisation'.
95
 
96
  ## Session Continuity
97
 
98
+ Last session: 2026-08-27T00:43:56.253Z
99
+ Stopped at: Completed 01-04-PLAN.md (Japanese voice). VOIC-01 satisfied. Plan 01-01 still gated on the human hosting decision.
100
  Resume file: None
.planning/phases/01-voice-avatar-loop-skeleton/01-04-SUMMARY.md ADDED
@@ -0,0 +1,488 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ phase: 01-voice-avatar-loop-skeleton
3
+ plan: 04
4
+ subsystem: voice
5
+ tags: [voicevox, tts, japanese, mora-timing, git-lfs, onnxruntime, licensing, fixtures]
6
+
7
+ # Dependency graph
8
+ requires: ["01-02"]
9
+ provides:
10
+ - "synthesize() returning 24000 Hz WAV bytes plus a plain-dict AudioQuery with per-mora timings"
11
+ - "SynthResult / VisemeEvent / TurnTimings / AvatarDirective dataclasses consumed by 01-06 and 01-08"
12
+ - "zundamon.vvm and the Open JTalk dictionary on disk via Git LFS, loaded with no network at import"
13
+ - "requirements.txt - the Space's four-line dependency contract"
14
+ - "Three AudioQuery fixtures with true WAV durations and frame counts, the ground truth for AVTR-02"
15
+ - "The VERIFIED frame-quantisation formula, which corrects 01-RESEARCH.md"
16
+ - "docs/VOICEVOX-SETUP.md - executed API surface, asset hashes, and the full voice licence chain"
17
+ affects: [01-05, 01-06, 01-07, 01-08, 01-09]
18
+
19
+ # Tech tracking
20
+ tech-stack:
21
+ added: [voicevox_core 0.17.0, voicevox_vvm 0.17.0, VOICEVOX ONNX Runtime 1.23.2, open_jtalk_dic_utf_8-1.11]
22
+ patterns:
23
+ - "Lazy lru_cache(maxsize=1) synthesizer with a warmup() hook: import stays free, startup pays the cost"
24
+ - "AudioQuery converted to a plain ENGINE-schema dict at the module boundary, never handed out live"
25
+ - "Large runtime assets committed through LFS; software binaries referenced by release URL, never vendored"
26
+ - "Directory-scoped .gitattributes to extend LFS coverage without touching another plan's root file"
27
+ - "Fixture generators assert the invariants their fixtures exist to prove before writing anything"
28
+
29
+ key-files:
30
+ created:
31
+ - docs/VOICEVOX-SETUP.md
32
+ - requirements.txt
33
+ - src/japanese_avatar/voice/tts.py
34
+ - src/japanese_avatar/voice/models.py
35
+ - tests/test_tts_contract.py
36
+ - tests/fixtures/make_synth_fixtures.py
37
+ - tests/fixtures/synth_meta.json
38
+ - tests/fixtures/audio_query_short.json
39
+ - tests/fixtures/audio_query_long.json
40
+ - tests/fixtures/audio_query_slow.json
41
+ - tests/fixtures/speech_ja.wav
42
+ - tests/fixtures/speech_ja_long.wav
43
+ - tests/fixtures/speech_ja_slow.wav
44
+ - voicevox/.gitattributes
45
+ - voicevox/model/zundamon.vvm
46
+ - voicevox/open_jtalk_dic_utf_8-1.11/
47
+ modified:
48
+ - pyproject.toml
49
+ - uv.lock
50
+
51
+ key-decisions:
52
+ - "voicevox_vvm 0.17.0, not 0.16.4: the matched release exists and ships vvm_format_version 2, which core 0.17.0 is built around"
53
+ - "ONNX Runtime option (c) lazy first-use download: (a) has no pip distribution at all and (b) has no build hook on a Gradio-SDK Space"
54
+ - "Directory-scoped voicevox/.gitattributes for the dictionary's .dic/.bin rather than editing plan 01-02's root file"
55
+ - "voicevox_core wired into pyproject as a marker-differentiated 'voice' extra so uv sync cannot silently prune it"
56
+ - "AudioQuery serialised to the VOICEVOX ENGINE mixed-case JSON schema, since 0.17.0 exposes no .json() method"
57
+
58
+ patterns-established:
59
+ - "Every API name in a doc carries the file or URL it was read from, and was executed before being written down"
60
+ - "Fixture generators fail loudly rather than emit a fixture that does not exercise its target branch"
61
+ - "Miscounted or mis-shaped acceptance criteria are verified substantively and the discrepancy recorded, never satisfied by contorting code"
62
+
63
+ requirements-completed: [VOIC-01]
64
+ requirements-advanced: [AVTR-02, VOIC-03]
65
+
66
+ # Metrics
67
+ duration: 33min
68
+ completed: 2026-08-27
69
+ ---
70
+
71
+ # Phase 01 Plan 04: Japanese Voice Summary
72
+
73
+ **The server speaks Japanese: `synthesize()` returns 24000 Hz WAV bytes plus a plain-dict `AudioQuery` carrying per-mora `consonant_length`/`vowel_length`, loaded from LFS-committed local assets on zero GPU β€” and the three committed fixtures exposed a frame-quantisation error in `01-RESEARCH.md` that would have broken plan 01-06's slow re-read.**
74
+
75
+ ## Performance
76
+
77
+ - **Duration:** ~33 min
78
+ - **Started:** 2026-08-27T00:07Z
79
+ - **Completed:** 2026-08-27T00:40Z
80
+ - **Tasks:** 3 of 3
81
+ - **Files created/modified:** 26, verified by `git diff --name-only 5c6321a HEAD`
82
+
83
+ ## Task Commits
84
+
85
+ 1. **Task 1: Resolve the 0.17.0 API surface, acquire the runtime assets, write requirements.txt** β€” `54edb09` (feat)
86
+ 2. **Task 2: Implement synthesize() and the shared voice dataclasses** β€” `b0f00f6` (feat)
87
+ 3. **Task 3: Capture the ground-truth synthesis fixtures** β€” `122fc80` (test)
88
+ 4. *(follow-up)* **Record the measured cold-start cost of the lazy runtime fetch** β€” `84bdf64` (docs)
89
+
90
+ ## Everything the plan's `<output>` block required
91
+
92
+ ### Corrections to RESEARCH's inferred 0.17.0 API names
93
+
94
+ RESEARCH's construction chain was **entirely correct** β€” `Onnxruntime.load_once(filename=...)`,
95
+ `OpenJtalk(dict_dir)`, `Synthesizer(ort, ojt)`, `VoiceModelFile.open(path)`,
96
+ `load_voice_model(model)`, `create_audio_query(text, style_id)`, `synthesis(query, style_id)` all
97
+ exist and all work as written. Six things differed:
98
+
99
+ | # | RESEARCH said | Reality | Who it affects |
100
+ |---|---|---|---|
101
+ | 1 | `voicevox_vvm` latest is 0.16.4; confirm the mismatch | **0.17.0 exists** (2026-08-12), one day before core 0.17.0. No mismatch. | this plan |
102
+ | 2 | `json.loads(query.json())` | **`AudioQuery` has no `.json()`** (nor `.to_json`). It *is* a real `dataclass`, so `dataclasses.asdict()` is exact. | 01-06 |
103
+ | 3 | `pauseLength` / `pauseLengthScale` are AudioQuery fields | **They do not exist in CORE 0.17.0** (they are ENGINE-only). Pipeline step 3 is a no-op. | 01-06 |
104
+ | 4 | `SpeakerMeta` | **`CharacterMeta`** (the `.speaker_uuid` attribute is unchanged) | anyone importing it |
105
+ | 5 | Open JTalk dict Β© Nagoya Institute of Technology | **Nara** Institute of Science and Technology β€” see licence note below | 01-09 |
106
+ | 6 | `round(length / speedScale * 93.75)` | **Wrong.** See the next section β€” the most consequential finding in this plan. | 01-06 |
107
+
108
+ ### The frame-quantisation correction (please read this before plan 01-06)
109
+
110
+ RESEARCH's `build_timeline` divides each phoneme length by `speedScale` and *then* quantises.
111
+ VOICEVOX CORE 0.17.0 actually quantises **first**, at speed 1.0, then divides the resulting **frame
112
+ count** and rounds again:
113
+
114
+ ```python
115
+ frames = round(length / speed * 93.75) # 01-RESEARCH.md - WRONG
116
+ frames = round(round(length * 93.75) / speed) # CORRECT
117
+ ```
118
+
119
+ At `speedScale == 1.0` the two coincide exactly, which is why RESEARCH's version looked verified β€”
120
+ and it is exactly why this was worth catching here rather than in 01-06. Measured across 4
121
+ sentences Γ— 6 speed values (1.0, 0.9, 0.75, 0.5, 1.25, 1.5), predicted total frames against the
122
+ true frame count of the synthesised WAV:
123
+
124
+ | Formula | Correct |
125
+ |---|---|
126
+ | `round(length / speed * 93.75)` | **8 / 24** |
127
+ | `round(round(length * 93.75) / speed)` | **24 / 24** |
128
+
129
+ Worst observed error from the wrong form: **5 frames β‰ˆ 53 ms**, five times the Β±1 frame
130
+ (10.667 ms) tolerance `test_no_drift_long_utterance` is specified with. This is `01-RESEARCH.md`
131
+ Pitfall 5 ("normal speed syncs, slow speed drifts") arrived at from the other direction.
132
+
133
+ Two consequences for plan 01-06:
134
+
135
+ - Implement the corrected formula. `tests/fixtures/make_synth_fixtures.py` now asserts it
136
+ reproduces each fixture's true frame count before it will write anything, so the fixtures cannot
137
+ drift away from the engine unnoticed.
138
+ - **`test_speed_scale` must not assert "the timeline is exactly 1/0.75Γ— longer" against durations.**
139
+ Re-quantisation after scaling means the realised ratio lands *near* the requested one, not on it:
140
+ the committed `slow`/`long` ratio is **1.341085**, not 1.333333. Assert that the timeline matches
141
+ the WAV, which is the property lip-sync actually depends on.
142
+
143
+ Everything else in RESEARCH's pipeline is confirmed: pause mora after its accent phrase, pre/post
144
+ silence included in scaling, consonant before vowel, accumulate quantised frames not floats,
145
+ banker's rounding (Python's built-in `round` already is).
146
+
147
+ ### Voice model identity
148
+
149
+ | | |
150
+ |---|---|
151
+ | Original filename | **`0.vvm`** (from `voicevox_vvm` 0.17.0) |
152
+ | Repo path | `voicevox/model/zundamon.vvm`, 59,308,488 B, Git LFS |
153
+ | SHA256 | `ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318` |
154
+ | `speaker_uuid` | `388f246b-8c41-4ac1-8e2d-5d79f3ff56d9` |
155
+ | **`style_id` (γšγ‚“γ γ‚‚γ‚“ γƒŽγƒΌγƒžγƒ«)** | **3** |
156
+
157
+ Resolved from the release's own `README.txt` character↔style table and confirmed against the
158
+ `metas.json` inside the downloaded file. The LFS pointer's `oid sha256` matches the hash above
159
+ exactly, confirming the clean filter ran rather than the raw blob being stored.
160
+
161
+ ### Version pairing: yes, and it is not optional in the direction that matters
162
+
163
+ **voicevox_core 0.17.0 + voicevox_vvm 0.17.0 β€” verified by executing it.** `0.vvm` reports
164
+ `vvm_format_version: 2` in its internal `manifest.json` and every style has `type:
165
+ "streaming_talk"`; both the format version and the `StyleType::StreamingTalk` variant were
166
+ introduced *in core 0.17.0*. So core 0.17.0 reads vvm 0.16.4 fine, but core 0.16.x could not read
167
+ vvm 0.17.0. Both are pinned to 0.17.0.
168
+
169
+ ### ONNX Runtime acquisition: option **(c)**
170
+
171
+ - **(a) pip-installable companion β€” does not exist.** `pypi.org/pypi/voicevox-onnxruntime/json`,
172
+ `voicevox_onnxruntime` and `voicevox-core` all return **HTTP 404**. The runtime ships only as
173
+ `.tgz`/`.zip` archives from `VOICEVOX/onnxruntime-builder`. Nothing could go in `requirements.txt`.
174
+ - **(b) downloader binary at build time β€” no hook exists.** A Gradio-SDK Space's build phase runs
175
+ pip against `requirements.txt` and apt against `packages.txt`; neither can execute `./download`.
176
+ Option (b) is documented as the local-development path.
177
+ - **(c) chosen and exercised.** `get_synthesizer()` resolves `VOICEVOX_ORT_PATH` β†’ an existing copy
178
+ under `voicevox_runtime/` β†’ download the official archive β†’ fall back to `load_once()`'s own
179
+ search path. Pointing `VOICEVOX_ORT_DIR` at an empty directory and calling `warmup()` fetched,
180
+ unpacked, loaded and synthesised successfully. The runtime is always referenced from its official
181
+ release URL and never copied into the repo, per the software terms' 禁歒事項.
182
+
183
+ ### Measured warm-up
184
+
185
+ | Stage | Time |
186
+ |---|---|
187
+ | `import japanese_avatar.voice.tts` | 0.110 s (no file or socket touched) |
188
+ | `Onnxruntime.load_once(...)` | 0.093 s |
189
+ | `OpenJtalk(dict_dir)` | 0.001 s |
190
+ | `VoiceModelFile.open` + `load_voice_model` | 1.742 s |
191
+ | **`warmup()` total, runtime present** | **1.73 s** |
192
+ | **`warmup()` total, empty runtime cache (includes download)** | **2.66 s** |
193
+ | Second `warmup()` (cache hit) | 0.000001 s |
194
+
195
+ Per turn once warm: `create_audio_query` **1.2 ms**, `synthesis` **1215 ms** for γ€Œγ“γ‚“γ«γ‘γ―γ€. The
196
+ cold-start penalty this design carries is therefore about **0.9 s** β€” the 8.2 MB runtime fetch. The
197
+ ~166 MB of dictionary and voice model that would otherwise dominate every 48 h wake arrive with the
198
+ clone via LFS instead.
199
+
200
+ ### The `long` fixture
201
+
202
+ **Text:** γ€Œδ»Šζ—₯γ―γ„γ„ε€©ζ°—γ§γ™γ‹γ‚‰γ€ε…¬εœ’γ‚’ζ•£ζ­©γ—γ¦γ‹γ‚‰γ€θ²·γ„η‰©γ«θ‘ŒγγΎγ—γŸγ€‚γ€ β€” the plan's suggested
203
+ sentence, kept because it passed every assertion on the first try.
204
+
205
+ | Property | Value |
206
+ |---|---|
207
+ | Mora count (incl. pause moras) | **36** (required β‰₯ 20) |
208
+ | Phoneme count | 61 |
209
+ | Vowel symbol set | `['I', 'N', 'U', 'a', 'e', 'i', 'o', 'pau']` |
210
+ | **Devoiced vowels present** | **`I`, `U`** β€” from ですから and してから / γ—γΎγ—γŸ |
211
+ | `pause_mora` count | **2** |
212
+ | Frames / duration | 516 / **5.504 s** |
213
+
214
+ All three fixture cases, for the record:
215
+
216
+ | Case | speedScale | Moras | Frames | Duration (s) |
217
+ |---|---|---|---|---|
218
+ | `short` (こんにけは) | 1.0 | 5 | 99 | 1.056 |
219
+ | `long` | 1.0 | 36 | 516 | 5.504 |
220
+ | `slow` (same text) | 0.75 | 36 | 692 | 7.381333333333333 |
221
+
222
+ Regeneration is **byte-identical**: SHA-256 of all seven fixture files is unchanged across a second
223
+ `make_synth_fixtures.py` run.
224
+
225
+ ## Accomplishments
226
+
227
+ - **VOIC-01 is genuinely satisfied.** `synthesize("こんにけは")` returns a 50,732-byte RIFF WAV,
228
+ 24000 Hz / mono / 16-bit, 1.056 s, with a JSON-round-trippable `AudioQuery` carrying per-mora
229
+ `consonant_length` and `vowel_length`. `tests/test_tts_contract.py` is **7 passed** against a real
230
+ installed `voicevox_core` β€” quoted verbatim below.
231
+ - **Zero GPU on the synthesis path, proven by AST scan, not by grep.**
232
+ `test_no_gpu_imports_on_synthesis_path` parses every module under
233
+ `src/japanese_avatar/voice/` and rejects any import of the ZeroGPU allocation package or `torch`.
234
+ - **The licensing branch is closed and written down.** Every fact plan 01-09's `LICENSES.md` needs β€”
235
+ three terms URLs, the exact `VOICEVOX:γšγ‚“γ γ‚‚γ‚“` credit string, the placement standard, the
236
+ flow-down obligation, the AudioQuery-triggers-credit Q&A, the Nemo pre-cleared swap with its style
237
+ IDs, and the BSD-3-Clause dictionary notice β€” lives in `docs/VOICEVOX-SETUP.md`. No "email the
238
+ rights holder" task was created; SSS LLC 免責村項 2 is quoted so it is never re-litigated.
239
+ - **Cold start retired ahead of time.** 166 MB of assets arrive via LFS; only 8.2 MB is fetched, and
240
+ only inside a lazily-built synthesizer that `warmup()` pre-pays at startup.
241
+ - **Quick loop kept inside its budget** at ~11 s against the 15 s Nyquist limit, after removing six
242
+ redundant syntheses.
243
+
244
+ ### Verbatim test output required by the plan
245
+
246
+ ```
247
+ $ uv run pytest tests/test_tts_contract.py -q
248
+ ....... [100%]
249
+ ```
250
+
251
+ Seven tests: the six the plan specifies plus `test_fixtures_match_current_engine` added in Task 3.
252
+ Full quick loop (`pytest tests/ -q --ignore=tests/e2e`): **13 passed**, timed at 10,890 / 11,633 /
253
+ 11,352 ms over three runs.
254
+
255
+ ## Decisions Made
256
+
257
+ - **`voicevox/.gitattributes` instead of editing the root file.** `sys.dic` is 103 MB and the root
258
+ `.gitattributes` (plan 01-02) arms only `.vrm/.vvm/.wav/.onnx`. Git applies attributes
259
+ hierarchically, so a directory-scoped file adds `*.dic`/`*.bin` coverage without competing for
260
+ ownership of 01-02's file. Verified with `git check-attr filter`, which flipped from
261
+ `unspecified` to `lfs` for all four dictionary binaries.
262
+ - **`voicevox_core` is a `pyproject.toml` extra, not a manual install.** Left out of the project
263
+ metadata, the next `uv sync` would prune it and `tests/test_tts_contract.py` would silently
264
+ degrade to "skipped" for every later plan in the phase. It is wired as the `voice` extra with
265
+ marker-differentiated `[tool.uv.sources]`, so one lockfile serves Windows dev and the Linux Space
266
+ builder. `requirements.txt` remains the Space's independent contract and carries only the
267
+ manylinux wheel.
268
+ - **ENGINE mixed-case JSON schema for the serialised query.** `speedScale` camel, `accent_phrases`
269
+ and all mora keys snake. Not a style choice β€” it is the wire format the plan's own acceptance
270
+ criteria, RESEARCH's schema block and every VOICEVOX consumer expect.
271
+ - **All three fixture WAVs committed, not just the short one.** `speech_ja_long.wav` and
272
+ `speech_ja_slow.wav` are 264 KB and 354 KB; having the real audio next to the query lets 01-06
273
+ assert against the signal rather than only against a recorded number.
274
+
275
+ ## Deviations from Plan
276
+
277
+ ### Auto-fixed Issues
278
+
279
+ **1. [Rule 1 - Bug] `01-RESEARCH.md`'s frame-quantisation formula is wrong for `speedScale != 1.0`**
280
+
281
+ - **Found during:** Task 3, when the `slow` fixture's rebuilt timeline missed its WAV by exactly one
282
+ frame while `long` matched to 2.7e-15 s.
283
+ - **Expected:** the plan's fixture table asserts the slow timeline is "exactly `1/0.75x` longer".
284
+ - **Found:** `round(length / speed * 93.75)` is right 8/24 across 6 speeds Γ— 4 sentences;
285
+ `round(round(length * 93.75) / speed)` is right 24/24. Worst error 5 frames β‰ˆ 53 ms.
286
+ - **Fix:** documented under a dedicated "Frame quantisation" section in `docs/VOICEVOX-SETUP.md`
287
+ with the measurement table, and enforced in `make_synth_fixtures.py`, which asserts the correct
288
+ formula reproduces each fixture's true frame count before writing. `visemes.py` itself is plan
289
+ 01-06's to write β€” this plan corrects the specification it will be written from, and does not
290
+ pre-empt it.
291
+ - **Files modified:** `docs/VOICEVOX-SETUP.md`, `tests/fixtures/make_synth_fixtures.py`
292
+ - **Committed in:** `122fc80`
293
+
294
+ **2. [Rule 3 - Blocking] The Open JTalk dictionary had no LFS coverage**
295
+
296
+ - **Issue:** `sys.dic` is 103,073,776 B. `git check-attr filter` returned `unspecified` for every
297
+ `.dic`/`.bin` file, so they would have been committed as raw blobs β€” rejected outright by a Hugging
298
+ Face Space repo, which requires LFS above 10 MB.
299
+ - **Constraint tension:** the executor brief forbids modifying the root `.gitattributes` (owned by
300
+ 01-02) while also requiring the dictionary to land through LFS.
301
+ - **Fix:** a new `voicevox/.gitattributes` scoped to the directory this plan owns. Root file
302
+ untouched β€” confirmed by `git diff --cached --name-only | grep -cx '.gitattributes'` β†’ `0`.
303
+ - **Verification:** all four dictionary binaries flipped to `filter: lfs`; `git lfs status` shows
304
+ `sys.dic`, `char.bin`, `matrix.bin`, `unk.dic` and `zundamon.vvm` as LFS objects.
305
+ - **Committed in:** `54edb09`
306
+
307
+ **3. [Rule 2 - Missing Critical] `voicevox_core` would have been pruned by the next `uv sync`**
308
+
309
+ - **Issue:** the plan says to install the Windows wheel "separately". Anything not in the lockfile is
310
+ removed by `uv sync`, which would have turned the VOIC-01 suite into a silent skip for plans
311
+ 01-06/07/08.
312
+ - **Fix:** `voice` extra in `pyproject.toml` with marker-differentiated `[tool.uv.sources]`. First
313
+ attempt failed universal resolution for `sys_platform == 'emscripten'`; corrected by putting a
314
+ disjunctive marker on the requirement itself.
315
+ - **Files modified:** `pyproject.toml`, `uv.lock`
316
+ - **Committed in:** `54edb09`
317
+
318
+ **4. [Rule 1 - Bug] Quick loop had grown to ~20 s against a 15 s budget**
319
+
320
+ - **Issue:** the test module as specified performs six syntheses of which three are exact duplicates
321
+ (`LONG_TEXT` at speed 1.0 three times, at 0.75 twice). `01-VALIDATION.md` caps quick-loop feedback
322
+ latency at ~15 s.
323
+ - **Fix:** a module-scoped `synth` fixture memoising by `(text, speed)`. Three distinct syntheses now
324
+ cover every test. No assertion was weakened or removed.
325
+ - **Verification:** 10,890 / 11,633 / 11,352 ms over three timed runs, 13 passed.
326
+ - **Committed in:** `122fc80`
327
+
328
+ **5. [Rule 1 - Bug] Fixture JSON was written with CRLF on Windows**
329
+
330
+ - **Issue:** `Path.write_text` newline-translates by default, so the committed fixtures were CRLF and
331
+ git warned it would normalise them. A Linux regeneration would then produce different bytes,
332
+ destroying the byte-reproducibility property the fixtures depend on.
333
+ - **Fix:** explicit `newline="\n"` on both writes. Confirmed 0 CRLF bytes in all four JSON files, and
334
+ regeneration is byte-identical.
335
+ - **Committed in:** `122fc80`
336
+
337
+ **6. [Rule 1 - Doc bug] Open JTalk dictionary attribution named the wrong institution**
338
+
339
+ - **Expected:** RESEARCH and the plan's acceptance criterion both say "Nagoya Institute of
340
+ Technology".
341
+ - **Found:** the shipped `COPYING` reads *"Copyright (c) 2009, **Nara** Institute of Science and
342
+ Technology, Japan"*. Both institutions are genuinely involved β€” Nara (NAIST) holds the dictionary
343
+ data; Nagoya Institute of Technology and the HTS Working Group hold Open JTalk itself.
344
+ - **Fix:** `docs/VOICEVOX-SETUP.md` quotes the NAIST notice verbatim, states both attributions, and
345
+ flags that `LICENSES.md` must carry both β€” reproducing only Nagoya would fail the BSD notice on the
346
+ files actually redistributed here. The acceptance criterion's literal string is still present.
347
+ - **Committed in:** `54edb09`
348
+
349
+ ### Acceptance criteria restated rather than skipped
350
+
351
+ Two criteria are mis-shaped in the same way `01-02` documented. In both cases the substantive
352
+ property was verified directly and no code was contorted.
353
+
354
+ - **`grep -rc "import spaces\|@spaces" src/japanese_avatar/voice/` returns 0.** It initially returned
355
+ `1` β€” matching my own module docstring, which named the prohibition literally
356
+ (``@spaces.GPU``). The docstring was reworded so the grep is literally true as well as
357
+ substantively true; it now returns 0 for every file. The real guarantee is the AST scan in
358
+ `test_no_gpu_imports_on_synthesis_path`, which a grep cannot provide.
359
+ - **`cases.*.duration_seconds` are "full-precision floats … (more than 3 decimal places)".** `slow`
360
+ is `7.381333333333333`, but `short` is `1.056` and `long` is `5.504` β€” exactly 3 decimal places.
361
+ These *are* the unrounded header values; VOICEVOX frame counts happen to be divisible by 3 here, so
362
+ `nΒ·256/24000` terminates at 3 dp. The substantive property β€” that the recorded value equals
363
+ `getnframes()/getframerate()` exactly, with no rounding applied β€” was asserted directly and holds
364
+ for all three cases (`exact=True`). The generator additionally asserts the written WAV's header
365
+ duration equals what `synthesize()` reported, bit for bit.
366
+
367
+ ---
368
+
369
+ **Total deviations:** 6 auto-fixed (3 bugs, 1 blocking, 1 missing-critical, 1 doc bug) + 2 acceptance
370
+ criteria restated
371
+ **Impact on plan:** No scope creep and no design improvisation. Deviations 2 and 3 were required for
372
+ the plan's own criteria to be satisfiable; 1 and 6 correct specifications that later plans would
373
+ otherwise have built on; 4 and 5 protect the phase's feedback loop and fixture reproducibility.
374
+
375
+ ## Issues Encountered
376
+
377
+ - **The Open JTalk mirror in the VOICEVOX guide is dead.** `jaist.dl.sourceforge.net` fails DNS
378
+ resolution; `downloads.sourceforge.net` served the same archive (23,643,819 B) without incident.
379
+ Recorded in the setup doc as the working URL.
380
+ - **A ruff lint/format standoff.** `SIM108` demanded a ternary for the ONNX Runtime branch while
381
+ `ruff format` insisted on collapsing that ternary onto a >100-char line. Resolved by shortening the
382
+ expression to `ort_kwargs = {"filename": ort_path} if ort_path else {}` β€” satisfying both and, as it
383
+ happens, reading better than either.
384
+ - One transient `.ruff_cache` "Access is denied" warning on Windows during a concurrent invocation,
385
+ as in plan 01-02. Cosmetic; the run reported "All checks passed!".
386
+
387
+ ## Requirement Status
388
+
389
+ **VOIC-01 β†’ Complete.** `01-VALIDATION.md` binds it to `tests/test_tts_contract.py`, which this plan
390
+ creates and which passes 7/7 against a real installed `voicevox_core`. The behaviour it names β€”
391
+ "`synthesize()` returns non-empty WAV + parseable `AudioQuery`; zero GPU on the call path" β€” is
392
+ demonstrated, not asserted. This is the one requirement in this plan's frontmatter and it is
393
+ genuinely satisfied.
394
+
395
+ **Deliberately NOT marked:** the plan's frontmatter lists only `VOIC-01`, so no over-claiming was
396
+ possible here. For the record, this plan *advances* but does not satisfy:
397
+
398
+ - **AVTR-02** β€” needs `tests/test_visemes.py`, which plan 01-06 creates. This plan supplies its
399
+ ground-truth fixtures and corrects its algorithm.
400
+ - **VOIC-03** β€” the `speedScale` mechanism works and is fixture-backed, but the requirement's tests
401
+ (`test_speed_scale`, and the deployed `test_slower`) belong to plans 01-06 and 01-10.
402
+
403
+ ## Constraint Compliance
404
+
405
+ - **No Space was created.** No `hf` CLI call, no `huggingface_hub` call, no Space README front-matter
406
+ written. Plan 01-01's human hosting gate remains uncleared and untouched.
407
+ - **No `git push` occurred β€” and could not have.** `git remote -v` is still empty; this repo has no
408
+ remote configured.
409
+ - **No money was spent.** No PRO subscription, no paid resource, no API with a meter.
410
+ - **Python is 3.12.12** throughout (`uv run python -V`).
411
+ - **`voicevox_core` 0.17.0**, wheel filename confirmed against the GitHub Releases API rather than
412
+ guessed. `requirements.txt` carries the **manylinux** wheel; the **win_amd64** wheel is confined to
413
+ the `pyproject.toml` `voice` extra behind a `sys_platform == 'win32'` marker.
414
+ - **File ownership honoured.** `docs/ASSETS.md`, `README.md` and the root `.gitattributes` are all
415
+ absent from this plan's diff β€” verified with `git diff --cached --name-only`.
416
+ - **Nothing vendored.** `find . -name "*.whl" -not -path "./.venv/*"` β†’ 0. No `.so`, `.dll` or ORT
417
+ binary is tracked; `voicevox_runtime/` was already gitignored by plan 01-02.
418
+ - **No GPU anywhere near synthesis.** AST-verified, and `grep -r "import spaces\|@spaces"` over
419
+ `src/japanese_avatar/voice/` returns 0 matches.
420
+ - **Zero AI tutoring in scope.** No LLM, no level gating, no grammar database, no accounts, no
421
+ database. The discarded TTS fallback branch is absent from `docs/` and `src/` entirely (0 matches),
422
+ deleted rather than deferred.
423
+ - Network egress was limited to what the plan's `<action>` blocks instruct: the GitHub Releases API,
424
+ the voice model, the Open JTalk dictionary, the ONNX Runtime archive, and the `voicevox_core` wheel.
425
+
426
+ ## Known Stubs
427
+
428
+ None. Every artifact this plan claims is wired to a real data source and exercised by a passing test.
429
+
430
+ `src/japanese_avatar/voice/models.py` defines `VisemeEvent` and `AvatarDirective`, which nothing
431
+ consumes *yet* β€” but they are interface definitions the plan explicitly assigns to this file for
432
+ plans 01-06 and 01-08 to consume, not placeholders. They contain no fabricated data and no
433
+ hardcoded empty values flowing to UI.
434
+
435
+ ## User Setup Required
436
+
437
+ None. `uv sync --extra dev --extra voice` is the only local step and it has already been run and
438
+ committed to the lockfile.
439
+
440
+ ## Next Phase Readiness
441
+
442
+ - **Plan 01-06** is unblocked and materially de-risked: three real `AudioQuery` fixtures with true
443
+ WAV durations *and* frame counts, a long case with 36 moras / 61 phonemes / two devoiced vowels /
444
+ two pause moras, the corrected quantisation formula with its evidence, and an explicit warning that
445
+ `pauseLength`/`pauseLengthScale` do not exist. **Read the "Frame quantisation" section of
446
+ `docs/VOICEVOX-SETUP.md` before writing `visemes.py`.**
447
+ - **Plan 01-05** has `requirements.txt`. Note that the Space also needs `VOICEVOX_ORT_*` to be left at
448
+ defaults so the lazy fetch works, and that `git lfs pull` must succeed on the Space or the assets
449
+ arrive as 130-byte pointers β€” `tts.py` raises a message naming that case explicitly.
450
+ - **Plan 01-07** has `tests/fixtures/speech_ja.wav` (24000 Hz mono 16-bit), the file
451
+ `tests/conftest.py`'s `speech_wav` fixture has been skipping on since 01-02.
452
+ - **Plan 01-08** has `TurnTimings` and the `warmup()` hook to call at startup, plus real numbers to
453
+ budget against: 1.73 s warm-up, 1.2 ms query, ~1.2 s synthesis.
454
+ - **Plan 01-09** has every voice-side licence fact `LICENSES.md` needs, including the dual
455
+ Nara/Nagoya attribution correction.
456
+
457
+ Carried forward, unchanged by this plan:
458
+
459
+ - The hosting model is still unresolved (ZeroGPU eligibility vs PRO) β€” plan 01-01's human gate.
460
+ - Per-turn GPU-second cost remains unmeasured (and remains zero for speech output, by construction).
461
+ - Resolved: "VOICEVOX character terms still need a human read" is now **closed** β€” all three licence
462
+ layers were read in full and recorded, and the rights holder declines qualification questions as a
463
+ matter of policy, so there is no further human step available or required.
464
+
465
+ ## Self-Check: PASSED
466
+
467
+ All 16 files claimed created exist on disk; all 2 claimed modified exist. All 4 commit hashes
468
+ (`54edb09`, `b0f00f6`, `122fc80`, `84bdf64`) resolve in `git log`. File count stated as **26** against
469
+ `git diff --name-only 5c6321a HEAD`, not estimated.
470
+
471
+ Full plan verification block re-run after the final commit:
472
+
473
+ | Check | Result |
474
+ |---|---|
475
+ | `pytest tests/test_tts_contract.py -q` | **7 passed** |
476
+ | `pytest tests/ -q --ignore=tests/e2e` | **13 passed**, ~11.0 s |
477
+ | `python -c "…synthesize('こんにけは').duration"` | `1.056` |
478
+ | `git lfs ls-files` | lists `voicevox/model/zundamon.vvm`, `tests/fixtures/speech_ja.wav` (+ 8 more) |
479
+ | `grep -rc "import spaces\|@spaces" src/japanese_avatar/voice/` | **0** for every file |
480
+ | `find . -name "*.whl" -not -path "./.venv/*"` | empty |
481
+ | `ruff check . && ruff format --check .` | "All checks passed!" / "13 files already formatted" |
482
+ | `requirements.txt` non-empty lines | **4**, one `.whl`, zero vendored |
483
+ | `git status --short` | clean β€” no untracked files |
484
+ | `git remote -v` | empty β€” no push was possible |
485
+
486
+ ---
487
+ *Phase: 01-voice-avatar-loop-skeleton*
488
+ *Completed: 2026-08-27*