Spaces:
Running on Zero
Running on Zero
docs(01-04): complete Japanese voice plan
Browse files
.planning/REQUIREMENTS.md
CHANGED
|
@@ -13,7 +13,7 @@
|
|
| 13 |
|
| 14 |
### Voice Loop
|
| 15 |
|
| 16 |
-
- [
|
| 17 |
- [ ] **VOIC-02**: Learner can speak Japanese via push-to-talk; ASR transcribes with visible transcript
|
| 18 |
- [ ] **VOIC-03**: Learner can replay the avatar's last utterance and request it slower
|
| 19 |
- [ ] **VOIC-04**: Learner can type instead of speak at any point (full text fallback)
|
|
@@ -102,7 +102,7 @@ Which phases cover which requirements. Updated during roadmap creation.
|
|
| 102 |
| AVTR-01 | Phase 1 | Pending |
|
| 103 |
| AVTR-02 | Phase 1 | Pending |
|
| 104 |
| AVTR-03 | Phase 3 | Pending |
|
| 105 |
-
| VOIC-01 | Phase 1 |
|
| 106 |
| VOIC-02 | Phase 1 | Pending |
|
| 107 |
| VOIC-03 | Phase 1 | Pending |
|
| 108 |
| VOIC-04 | Phase 1 | Pending |
|
|
|
|
| 13 |
|
| 14 |
### Voice Loop
|
| 15 |
|
| 16 |
+
- [x] **VOIC-01**: Avatar speaks Japanese aloud via TTS that emits per-mora timings (CPU-path, zero GPU quota)
|
| 17 |
- [ ] **VOIC-02**: Learner can speak Japanese via push-to-talk; ASR transcribes with visible transcript
|
| 18 |
- [ ] **VOIC-03**: Learner can replay the avatar's last utterance and request it slower
|
| 19 |
- [ ] **VOIC-04**: Learner can type instead of speak at any point (full text fallback)
|
|
|
|
| 102 |
| AVTR-01 | Phase 1 | Pending |
|
| 103 |
| AVTR-02 | Phase 1 | Pending |
|
| 104 |
| AVTR-03 | Phase 3 | Pending |
|
| 105 |
+
| VOIC-01 | Phase 1 | Complete |
|
| 106 |
| VOIC-02 | Phase 1 | Pending |
|
| 107 |
| VOIC-03 | Phase 1 | Pending |
|
| 108 |
| VOIC-04 | Phase 1 | Pending |
|
.planning/ROADMAP.md
CHANGED
|
@@ -34,7 +34,7 @@ Decimal phases appear between their surrounding integers in numeric order.
|
|
| 34 |
**Plans**: 10 plans in 7 waves
|
| 35 |
- [x] 01-02-PLAN.md β Toolchain, LFS arming, package skeleton, test scaffolding (local-only) [wave 1]
|
| 36 |
- [ ] 01-01-PLAN.md β Hosting decision + Space creation + VRM sourcing (user-gated) [wave 2]
|
| 37 |
-
- [
|
| 38 |
- [ ] 01-03-PLAN.md β Avatar stage, shared facade + turn loop, both transports [wave 3]
|
| 39 |
- [ ] 01-05-PLAN.md β Space manifest, deploy the spike, first deployed E2E, verdict (user-gated) [wave 4]
|
| 40 |
- [ ] 01-06-PLAN.md β Mora-to-viseme timeline builder, test-first [wave 4]
|
|
@@ -109,7 +109,7 @@ Phases execute in numeric order: 1 β 2 β 3 β 4 β 5 β 6
|
|
| 109 |
|
| 110 |
| Phase | Plans Complete | Status | Completed |
|
| 111 |
|-------|----------------|--------|-----------|
|
| 112 |
-
| 1. Voice + Avatar Loop Skeleton |
|
| 113 |
| 2. Japanese Language Core | 0/TBD | Not started | - |
|
| 114 |
| 3. Tutoring Brain | 0/TBD | Not started | - |
|
| 115 |
| 4. Accounts & Persistence | 0/TBD | Not started | - |
|
|
|
|
| 34 |
**Plans**: 10 plans in 7 waves
|
| 35 |
- [x] 01-02-PLAN.md β Toolchain, LFS arming, package skeleton, test scaffolding (local-only) [wave 1]
|
| 36 |
- [ ] 01-01-PLAN.md β Hosting decision + Space creation + VRM sourcing (user-gated) [wave 2]
|
| 37 |
+
- [x] 01-04-PLAN.md β VOICEVOX TTS, requirements.txt, ground-truth AudioQuery fixtures [wave 2]
|
| 38 |
- [ ] 01-03-PLAN.md β Avatar stage, shared facade + turn loop, both transports [wave 3]
|
| 39 |
- [ ] 01-05-PLAN.md β Space manifest, deploy the spike, first deployed E2E, verdict (user-gated) [wave 4]
|
| 40 |
- [ ] 01-06-PLAN.md β Mora-to-viseme timeline builder, test-first [wave 4]
|
|
|
|
| 109 |
|
| 110 |
| Phase | Plans Complete | Status | Completed |
|
| 111 |
|-------|----------------|--------|-----------|
|
| 112 |
+
| 1. Voice + Avatar Loop Skeleton | 2/10 | In Progress | - |
|
| 113 |
| 2. Japanese Language Core | 0/TBD | Not started | - |
|
| 114 |
| 3. Tutoring Brain | 0/TBD | Not started | - |
|
| 115 |
| 4. Accounts & Persistence | 0/TBD | Not started | - |
|
.planning/STATE.md
CHANGED
|
@@ -3,13 +3,13 @@ gsd_state_version: 1.0
|
|
| 3 |
milestone: v1.0
|
| 4 |
milestone_name: milestone
|
| 5 |
status: unknown
|
| 6 |
-
stopped_at: Completed 01-
|
| 7 |
-
last_updated: "2026-08-27T00:
|
| 8 |
progress:
|
| 9 |
total_phases: 6
|
| 10 |
completed_phases: 0
|
| 11 |
total_plans: 10
|
| 12 |
-
completed_plans:
|
| 13 |
---
|
| 14 |
|
| 15 |
# Project State
|
|
@@ -24,36 +24,39 @@ See: .planning/PROJECT.md (updated 2026-08-08)
|
|
| 24 |
## Current Position
|
| 25 |
|
| 26 |
Phase: 01 (voice-avatar-loop-skeleton) β EXECUTING
|
| 27 |
-
Plan:
|
| 28 |
|
| 29 |
-
**Completed:** 01-02 (local foundations β pytest/ruff on Python 3.12.12, Git LFS armed, `--space-url` plumbing, audio fixtures).
|
| 30 |
-
**Not started:** 01-01 β it is deliberately out of order. 01-02
|
| 31 |
-
**Next actionable:** 01-03 (avatar transport seam)
|
|
|
|
|
|
|
| 32 |
|
| 33 |
## Performance Metrics
|
| 34 |
|
| 35 |
**Velocity:**
|
| 36 |
|
| 37 |
-
- Total plans completed:
|
| 38 |
-
- Average duration:
|
| 39 |
-
- Total execution time: 0.
|
| 40 |
|
| 41 |
**By Phase:**
|
| 42 |
|
| 43 |
| Phase | Plans | Total | Avg/Plan |
|
| 44 |
|-------|-------|-------|----------|
|
| 45 |
-
| 01 |
|
| 46 |
|
| 47 |
**Per-plan detail:**
|
| 48 |
|
| 49 |
| Plan | Duration | Tasks | Files |
|
| 50 |
|------|----------|-------|-------|
|
| 51 |
| Phase 01 P02 | 21min | 3 tasks | 16 files |
|
|
|
|
| 52 |
|
| 53 |
**Recent Trend:**
|
| 54 |
|
| 55 |
-
- Last 5 plans: 01-02 (21min)
|
| 56 |
-
- Trend:
|
| 57 |
|
| 58 |
*Updated after each plan completion*
|
| 59 |
|
|
@@ -71,6 +74,9 @@ Recent decisions affecting current work:
|
|
| 71 |
- [Phase 01]: Ruff excludes *.md: ruff 0.16 formats Python blocks embedded in Markdown, which made the phase's own planning docs fail the lint gate. Lint governs shipped Python, not prose.
|
| 72 |
- [Phase 01]: uv.lock is committed so a Space rebuild resolves the same dependency set that passed locally; requirements.txt (plan 01-04) remains the Space's contract.
|
| 73 |
- [Phase 01]: Git LFS armed for .vrm/.vvm/.wav/.onnx in Wave 1, before any binary exists - retires the 10 MiB non-LFS push rejection ahead of plans 01-01 and 01-04.
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
### Pending Todos
|
| 76 |
|
|
@@ -81,13 +87,14 @@ None yet.
|
|
| 81 |
### Blockers/Concerns
|
| 82 |
|
| 83 |
- **Hosting model unresolved (pre-Phase-1 decision).** STACK research found free Gradio Spaces on personal accounts now require PRO; the free path is ZeroGPU with per-visitor quota (~2 min/day anonymous). Confirm WolfDavid account eligibility (verified email, >30 days old, <2 existing ZeroGPU Spaces) during Phase 1 planning; if ineligible, PRO ($9/mo) becomes a decision before any build.
|
| 84 |
-
- **Asset licensing
|
| 85 |
- **Per-turn GPU-second cost is unmeasured.** The entire free-tier UX budget depends on this number β measure in Phase 1/3, do not assume.
|
| 86 |
- **REQUIREMENTS.md originally miscounted v1 as 31.** Actual count is 33; traceability and coverage corrected during roadmap creation.
|
| 87 |
- **Phase 1 plan frontmatter over-claims shared requirements β do not blind-run `requirements mark-complete`.** DPLY-01 appears in the `requirements:` field of plans 01-01, 01-02, 01-05, 01-09 and 01-10, but it is only *satisfied* by 01-05 (`test_space_reachable`). Executing 01-02 marked it Complete; this was reverted, and DPLY-01 is correctly `Pending`. The same over-claiming applies to AVTR-01/02, VOIC-02/03/04/05 and DPLY-04 across plans 01-09 and 01-10. Verify a requirement's actual acceptance test has run before checking it off.
|
|
|
|
| 88 |
|
| 89 |
## Session Continuity
|
| 90 |
|
| 91 |
-
Last session: 2026-08-27T00:
|
| 92 |
-
Stopped at: Completed 01-
|
| 93 |
Resume file: None
|
|
|
|
| 3 |
milestone: v1.0
|
| 4 |
milestone_name: milestone
|
| 5 |
status: unknown
|
| 6 |
+
stopped_at: Completed 01-04-PLAN.md (Japanese voice). VOIC-01 satisfied. Plan 01-01 still gated on the human hosting decision.
|
| 7 |
+
last_updated: "2026-08-27T00:44:59.201Z"
|
| 8 |
progress:
|
| 9 |
total_phases: 6
|
| 10 |
completed_phases: 0
|
| 11 |
total_plans: 10
|
| 12 |
+
completed_plans: 2
|
| 13 |
---
|
| 14 |
|
| 15 |
# Project State
|
|
|
|
| 24 |
## Current Position
|
| 25 |
|
| 26 |
Phase: 01 (voice-avatar-loop-skeleton) β EXECUTING
|
| 27 |
+
Plan: 3 of 10 (2 complete)
|
| 28 |
|
| 29 |
+
**Completed:** 01-02 (local foundations β pytest/ruff on Python 3.12.12, Git LFS armed, `--space-url` plumbing, audio fixtures) and 01-04 (Japanese voice β VOICEVOX assets via LFS, `synthesize()`, `requirements.txt`, ground-truth fixtures). **VOIC-01 is Complete.**
|
| 30 |
+
**Not started:** 01-01 β it is deliberately out of order. 01-02 and 01-04 are both local-only precisely so they could run in parallel with 01-01's **human hosting checkpoint, which has NOT cleared**. The plan counter is a count, not a plan ID: "Plan 3 of 10" does not mean 01-01 or 01-03 is done.
|
| 31 |
+
**Next actionable:** 01-03 (avatar transport seam) and 01-06 (viseme timeline) are both unblocked β 01-06's dependency was 01-04's fixtures, which now exist. 01-05 has its `requirements.txt` but still needs the hosting decision. 01-01 remains gated on the human checkpoint.
|
| 32 |
+
|
| 33 |
+
> **Before writing `visemes.py` in 01-06, read `docs/VOICEVOX-SETUP.md` Β§ "Frame quantisation".** 01-RESEARCH.md's formula is wrong for `speedScale != 1.0`; the corrected form and its evidence are recorded there and in the Blockers/Concerns section below.
|
| 34 |
|
| 35 |
## Performance Metrics
|
| 36 |
|
| 37 |
**Velocity:**
|
| 38 |
|
| 39 |
+
- Total plans completed: 2
|
| 40 |
+
- Average duration: 27 min
|
| 41 |
+
- Total execution time: 0.9 hours
|
| 42 |
|
| 43 |
**By Phase:**
|
| 44 |
|
| 45 |
| Phase | Plans | Total | Avg/Plan |
|
| 46 |
|-------|-------|-------|----------|
|
| 47 |
+
| 01 | 2 | 54min | 27min |
|
| 48 |
|
| 49 |
**Per-plan detail:**
|
| 50 |
|
| 51 |
| Plan | Duration | Tasks | Files |
|
| 52 |
|------|----------|-------|-------|
|
| 53 |
| Phase 01 P02 | 21min | 3 tasks | 16 files |
|
| 54 |
+
| Phase 01 P04 | 33min | 3 tasks | 26 files |
|
| 55 |
|
| 56 |
**Recent Trend:**
|
| 57 |
|
| 58 |
+
- Last 5 plans: 01-02 (21min), 01-04 (33min)
|
| 59 |
+
- Trend: slower, but 01-04 downloaded ~190 MB of assets and ran an empirical investigation the plan did not budget for
|
| 60 |
|
| 61 |
*Updated after each plan completion*
|
| 62 |
|
|
|
|
| 74 |
- [Phase 01]: Ruff excludes *.md: ruff 0.16 formats Python blocks embedded in Markdown, which made the phase's own planning docs fail the lint gate. Lint governs shipped Python, not prose.
|
| 75 |
- [Phase 01]: uv.lock is committed so a Space rebuild resolves the same dependency set that passed locally; requirements.txt (plan 01-04) remains the Space's contract.
|
| 76 |
- [Phase 01]: Git LFS armed for .vrm/.vvm/.wav/.onnx in Wave 1, before any binary exists - retires the 10 MiB non-LFS push rejection ahead of plans 01-01 and 01-04.
|
| 77 |
+
- [Phase 01]: voicevox_vvm 0.17.0 (not 0.16.4) pairs with core 0.17.0; the vvm ships format v2 + StyleType::StreamingTalk, both introduced in core 0.17.0
|
| 78 |
+
- [Phase 01]: ONNX Runtime uses lazy first-use download (option c): no pip distribution exists and a Gradio-SDK Space has no build hook to run the official downloader
|
| 79 |
+
- [Phase 01]: AudioQuery serialised to a plain ENGINE-schema dict at the module boundary; voicevox_core 0.17.0 exposes no .json() method, but the type is a real dataclass
|
| 80 |
|
| 81 |
### Pending Todos
|
| 82 |
|
|
|
|
| 87 |
### Blockers/Concerns
|
| 88 |
|
| 89 |
- **Hosting model unresolved (pre-Phase-1 decision).** STACK research found free Gradio Spaces on personal accounts now require PRO; the free path is ZeroGPU with per-visitor quota (~2 min/day anonymous). Confirm WolfDavid account eligibility (verified email, >30 days old, <2 existing ZeroGPU Spaces) during Phase 1 planning; if ineligible, PRO ($9/mo) becomes a decision before any build.
|
| 90 |
+
- **Asset licensing β VOICEVOX half RESOLVED (01-04), VRM half still open.** The VOICEVOX branch is closed: all three licence layers were read in full, third-party web synthesis is explicitly contemplated (software cl.3 attaches a flow-down obligation rather than prohibiting it), embedded `.vvm` redistribution is permitted (model cl.2), and the rights holder declines qualification questions as policy (SSS LLC ε
責ζ‘ι
2) β so there is no human step available or required. Compliance is a visible `VOICEVOX:γγγ γγ` credit plus a terms notice wherever audio is obtainable; all facts recorded in `docs/VOICEVOX-SETUP.md`. **Still open:** VRM model licences vary per author; a self-authored VRoid Studio character removes that question (plan 01-01).
|
| 91 |
- **Per-turn GPU-second cost is unmeasured.** The entire free-tier UX budget depends on this number β measure in Phase 1/3, do not assume.
|
| 92 |
- **REQUIREMENTS.md originally miscounted v1 as 31.** Actual count is 33; traceability and coverage corrected during roadmap creation.
|
| 93 |
- **Phase 1 plan frontmatter over-claims shared requirements β do not blind-run `requirements mark-complete`.** DPLY-01 appears in the `requirements:` field of plans 01-01, 01-02, 01-05, 01-09 and 01-10, but it is only *satisfied* by 01-05 (`test_space_reachable`). Executing 01-02 marked it Complete; this was reverted, and DPLY-01 is correctly `Pending`. The same over-claiming applies to AVTR-01/02, VOIC-02/03/04/05 and DPLY-04 across plans 01-09 and 01-10. Verify a requirement's actual acceptance test has run before checking it off.
|
| 94 |
+
- CORRECTION for plan 01-06 (not a blocker, a spec fix): 01-RESEARCH.md's frame quantisation is wrong for speedScale != 1.0. Correct form is round(round(length*93.75)/speed), NOT round(length/speed*93.75) - verified 24/24 vs 8/24 across 6 speeds x 4 sentences, worst error 5 frames (53ms). Also: pauseLength/pauseLengthScale do not exist in voicevox_core 0.17.0, and the slow/long duration ratio is 1.341085, not exactly 1/0.75. See docs/VOICEVOX-SETUP.md section 'Frame quantisation'.
|
| 95 |
|
| 96 |
## Session Continuity
|
| 97 |
|
| 98 |
+
Last session: 2026-08-27T00:43:56.253Z
|
| 99 |
+
Stopped at: Completed 01-04-PLAN.md (Japanese voice). VOIC-01 satisfied. Plan 01-01 still gated on the human hosting decision.
|
| 100 |
Resume file: None
|
.planning/phases/01-voice-avatar-loop-skeleton/01-04-SUMMARY.md
ADDED
|
@@ -0,0 +1,488 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
phase: 01-voice-avatar-loop-skeleton
|
| 3 |
+
plan: 04
|
| 4 |
+
subsystem: voice
|
| 5 |
+
tags: [voicevox, tts, japanese, mora-timing, git-lfs, onnxruntime, licensing, fixtures]
|
| 6 |
+
|
| 7 |
+
# Dependency graph
|
| 8 |
+
requires: ["01-02"]
|
| 9 |
+
provides:
|
| 10 |
+
- "synthesize() returning 24000 Hz WAV bytes plus a plain-dict AudioQuery with per-mora timings"
|
| 11 |
+
- "SynthResult / VisemeEvent / TurnTimings / AvatarDirective dataclasses consumed by 01-06 and 01-08"
|
| 12 |
+
- "zundamon.vvm and the Open JTalk dictionary on disk via Git LFS, loaded with no network at import"
|
| 13 |
+
- "requirements.txt - the Space's four-line dependency contract"
|
| 14 |
+
- "Three AudioQuery fixtures with true WAV durations and frame counts, the ground truth for AVTR-02"
|
| 15 |
+
- "The VERIFIED frame-quantisation formula, which corrects 01-RESEARCH.md"
|
| 16 |
+
- "docs/VOICEVOX-SETUP.md - executed API surface, asset hashes, and the full voice licence chain"
|
| 17 |
+
affects: [01-05, 01-06, 01-07, 01-08, 01-09]
|
| 18 |
+
|
| 19 |
+
# Tech tracking
|
| 20 |
+
tech-stack:
|
| 21 |
+
added: [voicevox_core 0.17.0, voicevox_vvm 0.17.0, VOICEVOX ONNX Runtime 1.23.2, open_jtalk_dic_utf_8-1.11]
|
| 22 |
+
patterns:
|
| 23 |
+
- "Lazy lru_cache(maxsize=1) synthesizer with a warmup() hook: import stays free, startup pays the cost"
|
| 24 |
+
- "AudioQuery converted to a plain ENGINE-schema dict at the module boundary, never handed out live"
|
| 25 |
+
- "Large runtime assets committed through LFS; software binaries referenced by release URL, never vendored"
|
| 26 |
+
- "Directory-scoped .gitattributes to extend LFS coverage without touching another plan's root file"
|
| 27 |
+
- "Fixture generators assert the invariants their fixtures exist to prove before writing anything"
|
| 28 |
+
|
| 29 |
+
key-files:
|
| 30 |
+
created:
|
| 31 |
+
- docs/VOICEVOX-SETUP.md
|
| 32 |
+
- requirements.txt
|
| 33 |
+
- src/japanese_avatar/voice/tts.py
|
| 34 |
+
- src/japanese_avatar/voice/models.py
|
| 35 |
+
- tests/test_tts_contract.py
|
| 36 |
+
- tests/fixtures/make_synth_fixtures.py
|
| 37 |
+
- tests/fixtures/synth_meta.json
|
| 38 |
+
- tests/fixtures/audio_query_short.json
|
| 39 |
+
- tests/fixtures/audio_query_long.json
|
| 40 |
+
- tests/fixtures/audio_query_slow.json
|
| 41 |
+
- tests/fixtures/speech_ja.wav
|
| 42 |
+
- tests/fixtures/speech_ja_long.wav
|
| 43 |
+
- tests/fixtures/speech_ja_slow.wav
|
| 44 |
+
- voicevox/.gitattributes
|
| 45 |
+
- voicevox/model/zundamon.vvm
|
| 46 |
+
- voicevox/open_jtalk_dic_utf_8-1.11/
|
| 47 |
+
modified:
|
| 48 |
+
- pyproject.toml
|
| 49 |
+
- uv.lock
|
| 50 |
+
|
| 51 |
+
key-decisions:
|
| 52 |
+
- "voicevox_vvm 0.17.0, not 0.16.4: the matched release exists and ships vvm_format_version 2, which core 0.17.0 is built around"
|
| 53 |
+
- "ONNX Runtime option (c) lazy first-use download: (a) has no pip distribution at all and (b) has no build hook on a Gradio-SDK Space"
|
| 54 |
+
- "Directory-scoped voicevox/.gitattributes for the dictionary's .dic/.bin rather than editing plan 01-02's root file"
|
| 55 |
+
- "voicevox_core wired into pyproject as a marker-differentiated 'voice' extra so uv sync cannot silently prune it"
|
| 56 |
+
- "AudioQuery serialised to the VOICEVOX ENGINE mixed-case JSON schema, since 0.17.0 exposes no .json() method"
|
| 57 |
+
|
| 58 |
+
patterns-established:
|
| 59 |
+
- "Every API name in a doc carries the file or URL it was read from, and was executed before being written down"
|
| 60 |
+
- "Fixture generators fail loudly rather than emit a fixture that does not exercise its target branch"
|
| 61 |
+
- "Miscounted or mis-shaped acceptance criteria are verified substantively and the discrepancy recorded, never satisfied by contorting code"
|
| 62 |
+
|
| 63 |
+
requirements-completed: [VOIC-01]
|
| 64 |
+
requirements-advanced: [AVTR-02, VOIC-03]
|
| 65 |
+
|
| 66 |
+
# Metrics
|
| 67 |
+
duration: 33min
|
| 68 |
+
completed: 2026-08-27
|
| 69 |
+
---
|
| 70 |
+
|
| 71 |
+
# Phase 01 Plan 04: Japanese Voice Summary
|
| 72 |
+
|
| 73 |
+
**The server speaks Japanese: `synthesize()` returns 24000 Hz WAV bytes plus a plain-dict `AudioQuery` carrying per-mora `consonant_length`/`vowel_length`, loaded from LFS-committed local assets on zero GPU β and the three committed fixtures exposed a frame-quantisation error in `01-RESEARCH.md` that would have broken plan 01-06's slow re-read.**
|
| 74 |
+
|
| 75 |
+
## Performance
|
| 76 |
+
|
| 77 |
+
- **Duration:** ~33 min
|
| 78 |
+
- **Started:** 2026-08-27T00:07Z
|
| 79 |
+
- **Completed:** 2026-08-27T00:40Z
|
| 80 |
+
- **Tasks:** 3 of 3
|
| 81 |
+
- **Files created/modified:** 26, verified by `git diff --name-only 5c6321a HEAD`
|
| 82 |
+
|
| 83 |
+
## Task Commits
|
| 84 |
+
|
| 85 |
+
1. **Task 1: Resolve the 0.17.0 API surface, acquire the runtime assets, write requirements.txt** β `54edb09` (feat)
|
| 86 |
+
2. **Task 2: Implement synthesize() and the shared voice dataclasses** β `b0f00f6` (feat)
|
| 87 |
+
3. **Task 3: Capture the ground-truth synthesis fixtures** β `122fc80` (test)
|
| 88 |
+
4. *(follow-up)* **Record the measured cold-start cost of the lazy runtime fetch** β `84bdf64` (docs)
|
| 89 |
+
|
| 90 |
+
## Everything the plan's `<output>` block required
|
| 91 |
+
|
| 92 |
+
### Corrections to RESEARCH's inferred 0.17.0 API names
|
| 93 |
+
|
| 94 |
+
RESEARCH's construction chain was **entirely correct** β `Onnxruntime.load_once(filename=...)`,
|
| 95 |
+
`OpenJtalk(dict_dir)`, `Synthesizer(ort, ojt)`, `VoiceModelFile.open(path)`,
|
| 96 |
+
`load_voice_model(model)`, `create_audio_query(text, style_id)`, `synthesis(query, style_id)` all
|
| 97 |
+
exist and all work as written. Six things differed:
|
| 98 |
+
|
| 99 |
+
| # | RESEARCH said | Reality | Who it affects |
|
| 100 |
+
|---|---|---|---|
|
| 101 |
+
| 1 | `voicevox_vvm` latest is 0.16.4; confirm the mismatch | **0.17.0 exists** (2026-08-12), one day before core 0.17.0. No mismatch. | this plan |
|
| 102 |
+
| 2 | `json.loads(query.json())` | **`AudioQuery` has no `.json()`** (nor `.to_json`). It *is* a real `dataclass`, so `dataclasses.asdict()` is exact. | 01-06 |
|
| 103 |
+
| 3 | `pauseLength` / `pauseLengthScale` are AudioQuery fields | **They do not exist in CORE 0.17.0** (they are ENGINE-only). Pipeline step 3 is a no-op. | 01-06 |
|
| 104 |
+
| 4 | `SpeakerMeta` | **`CharacterMeta`** (the `.speaker_uuid` attribute is unchanged) | anyone importing it |
|
| 105 |
+
| 5 | Open JTalk dict Β© Nagoya Institute of Technology | **Nara** Institute of Science and Technology β see licence note below | 01-09 |
|
| 106 |
+
| 6 | `round(length / speedScale * 93.75)` | **Wrong.** See the next section β the most consequential finding in this plan. | 01-06 |
|
| 107 |
+
|
| 108 |
+
### The frame-quantisation correction (please read this before plan 01-06)
|
| 109 |
+
|
| 110 |
+
RESEARCH's `build_timeline` divides each phoneme length by `speedScale` and *then* quantises.
|
| 111 |
+
VOICEVOX CORE 0.17.0 actually quantises **first**, at speed 1.0, then divides the resulting **frame
|
| 112 |
+
count** and rounds again:
|
| 113 |
+
|
| 114 |
+
```python
|
| 115 |
+
frames = round(length / speed * 93.75) # 01-RESEARCH.md - WRONG
|
| 116 |
+
frames = round(round(length * 93.75) / speed) # CORRECT
|
| 117 |
+
```
|
| 118 |
+
|
| 119 |
+
At `speedScale == 1.0` the two coincide exactly, which is why RESEARCH's version looked verified β
|
| 120 |
+
and it is exactly why this was worth catching here rather than in 01-06. Measured across 4
|
| 121 |
+
sentences Γ 6 speed values (1.0, 0.9, 0.75, 0.5, 1.25, 1.5), predicted total frames against the
|
| 122 |
+
true frame count of the synthesised WAV:
|
| 123 |
+
|
| 124 |
+
| Formula | Correct |
|
| 125 |
+
|---|---|
|
| 126 |
+
| `round(length / speed * 93.75)` | **8 / 24** |
|
| 127 |
+
| `round(round(length * 93.75) / speed)` | **24 / 24** |
|
| 128 |
+
|
| 129 |
+
Worst observed error from the wrong form: **5 frames β 53 ms**, five times the Β±1 frame
|
| 130 |
+
(10.667 ms) tolerance `test_no_drift_long_utterance` is specified with. This is `01-RESEARCH.md`
|
| 131 |
+
Pitfall 5 ("normal speed syncs, slow speed drifts") arrived at from the other direction.
|
| 132 |
+
|
| 133 |
+
Two consequences for plan 01-06:
|
| 134 |
+
|
| 135 |
+
- Implement the corrected formula. `tests/fixtures/make_synth_fixtures.py` now asserts it
|
| 136 |
+
reproduces each fixture's true frame count before it will write anything, so the fixtures cannot
|
| 137 |
+
drift away from the engine unnoticed.
|
| 138 |
+
- **`test_speed_scale` must not assert "the timeline is exactly 1/0.75Γ longer" against durations.**
|
| 139 |
+
Re-quantisation after scaling means the realised ratio lands *near* the requested one, not on it:
|
| 140 |
+
the committed `slow`/`long` ratio is **1.341085**, not 1.333333. Assert that the timeline matches
|
| 141 |
+
the WAV, which is the property lip-sync actually depends on.
|
| 142 |
+
|
| 143 |
+
Everything else in RESEARCH's pipeline is confirmed: pause mora after its accent phrase, pre/post
|
| 144 |
+
silence included in scaling, consonant before vowel, accumulate quantised frames not floats,
|
| 145 |
+
banker's rounding (Python's built-in `round` already is).
|
| 146 |
+
|
| 147 |
+
### Voice model identity
|
| 148 |
+
|
| 149 |
+
| | |
|
| 150 |
+
|---|---|
|
| 151 |
+
| Original filename | **`0.vvm`** (from `voicevox_vvm` 0.17.0) |
|
| 152 |
+
| Repo path | `voicevox/model/zundamon.vvm`, 59,308,488 B, Git LFS |
|
| 153 |
+
| SHA256 | `ecd35374d4182cd883cba5040376f7f888cc6ba248b1c2f4cea07cdb34bb1318` |
|
| 154 |
+
| `speaker_uuid` | `388f246b-8c41-4ac1-8e2d-5d79f3ff56d9` |
|
| 155 |
+
| **`style_id` (γγγ γγ γγΌγγ«)** | **3** |
|
| 156 |
+
|
| 157 |
+
Resolved from the release's own `README.txt` characterβstyle table and confirmed against the
|
| 158 |
+
`metas.json` inside the downloaded file. The LFS pointer's `oid sha256` matches the hash above
|
| 159 |
+
exactly, confirming the clean filter ran rather than the raw blob being stored.
|
| 160 |
+
|
| 161 |
+
### Version pairing: yes, and it is not optional in the direction that matters
|
| 162 |
+
|
| 163 |
+
**voicevox_core 0.17.0 + voicevox_vvm 0.17.0 β verified by executing it.** `0.vvm` reports
|
| 164 |
+
`vvm_format_version: 2` in its internal `manifest.json` and every style has `type:
|
| 165 |
+
"streaming_talk"`; both the format version and the `StyleType::StreamingTalk` variant were
|
| 166 |
+
introduced *in core 0.17.0*. So core 0.17.0 reads vvm 0.16.4 fine, but core 0.16.x could not read
|
| 167 |
+
vvm 0.17.0. Both are pinned to 0.17.0.
|
| 168 |
+
|
| 169 |
+
### ONNX Runtime acquisition: option **(c)**
|
| 170 |
+
|
| 171 |
+
- **(a) pip-installable companion β does not exist.** `pypi.org/pypi/voicevox-onnxruntime/json`,
|
| 172 |
+
`voicevox_onnxruntime` and `voicevox-core` all return **HTTP 404**. The runtime ships only as
|
| 173 |
+
`.tgz`/`.zip` archives from `VOICEVOX/onnxruntime-builder`. Nothing could go in `requirements.txt`.
|
| 174 |
+
- **(b) downloader binary at build time β no hook exists.** A Gradio-SDK Space's build phase runs
|
| 175 |
+
pip against `requirements.txt` and apt against `packages.txt`; neither can execute `./download`.
|
| 176 |
+
Option (b) is documented as the local-development path.
|
| 177 |
+
- **(c) chosen and exercised.** `get_synthesizer()` resolves `VOICEVOX_ORT_PATH` β an existing copy
|
| 178 |
+
under `voicevox_runtime/` β download the official archive β fall back to `load_once()`'s own
|
| 179 |
+
search path. Pointing `VOICEVOX_ORT_DIR` at an empty directory and calling `warmup()` fetched,
|
| 180 |
+
unpacked, loaded and synthesised successfully. The runtime is always referenced from its official
|
| 181 |
+
release URL and never copied into the repo, per the software terms' η¦ζ’δΊι
.
|
| 182 |
+
|
| 183 |
+
### Measured warm-up
|
| 184 |
+
|
| 185 |
+
| Stage | Time |
|
| 186 |
+
|---|---|
|
| 187 |
+
| `import japanese_avatar.voice.tts` | 0.110 s (no file or socket touched) |
|
| 188 |
+
| `Onnxruntime.load_once(...)` | 0.093 s |
|
| 189 |
+
| `OpenJtalk(dict_dir)` | 0.001 s |
|
| 190 |
+
| `VoiceModelFile.open` + `load_voice_model` | 1.742 s |
|
| 191 |
+
| **`warmup()` total, runtime present** | **1.73 s** |
|
| 192 |
+
| **`warmup()` total, empty runtime cache (includes download)** | **2.66 s** |
|
| 193 |
+
| Second `warmup()` (cache hit) | 0.000001 s |
|
| 194 |
+
|
| 195 |
+
Per turn once warm: `create_audio_query` **1.2 ms**, `synthesis` **1215 ms** for γγγγ«γ‘γ―γ. The
|
| 196 |
+
cold-start penalty this design carries is therefore about **0.9 s** β the 8.2 MB runtime fetch. The
|
| 197 |
+
~166 MB of dictionary and voice model that would otherwise dominate every 48 h wake arrive with the
|
| 198 |
+
clone via LFS instead.
|
| 199 |
+
|
| 200 |
+
### The `long` fixture
|
| 201 |
+
|
| 202 |
+
**Text:** γδ»ζ₯γ―γγ倩ζ°γ§γγγγε
¬εγζ£ζ©γγ¦γγγθ²·γη©γ«θ‘γγΎγγγγ β the plan's suggested
|
| 203 |
+
sentence, kept because it passed every assertion on the first try.
|
| 204 |
+
|
| 205 |
+
| Property | Value |
|
| 206 |
+
|---|---|
|
| 207 |
+
| Mora count (incl. pause moras) | **36** (required β₯ 20) |
|
| 208 |
+
| Phoneme count | 61 |
|
| 209 |
+
| Vowel symbol set | `['I', 'N', 'U', 'a', 'e', 'i', 'o', 'pau']` |
|
| 210 |
+
| **Devoiced vowels present** | **`I`, `U`** β from γ§γγγ and γγ¦γγ / γγΎγγ |
|
| 211 |
+
| `pause_mora` count | **2** |
|
| 212 |
+
| Frames / duration | 516 / **5.504 s** |
|
| 213 |
+
|
| 214 |
+
All three fixture cases, for the record:
|
| 215 |
+
|
| 216 |
+
| Case | speedScale | Moras | Frames | Duration (s) |
|
| 217 |
+
|---|---|---|---|---|
|
| 218 |
+
| `short` (γγγ«γ‘γ―) | 1.0 | 5 | 99 | 1.056 |
|
| 219 |
+
| `long` | 1.0 | 36 | 516 | 5.504 |
|
| 220 |
+
| `slow` (same text) | 0.75 | 36 | 692 | 7.381333333333333 |
|
| 221 |
+
|
| 222 |
+
Regeneration is **byte-identical**: SHA-256 of all seven fixture files is unchanged across a second
|
| 223 |
+
`make_synth_fixtures.py` run.
|
| 224 |
+
|
| 225 |
+
## Accomplishments
|
| 226 |
+
|
| 227 |
+
- **VOIC-01 is genuinely satisfied.** `synthesize("γγγ«γ‘γ―")` returns a 50,732-byte RIFF WAV,
|
| 228 |
+
24000 Hz / mono / 16-bit, 1.056 s, with a JSON-round-trippable `AudioQuery` carrying per-mora
|
| 229 |
+
`consonant_length` and `vowel_length`. `tests/test_tts_contract.py` is **7 passed** against a real
|
| 230 |
+
installed `voicevox_core` β quoted verbatim below.
|
| 231 |
+
- **Zero GPU on the synthesis path, proven by AST scan, not by grep.**
|
| 232 |
+
`test_no_gpu_imports_on_synthesis_path` parses every module under
|
| 233 |
+
`src/japanese_avatar/voice/` and rejects any import of the ZeroGPU allocation package or `torch`.
|
| 234 |
+
- **The licensing branch is closed and written down.** Every fact plan 01-09's `LICENSES.md` needs β
|
| 235 |
+
three terms URLs, the exact `VOICEVOX:γγγ γγ` credit string, the placement standard, the
|
| 236 |
+
flow-down obligation, the AudioQuery-triggers-credit Q&A, the Nemo pre-cleared swap with its style
|
| 237 |
+
IDs, and the BSD-3-Clause dictionary notice β lives in `docs/VOICEVOX-SETUP.md`. No "email the
|
| 238 |
+
rights holder" task was created; SSS LLC ε
責ζ‘ι
2 is quoted so it is never re-litigated.
|
| 239 |
+
- **Cold start retired ahead of time.** 166 MB of assets arrive via LFS; only 8.2 MB is fetched, and
|
| 240 |
+
only inside a lazily-built synthesizer that `warmup()` pre-pays at startup.
|
| 241 |
+
- **Quick loop kept inside its budget** at ~11 s against the 15 s Nyquist limit, after removing six
|
| 242 |
+
redundant syntheses.
|
| 243 |
+
|
| 244 |
+
### Verbatim test output required by the plan
|
| 245 |
+
|
| 246 |
+
```
|
| 247 |
+
$ uv run pytest tests/test_tts_contract.py -q
|
| 248 |
+
....... [100%]
|
| 249 |
+
```
|
| 250 |
+
|
| 251 |
+
Seven tests: the six the plan specifies plus `test_fixtures_match_current_engine` added in Task 3.
|
| 252 |
+
Full quick loop (`pytest tests/ -q --ignore=tests/e2e`): **13 passed**, timed at 10,890 / 11,633 /
|
| 253 |
+
11,352 ms over three runs.
|
| 254 |
+
|
| 255 |
+
## Decisions Made
|
| 256 |
+
|
| 257 |
+
- **`voicevox/.gitattributes` instead of editing the root file.** `sys.dic` is 103 MB and the root
|
| 258 |
+
`.gitattributes` (plan 01-02) arms only `.vrm/.vvm/.wav/.onnx`. Git applies attributes
|
| 259 |
+
hierarchically, so a directory-scoped file adds `*.dic`/`*.bin` coverage without competing for
|
| 260 |
+
ownership of 01-02's file. Verified with `git check-attr filter`, which flipped from
|
| 261 |
+
`unspecified` to `lfs` for all four dictionary binaries.
|
| 262 |
+
- **`voicevox_core` is a `pyproject.toml` extra, not a manual install.** Left out of the project
|
| 263 |
+
metadata, the next `uv sync` would prune it and `tests/test_tts_contract.py` would silently
|
| 264 |
+
degrade to "skipped" for every later plan in the phase. It is wired as the `voice` extra with
|
| 265 |
+
marker-differentiated `[tool.uv.sources]`, so one lockfile serves Windows dev and the Linux Space
|
| 266 |
+
builder. `requirements.txt` remains the Space's independent contract and carries only the
|
| 267 |
+
manylinux wheel.
|
| 268 |
+
- **ENGINE mixed-case JSON schema for the serialised query.** `speedScale` camel, `accent_phrases`
|
| 269 |
+
and all mora keys snake. Not a style choice β it is the wire format the plan's own acceptance
|
| 270 |
+
criteria, RESEARCH's schema block and every VOICEVOX consumer expect.
|
| 271 |
+
- **All three fixture WAVs committed, not just the short one.** `speech_ja_long.wav` and
|
| 272 |
+
`speech_ja_slow.wav` are 264 KB and 354 KB; having the real audio next to the query lets 01-06
|
| 273 |
+
assert against the signal rather than only against a recorded number.
|
| 274 |
+
|
| 275 |
+
## Deviations from Plan
|
| 276 |
+
|
| 277 |
+
### Auto-fixed Issues
|
| 278 |
+
|
| 279 |
+
**1. [Rule 1 - Bug] `01-RESEARCH.md`'s frame-quantisation formula is wrong for `speedScale != 1.0`**
|
| 280 |
+
|
| 281 |
+
- **Found during:** Task 3, when the `slow` fixture's rebuilt timeline missed its WAV by exactly one
|
| 282 |
+
frame while `long` matched to 2.7e-15 s.
|
| 283 |
+
- **Expected:** the plan's fixture table asserts the slow timeline is "exactly `1/0.75x` longer".
|
| 284 |
+
- **Found:** `round(length / speed * 93.75)` is right 8/24 across 6 speeds Γ 4 sentences;
|
| 285 |
+
`round(round(length * 93.75) / speed)` is right 24/24. Worst error 5 frames β 53 ms.
|
| 286 |
+
- **Fix:** documented under a dedicated "Frame quantisation" section in `docs/VOICEVOX-SETUP.md`
|
| 287 |
+
with the measurement table, and enforced in `make_synth_fixtures.py`, which asserts the correct
|
| 288 |
+
formula reproduces each fixture's true frame count before writing. `visemes.py` itself is plan
|
| 289 |
+
01-06's to write β this plan corrects the specification it will be written from, and does not
|
| 290 |
+
pre-empt it.
|
| 291 |
+
- **Files modified:** `docs/VOICEVOX-SETUP.md`, `tests/fixtures/make_synth_fixtures.py`
|
| 292 |
+
- **Committed in:** `122fc80`
|
| 293 |
+
|
| 294 |
+
**2. [Rule 3 - Blocking] The Open JTalk dictionary had no LFS coverage**
|
| 295 |
+
|
| 296 |
+
- **Issue:** `sys.dic` is 103,073,776 B. `git check-attr filter` returned `unspecified` for every
|
| 297 |
+
`.dic`/`.bin` file, so they would have been committed as raw blobs β rejected outright by a Hugging
|
| 298 |
+
Face Space repo, which requires LFS above 10 MB.
|
| 299 |
+
- **Constraint tension:** the executor brief forbids modifying the root `.gitattributes` (owned by
|
| 300 |
+
01-02) while also requiring the dictionary to land through LFS.
|
| 301 |
+
- **Fix:** a new `voicevox/.gitattributes` scoped to the directory this plan owns. Root file
|
| 302 |
+
untouched β confirmed by `git diff --cached --name-only | grep -cx '.gitattributes'` β `0`.
|
| 303 |
+
- **Verification:** all four dictionary binaries flipped to `filter: lfs`; `git lfs status` shows
|
| 304 |
+
`sys.dic`, `char.bin`, `matrix.bin`, `unk.dic` and `zundamon.vvm` as LFS objects.
|
| 305 |
+
- **Committed in:** `54edb09`
|
| 306 |
+
|
| 307 |
+
**3. [Rule 2 - Missing Critical] `voicevox_core` would have been pruned by the next `uv sync`**
|
| 308 |
+
|
| 309 |
+
- **Issue:** the plan says to install the Windows wheel "separately". Anything not in the lockfile is
|
| 310 |
+
removed by `uv sync`, which would have turned the VOIC-01 suite into a silent skip for plans
|
| 311 |
+
01-06/07/08.
|
| 312 |
+
- **Fix:** `voice` extra in `pyproject.toml` with marker-differentiated `[tool.uv.sources]`. First
|
| 313 |
+
attempt failed universal resolution for `sys_platform == 'emscripten'`; corrected by putting a
|
| 314 |
+
disjunctive marker on the requirement itself.
|
| 315 |
+
- **Files modified:** `pyproject.toml`, `uv.lock`
|
| 316 |
+
- **Committed in:** `54edb09`
|
| 317 |
+
|
| 318 |
+
**4. [Rule 1 - Bug] Quick loop had grown to ~20 s against a 15 s budget**
|
| 319 |
+
|
| 320 |
+
- **Issue:** the test module as specified performs six syntheses of which three are exact duplicates
|
| 321 |
+
(`LONG_TEXT` at speed 1.0 three times, at 0.75 twice). `01-VALIDATION.md` caps quick-loop feedback
|
| 322 |
+
latency at ~15 s.
|
| 323 |
+
- **Fix:** a module-scoped `synth` fixture memoising by `(text, speed)`. Three distinct syntheses now
|
| 324 |
+
cover every test. No assertion was weakened or removed.
|
| 325 |
+
- **Verification:** 10,890 / 11,633 / 11,352 ms over three timed runs, 13 passed.
|
| 326 |
+
- **Committed in:** `122fc80`
|
| 327 |
+
|
| 328 |
+
**5. [Rule 1 - Bug] Fixture JSON was written with CRLF on Windows**
|
| 329 |
+
|
| 330 |
+
- **Issue:** `Path.write_text` newline-translates by default, so the committed fixtures were CRLF and
|
| 331 |
+
git warned it would normalise them. A Linux regeneration would then produce different bytes,
|
| 332 |
+
destroying the byte-reproducibility property the fixtures depend on.
|
| 333 |
+
- **Fix:** explicit `newline="\n"` on both writes. Confirmed 0 CRLF bytes in all four JSON files, and
|
| 334 |
+
regeneration is byte-identical.
|
| 335 |
+
- **Committed in:** `122fc80`
|
| 336 |
+
|
| 337 |
+
**6. [Rule 1 - Doc bug] Open JTalk dictionary attribution named the wrong institution**
|
| 338 |
+
|
| 339 |
+
- **Expected:** RESEARCH and the plan's acceptance criterion both say "Nagoya Institute of
|
| 340 |
+
Technology".
|
| 341 |
+
- **Found:** the shipped `COPYING` reads *"Copyright (c) 2009, **Nara** Institute of Science and
|
| 342 |
+
Technology, Japan"*. Both institutions are genuinely involved β Nara (NAIST) holds the dictionary
|
| 343 |
+
data; Nagoya Institute of Technology and the HTS Working Group hold Open JTalk itself.
|
| 344 |
+
- **Fix:** `docs/VOICEVOX-SETUP.md` quotes the NAIST notice verbatim, states both attributions, and
|
| 345 |
+
flags that `LICENSES.md` must carry both β reproducing only Nagoya would fail the BSD notice on the
|
| 346 |
+
files actually redistributed here. The acceptance criterion's literal string is still present.
|
| 347 |
+
- **Committed in:** `54edb09`
|
| 348 |
+
|
| 349 |
+
### Acceptance criteria restated rather than skipped
|
| 350 |
+
|
| 351 |
+
Two criteria are mis-shaped in the same way `01-02` documented. In both cases the substantive
|
| 352 |
+
property was verified directly and no code was contorted.
|
| 353 |
+
|
| 354 |
+
- **`grep -rc "import spaces\|@spaces" src/japanese_avatar/voice/` returns 0.** It initially returned
|
| 355 |
+
`1` β matching my own module docstring, which named the prohibition literally
|
| 356 |
+
(``@spaces.GPU``). The docstring was reworded so the grep is literally true as well as
|
| 357 |
+
substantively true; it now returns 0 for every file. The real guarantee is the AST scan in
|
| 358 |
+
`test_no_gpu_imports_on_synthesis_path`, which a grep cannot provide.
|
| 359 |
+
- **`cases.*.duration_seconds` are "full-precision floats β¦ (more than 3 decimal places)".** `slow`
|
| 360 |
+
is `7.381333333333333`, but `short` is `1.056` and `long` is `5.504` β exactly 3 decimal places.
|
| 361 |
+
These *are* the unrounded header values; VOICEVOX frame counts happen to be divisible by 3 here, so
|
| 362 |
+
`nΒ·256/24000` terminates at 3 dp. The substantive property β that the recorded value equals
|
| 363 |
+
`getnframes()/getframerate()` exactly, with no rounding applied β was asserted directly and holds
|
| 364 |
+
for all three cases (`exact=True`). The generator additionally asserts the written WAV's header
|
| 365 |
+
duration equals what `synthesize()` reported, bit for bit.
|
| 366 |
+
|
| 367 |
+
---
|
| 368 |
+
|
| 369 |
+
**Total deviations:** 6 auto-fixed (3 bugs, 1 blocking, 1 missing-critical, 1 doc bug) + 2 acceptance
|
| 370 |
+
criteria restated
|
| 371 |
+
**Impact on plan:** No scope creep and no design improvisation. Deviations 2 and 3 were required for
|
| 372 |
+
the plan's own criteria to be satisfiable; 1 and 6 correct specifications that later plans would
|
| 373 |
+
otherwise have built on; 4 and 5 protect the phase's feedback loop and fixture reproducibility.
|
| 374 |
+
|
| 375 |
+
## Issues Encountered
|
| 376 |
+
|
| 377 |
+
- **The Open JTalk mirror in the VOICEVOX guide is dead.** `jaist.dl.sourceforge.net` fails DNS
|
| 378 |
+
resolution; `downloads.sourceforge.net` served the same archive (23,643,819 B) without incident.
|
| 379 |
+
Recorded in the setup doc as the working URL.
|
| 380 |
+
- **A ruff lint/format standoff.** `SIM108` demanded a ternary for the ONNX Runtime branch while
|
| 381 |
+
`ruff format` insisted on collapsing that ternary onto a >100-char line. Resolved by shortening the
|
| 382 |
+
expression to `ort_kwargs = {"filename": ort_path} if ort_path else {}` β satisfying both and, as it
|
| 383 |
+
happens, reading better than either.
|
| 384 |
+
- One transient `.ruff_cache` "Access is denied" warning on Windows during a concurrent invocation,
|
| 385 |
+
as in plan 01-02. Cosmetic; the run reported "All checks passed!".
|
| 386 |
+
|
| 387 |
+
## Requirement Status
|
| 388 |
+
|
| 389 |
+
**VOIC-01 β Complete.** `01-VALIDATION.md` binds it to `tests/test_tts_contract.py`, which this plan
|
| 390 |
+
creates and which passes 7/7 against a real installed `voicevox_core`. The behaviour it names β
|
| 391 |
+
"`synthesize()` returns non-empty WAV + parseable `AudioQuery`; zero GPU on the call path" β is
|
| 392 |
+
demonstrated, not asserted. This is the one requirement in this plan's frontmatter and it is
|
| 393 |
+
genuinely satisfied.
|
| 394 |
+
|
| 395 |
+
**Deliberately NOT marked:** the plan's frontmatter lists only `VOIC-01`, so no over-claiming was
|
| 396 |
+
possible here. For the record, this plan *advances* but does not satisfy:
|
| 397 |
+
|
| 398 |
+
- **AVTR-02** β needs `tests/test_visemes.py`, which plan 01-06 creates. This plan supplies its
|
| 399 |
+
ground-truth fixtures and corrects its algorithm.
|
| 400 |
+
- **VOIC-03** β the `speedScale` mechanism works and is fixture-backed, but the requirement's tests
|
| 401 |
+
(`test_speed_scale`, and the deployed `test_slower`) belong to plans 01-06 and 01-10.
|
| 402 |
+
|
| 403 |
+
## Constraint Compliance
|
| 404 |
+
|
| 405 |
+
- **No Space was created.** No `hf` CLI call, no `huggingface_hub` call, no Space README front-matter
|
| 406 |
+
written. Plan 01-01's human hosting gate remains uncleared and untouched.
|
| 407 |
+
- **No `git push` occurred β and could not have.** `git remote -v` is still empty; this repo has no
|
| 408 |
+
remote configured.
|
| 409 |
+
- **No money was spent.** No PRO subscription, no paid resource, no API with a meter.
|
| 410 |
+
- **Python is 3.12.12** throughout (`uv run python -V`).
|
| 411 |
+
- **`voicevox_core` 0.17.0**, wheel filename confirmed against the GitHub Releases API rather than
|
| 412 |
+
guessed. `requirements.txt` carries the **manylinux** wheel; the **win_amd64** wheel is confined to
|
| 413 |
+
the `pyproject.toml` `voice` extra behind a `sys_platform == 'win32'` marker.
|
| 414 |
+
- **File ownership honoured.** `docs/ASSETS.md`, `README.md` and the root `.gitattributes` are all
|
| 415 |
+
absent from this plan's diff β verified with `git diff --cached --name-only`.
|
| 416 |
+
- **Nothing vendored.** `find . -name "*.whl" -not -path "./.venv/*"` β 0. No `.so`, `.dll` or ORT
|
| 417 |
+
binary is tracked; `voicevox_runtime/` was already gitignored by plan 01-02.
|
| 418 |
+
- **No GPU anywhere near synthesis.** AST-verified, and `grep -r "import spaces\|@spaces"` over
|
| 419 |
+
`src/japanese_avatar/voice/` returns 0 matches.
|
| 420 |
+
- **Zero AI tutoring in scope.** No LLM, no level gating, no grammar database, no accounts, no
|
| 421 |
+
database. The discarded TTS fallback branch is absent from `docs/` and `src/` entirely (0 matches),
|
| 422 |
+
deleted rather than deferred.
|
| 423 |
+
- Network egress was limited to what the plan's `<action>` blocks instruct: the GitHub Releases API,
|
| 424 |
+
the voice model, the Open JTalk dictionary, the ONNX Runtime archive, and the `voicevox_core` wheel.
|
| 425 |
+
|
| 426 |
+
## Known Stubs
|
| 427 |
+
|
| 428 |
+
None. Every artifact this plan claims is wired to a real data source and exercised by a passing test.
|
| 429 |
+
|
| 430 |
+
`src/japanese_avatar/voice/models.py` defines `VisemeEvent` and `AvatarDirective`, which nothing
|
| 431 |
+
consumes *yet* β but they are interface definitions the plan explicitly assigns to this file for
|
| 432 |
+
plans 01-06 and 01-08 to consume, not placeholders. They contain no fabricated data and no
|
| 433 |
+
hardcoded empty values flowing to UI.
|
| 434 |
+
|
| 435 |
+
## User Setup Required
|
| 436 |
+
|
| 437 |
+
None. `uv sync --extra dev --extra voice` is the only local step and it has already been run and
|
| 438 |
+
committed to the lockfile.
|
| 439 |
+
|
| 440 |
+
## Next Phase Readiness
|
| 441 |
+
|
| 442 |
+
- **Plan 01-06** is unblocked and materially de-risked: three real `AudioQuery` fixtures with true
|
| 443 |
+
WAV durations *and* frame counts, a long case with 36 moras / 61 phonemes / two devoiced vowels /
|
| 444 |
+
two pause moras, the corrected quantisation formula with its evidence, and an explicit warning that
|
| 445 |
+
`pauseLength`/`pauseLengthScale` do not exist. **Read the "Frame quantisation" section of
|
| 446 |
+
`docs/VOICEVOX-SETUP.md` before writing `visemes.py`.**
|
| 447 |
+
- **Plan 01-05** has `requirements.txt`. Note that the Space also needs `VOICEVOX_ORT_*` to be left at
|
| 448 |
+
defaults so the lazy fetch works, and that `git lfs pull` must succeed on the Space or the assets
|
| 449 |
+
arrive as 130-byte pointers β `tts.py` raises a message naming that case explicitly.
|
| 450 |
+
- **Plan 01-07** has `tests/fixtures/speech_ja.wav` (24000 Hz mono 16-bit), the file
|
| 451 |
+
`tests/conftest.py`'s `speech_wav` fixture has been skipping on since 01-02.
|
| 452 |
+
- **Plan 01-08** has `TurnTimings` and the `warmup()` hook to call at startup, plus real numbers to
|
| 453 |
+
budget against: 1.73 s warm-up, 1.2 ms query, ~1.2 s synthesis.
|
| 454 |
+
- **Plan 01-09** has every voice-side licence fact `LICENSES.md` needs, including the dual
|
| 455 |
+
Nara/Nagoya attribution correction.
|
| 456 |
+
|
| 457 |
+
Carried forward, unchanged by this plan:
|
| 458 |
+
|
| 459 |
+
- The hosting model is still unresolved (ZeroGPU eligibility vs PRO) β plan 01-01's human gate.
|
| 460 |
+
- Per-turn GPU-second cost remains unmeasured (and remains zero for speech output, by construction).
|
| 461 |
+
- Resolved: "VOICEVOX character terms still need a human read" is now **closed** β all three licence
|
| 462 |
+
layers were read in full and recorded, and the rights holder declines qualification questions as a
|
| 463 |
+
matter of policy, so there is no further human step available or required.
|
| 464 |
+
|
| 465 |
+
## Self-Check: PASSED
|
| 466 |
+
|
| 467 |
+
All 16 files claimed created exist on disk; all 2 claimed modified exist. All 4 commit hashes
|
| 468 |
+
(`54edb09`, `b0f00f6`, `122fc80`, `84bdf64`) resolve in `git log`. File count stated as **26** against
|
| 469 |
+
`git diff --name-only 5c6321a HEAD`, not estimated.
|
| 470 |
+
|
| 471 |
+
Full plan verification block re-run after the final commit:
|
| 472 |
+
|
| 473 |
+
| Check | Result |
|
| 474 |
+
|---|---|
|
| 475 |
+
| `pytest tests/test_tts_contract.py -q` | **7 passed** |
|
| 476 |
+
| `pytest tests/ -q --ignore=tests/e2e` | **13 passed**, ~11.0 s |
|
| 477 |
+
| `python -c "β¦synthesize('γγγ«γ‘γ―').duration"` | `1.056` |
|
| 478 |
+
| `git lfs ls-files` | lists `voicevox/model/zundamon.vvm`, `tests/fixtures/speech_ja.wav` (+ 8 more) |
|
| 479 |
+
| `grep -rc "import spaces\|@spaces" src/japanese_avatar/voice/` | **0** for every file |
|
| 480 |
+
| `find . -name "*.whl" -not -path "./.venv/*"` | empty |
|
| 481 |
+
| `ruff check . && ruff format --check .` | "All checks passed!" / "13 files already formatted" |
|
| 482 |
+
| `requirements.txt` non-empty lines | **4**, one `.whl`, zero vendored |
|
| 483 |
+
| `git status --short` | clean β no untracked files |
|
| 484 |
+
| `git remote -v` | empty β no push was possible |
|
| 485 |
+
|
| 486 |
+
---
|
| 487 |
+
*Phase: 01-voice-avatar-loop-skeleton*
|
| 488 |
+
*Completed: 2026-08-27*
|