Spaces:
Running on Zero
Running on Zero
| gsd_state_version: 1.0 | |
| milestone: v1.0 | |
| milestone_name: milestone | |
| status: unknown | |
| stopped_at: "02-10 Tasks 1 and 2 COMPLETE; Task 3 (phone rows) is an OPEN human checkpoint - the only thing left in Phase 2. Owner pushed df2c231..7eb3e94 (55 commits); executor read back sha 7eb3e94d4bbe4feeafc9079a1aa33da35f05a010, stage RUNNING, hardware zero-a10g, DISABLE_GPU=1, and git log space/main..master = 0 after git fetch space. Deployed suite 23 passed / 0 failed on that revision (one failure first time round - test_lookup_popover caught a webfont inside its zero-requests window; scoped in 9e2b095 and the whole file re-run green). docs/LATENCY.md Translation written from the Space: idle p50 276 / p95 357 ms (N=10), mid-turn p50 772 ms (N=3), click->shown p50 1344 ms (N=3), server translate_ms 34 ms idle / 303 ms mid-turn. Warm turns regenerated with the analyze stage (p50 1 ms, N=30); 7257 and 8031 ms on two runs of the SAME revision vs Phase 1's 7107 - no regression claimed. docs/HOSTING.md Phase 2 container record: cpu_cores 16, memory 104.0G, RSS 1334.5 after warm, warm total 14.0 s (translator 11.65 s = 34x the local cost), rss_delta 630.6 vs 410 expected (1.54x, under the 2x flag), four LFS objects proven resolved. JPN-01..04 MARKED on their deployed rows; AVTR-01 deliberately NOT marked (embed probe says the 01-11 audio fix is live - RESULT: speech-end arrived - audible - but a laptop Chromium is not a phone). 02-06's deferred first-turn question ANSWERED: no penalty, the warm-up is paid at Blocks.load (tap 1.2 s after ready -> speech in 7,983 / 8,672 ms; languageInfo() answers in 250 ms). 02-VALIDATION.md: every row filled, no pending markers, nyquist_compliant true, wave_0_complete true, status EXECUTING (signed-off is Task 3's). Local: full suite 399 passed / 0 failed / 0 skipped, quick loop 32.24 s, ruff clean 58 files. Nothing pushed by the executor. Next: the owner's four phone answers." | |
| last_updated: "2026-09-12T14:40:00.000Z" | |
| progress: | |
| total_phases: 6 | |
| completed_phases: 1 | |
| total_plans: 22 | |
| completed_plans: 21 | |
| # Project State | |
| ## Project Reference | |
| See: .planning/PROJECT.md (updated 2026-08-08) | |
| **Core value:** A learner can hold a real, level-appropriate spoken Japanese conversation with an animated avatar that talks back — and measurably improve over time because the avatar remembers them. | |
| **Current focus:** Phase 02 — japanese-language-core | |
| ## Current Position | |
| Phase: 02 (japanese-language-core) — EXECUTING | |
| Plan: 10 of 11 (Task 2 complete; Task 3 human checkpoint OPEN) | |
| > **02-10 Tasks 1 and 2 are DONE; Task 3's phone rows are the only thing left in Phase 2.** | |
| > | |
| > **Task 1 (owner push) — resolved 2026-09-12.** The owner ran `git push space master:main` | |
| > (`df2c231..7eb3e94`, 55 commits, LFS 4 objects / 87 MB). Read back first-hand by the executor: | |
| > `space_info().sha` = `7eb3e94d4bbe4feeafc9079a1aa33da35f05a010`, `.runtime.stage` = `RUNNING`, | |
| > hardware `zero-a10g`, `DISABLE_GPU` = `1`, and `git log space/main..master` = **0** after | |
| > `git fetch space` — the Space runs local `master`'s head exactly. Nothing was pushed by the | |
| > executor. | |
| > | |
| > **Task 2 (deployed verification) — complete 2026-09-12,** commits `9e2b095` (test scoping fix) | |
| > and `b1236c6`. Deployed suite **23 passed / 0 failed** on `7eb3e94`. | |
| > `docs/LATENCY.md § Translation`: idle p50 276 / p95 357 ms (N=10), mid-turn p50 772 ms (N=3), | |
| > click→shown p50 1344 ms (N=3), server `translate_ms` 34 ms idle / 303 ms mid-turn. Warm turns | |
| > re-measured with the new `analyze` stage (p50 **1 ms** over N=30): two runs on the same revision | |
| > read 7257 and 8031 ms against Phase 1's 7107 ms, a spread wider than the difference, so **no | |
| > regression is claimed**. `docs/HOSTING.md § Phase 2 container record`: `cpu_cores` 16, `memory` | |
| > 104.0G, post-warm RSS 1334.5 MB, warm total 14.0 s, `rss_delta` 630.6 vs 410 expected, and all | |
| > four Phase 2 LFS objects proven resolved (`mt_model_bytes` 77,339,435, `jmdict_entries` | |
| > 218,672). **JPN-01..04 marked complete.** Local gates alongside: full suite 399 passed / 0 | |
| > failed / 0 skipped; quick loop 32.24 s; ruff clean on 58 files. | |
| > | |
| > **Task 3 (phone rows) — OPEN, and it needs the owner with a real phone.** Commit `f10c271` | |
| > prepared the exact rows in `docs/LANGUAGE.md § Phone verification`, each with what the deployed | |
| > suite already measured beside it and why that is not the answer. Row 4 (the 人気 override) was | |
| > answered by measurement on the Space — ひとけ before の, にんき otherwise. Rows 1, 2, 3 and 5 | |
| > read *pending owner*; **no verdict was invented**. | |
| > | |
| > The owner opens <https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar> on a phone, | |
| > silent switch OFF, and answers: (1) furigana legibility across the three modes and N5→N2, with | |
| > two screenshots; (2) popover thumb reach at the top, middle and bottom of the column; | |
| > (3) which of the ten deployed English outputs would mislead a learner; (5) Phase 1's carried | |
| > row — was the greeting **audible on the first tap**, does the avatar render, does push-to-talk | |
| > work, on which device / OS / browser. Then: fill those rows, add a **non-EMULATED** row to | |
| > `docs/LATENCY.md § Mobile`, `requirements mark-complete AVTR-01` **only if** audible + renders, | |
| > set `02-VALIDATION.md` to `status: signed-off`, and Phase 2 is closed. | |
| ## Performance Metrics | |
| **Velocity:** | |
| - Total plans completed: 3 | |
| - Average duration: 23 min | |
| - Total execution time: 1.2 hours | |
| **By Phase:** | |
| | Phase | Plans | Total | Avg/Plan | | |
| |-------|-------|-------|----------| | |
| | 01 | 3 | 70min | 23min | | |
| | 02 | 1 (02-10; the other ten were executed in earlier sessions) | 118min | 118min | | |
| **Per-plan detail:** | |
| | Plan | Duration | Tasks | Files | | |
| |------|----------|-------|-------| | |
| | Phase 01 P02 | 21min | 3 tasks | 16 files | | |
| | Phase 01 P04 | 33min | 3 tasks | 26 files | | |
| | Phase 01 P01 | 16min | 3 tasks | 3 files | | |
| **Recent Trend:** | |
| - Last 5 plans: 01-02 (21min), 01-04 (33min), 01-01 (16min) | |
| - Trend: 01-01 was fast because its two expensive tasks were human checkpoints answered outside the executor; the executor's own work was verification and recording | |
| *Updated after each plan completion* | |
| | Phase 01 P03 | 62 | 4 tasks | 20 files | | |
| | Phase 01 P06 | 9 | 3 tasks | 7 files | | |
| | Phase 01 P07 | 118 | 3 tasks | 11 files | | |
| | Phase 01 P05 | multi-session | 3 tasks + T-pose fix | 15 files | | |
| | Phase 01 P08 | 60 | 3 tasks | 17 files | | |
| | Phase 01 P09 | 30 | 3 tasks | 10 files | | |
| | Phase 01 P11 | 65 | 2 tasks | 17 files | | |
| | Phase 02 P01 | 28 | 3 tasks | 18 files | | |
| | Phase 02 P02 | 12 | 3 tasks | 8 files | | |
| | Phase 02 P03 | 11 | 2 tasks | 5 files | | |
| | Phase 02 P04 | 17 | 2 tasks | 13 files | | |
| | Phase 02 P05 | 12 | 2 tasks | 7 files | | |
| | Phase 02 P06 | 38 | 3 tasks | 10 files | | |
| | Phase 02 P07 | multi-session | 3 tasks | 13 files | | |
| | Phase 02 P08 | 95 | 2 tasks | 8 files | | |
| | Phase 02 P09 | 80 | 2 tasks | 10 files | | |
| | Phase 02 P11 | 45min | 3 tasks | 6 files | | |
| ## Accumulated Context | |
| ### Decisions | |
| Decisions are logged in PROJECT.md Key Decisions table. | |
| Recent decisions affecting current work: | |
| - [Roadmap]: Risk-retirement ordering — the avatar/voice transport is proven first with zero AI, because it has the least prior art in a Gradio context and cannot be worked around if it fails. | |
| - [Roadmap]: SudachiPy tokenization (Phase 2) lands before any level gating or grammar grounding — furigana, lookup, level gating, and vocab tracking all collapse without it. | |
| - [Roadmap]: Persistence (Phase 4) is deliberately after the tutor core so the event-sourced schema reflects what the assessment engine actually emits, avoiding a Phase 5 migration. | |
| - [Roadmap]: BYOK (LLMB-02) is owned by Phase 4, not Phase 3 — its session-only, never-leaked guarantee is verified by the same concurrent-session isolation gate as account data. | |
| - [Phase 01]: Ruff excludes *.md: ruff 0.16 formats Python blocks embedded in Markdown, which made the phase's own planning docs fail the lint gate. Lint governs shipped Python, not prose. | |
| - [Phase 01]: uv.lock is committed so a Space rebuild resolves the same dependency set that passed locally; requirements.txt (plan 01-04) remains the Space's contract. | |
| - [Phase 01]: Git LFS armed for .vrm/.vvm/.wav/.onnx in Wave 1, before any binary exists - retires the 10 MiB non-LFS push rejection ahead of plans 01-01 and 01-04. | |
| - [Phase 01]: voicevox_vvm 0.17.0 (not 0.16.4) pairs with core 0.17.0; the vvm ships format v2 + StyleType::StreamingTalk, both introduced in core 0.17.0 | |
| - [Phase 01]: ONNX Runtime uses lazy first-use download (option c): no pip distribution exists and a Gradio-SDK Space has no build hook to run the official downloader | |
| - [Phase 01]: AudioQuery serialised to a plain ENGINE-schema dict at the module boundary; voicevox_core 0.17.0 exposes no .json() method, but the type is a real dataclass | |
| - [Phase 01]: Hosting is zerogpu-free and the alternative is gone: creating the Gradio Space on cpu-basic returned HTTP 402 Payment Required, the identical request on zero-a10g succeeded free. No future phase can fall back to cpu-basic without a PRO purchase. | |
| - [Phase 01]: 1 of 2 free ZeroGPU slots consumed by WolfDavid/japanese-learning-avatar; the remaining slot is the last free Gradio Space this account can create without PRO. | |
| - [Phase 01]: tutor.vrm is VRM1_Constraint_Twist_Sample (pixiv Inc., VRM 1.0), adopted in 01-01 as a dev asset with the sourcing gate armed; embedded VRMC_vrm.meta grants redistribution/avatar-use/modification with creditNotation unnecessary. **Gate closed in 01-10 (2026-09-06): kept as the FINAL shipped character** (docs/ASSETS.md § Sourcing decision). | |
| - [Phase 01]: Audit Space hardware via runtime.hardware.requested, not .current — .current is null for every SLEEPING Space, which made six of eight look unallocated. | |
| - [Phase 01]: gr.HTML js_on_load must be wrapped in an async IIFE: Gradio 6.22.0 compiles it with a plain non-async Function(), so the documented top-level await is a SyntaxError and the component silently never boots | |
| - [Phase 01]: Custom gr.HTML props are **kwargs, not a props= dict; the dict form creates one prop literally named 'props' and props.vrmUrl arrives undefined in the browser | |
| - [Phase 01]: Gradio's auto-generated .pyi component stub is build output: gitignored and ruff-excluded, because create_or_modify_pyi runs unconditionally with no opt-out | |
| - [Phase 01]: Browser assertions sample in-page at animation-frame rate; a 250 ms Python poll cannot reliably observe a 120 ms blink and would be flaky by construction | |
| - [Phase 01]: Viseme frame quantisation ships the CORRECTED order round(round(sec*93.75)/speed) from docs/VOICEVOX-SETUP.md, not 01-RESEARCH.md's round(sec/speed*93.75); all three fixtures now match their WAV to 0 ms | |
| - [Phase 01]: test_speed_scale pins the MEASURED slow/normal ratio 1.341085 plus a per-WAV one-frame match, not the arithmetic ideal 1/0.75 - re-quantisation after scaling puts the ideal 4 frames out on the long sentence | |
| - [Phase 01]: Timeline t accumulates as integer VOICEVOX frames divided by FRAMERATE once, which is what makes exact float equality the correct comparison in the golden-timeline regression guard | |
| - [Phase 01]: Browser ASR dtype is q4, NOT q8: q8/int8/quantized cannot create an ONNX Runtime session on the WASM backend for whisper-base, -small or -large-v3-turbo, while loading fine on WebGPU. 01-RESEARCH.md's q8 recommendation would have shipped green on a GPU machine and dead everywhere else. | |
| - [Phase 01]: WebGPU must be adapter-probed (requestAdapter) before pipeline(device:'webgpu'): a failed WebGPU init poisons the ONNX Runtime backend registry for the whole page, so the researched catch-and-re-instantiate-on-WASM fallback does not fall back. | |
| - [Phase 01]: ASR default is onnx-community/whisper-base q4 (135.8 MB, 779ms WebGPU / 3636ms WASM, CER 0.023 no-punct); better-accuracy toggle is whisper-large-v3-turbo q4f16 WebGPU-only (537.4 MB, CER 0.000). whisper-small rejected as dominated: 2.1x bytes, identical CER, 4.3x WASM inference. | |
| - [Phase 01]: turn-loop.js imports and constructs mic.js/asr.js itself rather than receiving them from a transport, so avatar.js and avatar-iframe.js changed by ZERO lines and both still gained push-to-talk. The iframe transport has no AudioContext to lend, so injection would have made the call sites asymmetric. | |
| - [Phase 01]: Spike verdict CONFIRMED — AVATAR_TRANSPORT=inline ships; gr.HTML hosts three.js + three-vrm on a public ZeroGPU Space with one module instance, and the canvas survives 20 Gradio interactions and the huggingface.co cross-origin embed. The iframe transport stays as the parity-tested fallback. | |
| - [Phase 01]: The VRM's authored rest pose is a T-pose and is posed by the stage at mount (1.2 rad upper arm, 0.2 rad forearm); armDown is measured on the RAW skeleton after render and asserted < -0.7 at three layers. A .vrma idle clip is deferred to a polish plan. | |
| - [Phase 01]: The platform owns the `spaces` version, demands a @spaces.GPU function, and requires a module-level `demo`; none is a project choice (docs/HOSTING.md § First deploy record). | |
| - [Phase 01]: Evidence PNGs are LFS objects — the Hub rejects non-LFS binaries over ~100 KB. | |
| - [Phase 01]: The pre-ASR gate is verified with browser audio processing OFF: Chromium's noise suppression drops the cafe fixture from 0.0577 to 0.0055 RMS, below the floor, leaving the envelope-modulation condition unexercised while the suite looks green. Production keeps all three processing flags on. | |
| - [Phase 01]: gr.HTML server functions receive ONE positional: Gradio 6.22.0's bridge sends a lone argument as itself, several as a list and none as [], so turn() takes a {text, speed} payload and greeting() tolerates a stray positional; server functions never raise because the client turns an HTTP error into undefined | |
| - [Phase 01]: The page's controls are bound in avatar/host.js by elem_id, loaded from the shared boot template after boot, so both transports get identical controls and the only server traffic a turn generates is the server_functions call; transports stay plumbing and turn-loop.js stays host-free | |
| - [Phase 01]: The thinking pose is published as headPitch and relaxedValue (rendered numbers, not a flag) and asserted at the standalone and both-transport layers with shared constants (>0.05 thinking, <0.03 idle); the deployed layer is 01-09's | |
| - [Phase 01]: getDebug() resolves to the facade's LIVE merged object; browser probes read scalars out the moment they sample or they read the final state | |
| - [Phase 01]: Latency is measured in a headed browser: headless SwiftShader rendering starves the CPU synthesiser 10x (synthesis 1.6 s headed vs 10-19 s headless), so the parity suite's lastTurnMs is a correctness signal only; the deployed number is the one that counts | |
| - [Phase 01]: The deployed "server-rendered credit" assertion targets GET /config, not the root HTML: Gradio 6 serves an SSR shell with no component values, so the root carries 0 hits on a correct deployment while /config carries the credit before any project script runs | |
| - [Phase 01]: Deployed browser profiles are fixtures in tests/e2e/conftest.py (fake microphone, no-WebGPU, request counter, event tap installed before navigation, first-frame-aware ready wait); a deployed test requests what it needs rather than repeating launch flags | |
| - [Phase 01]: A probe that reads state inside a facade listener sees the instant BEFORE the turn loop's observe() runs; deployed probes read one macrotask after the event and stamp samples when the snapshot is taken | |
| - [Phase 01]: LICENSES.md carries machine-readable vrm_meta and credits blocks that the deployed suite compares with the shipped VRM's embedded meta and the rendered page, so the document, the binary and the page cannot silently disagree | |
| - [Phase 01]: voicevox_core 0.17.0 is MIT (README § ライセンス at the tag; the LGPL v3 / commercial dual licence applied to prebuilt cores below 0.16) and the Whisper WEIGHTS are Apache-2.0 per the model card while the openai/whisper code is MIT - recorded from primary sources over the wording in CLAUDE.md | |
| - [Phase 01]: VOIC-02/03/04 and DPLY-04 marked Complete on deployed acceptance; AVTR-01/02 and VOIC-05 deliberately left Pending on their manual and 01-10 rows | |
| - [Phase 01]: The 3D runtime is vendored (avatar/vendor/, 930,045 B, text not LFS) from the exact esm.sh es2022 builds the browser executed, rewritten to one ./three.mjs and served by the Space; scripts/vendor_modules.py resolves inner paths from the entry points' own re-export lines and is idempotent. transformers.js stays on esm.sh deliberately: its WASM runtime and the Whisper model come from the HF CDN regardless, and it is the push-to-talk path, not the render path | |
| - [Phase 01]: .mjs is registered as text/javascript in avatar_component.py because module scripts are MIME-checked strictly and the type a host guesses for .mjs is a property of the box (Windows registry says text/plain; measured: every module refused); the served type is now a property of the app | |
| - [Phase 01]: The stage resolves vendored modules with new URL('./vendor/x', import.meta.url) — one rule for the app host and the standalone stage.html — and the stage core still names no host; THREE.Clock → THREE.Timer, removeUnnecessaryJoints dropped | |
| - [Phase 01]: Latency percentiles are nearest-rank, stated in docs/LATENCY.md; the harness asserts no threshold, and a loopback --space-url writes to tmp_path so a rehearsal can never overwrite the Space's record; the owner's Cold start / Mobile / Lip-sync sections are carried forward verbatim on re-run | |
| - [Phase 01]: Deployed p50 7107 ms / p95 10713 ms on c911d74 fired 01-RESEARCH Open Question 6's re-plan trigger; recorded as a roadmap decision (CPU share / cpu_num_threads / streaming synthesis are the levers), not absorbed and not fixed in Phase 1 | |
| - [Phase 01]: The Task 3 manual rows were performed by the execution agent at the owner's instruction and are labelled with who did them and how (Hub API pause + restart; attributed lip-sync review from frames + telemetry; EMULATED mobile; reviewed licence checklist) - never presented as the owner's own observation | |
| - [Phase 01]: VOIC-05 and AVTR-02 marked complete on their bound rows; AVTR-01 deliberately left Pending because no real device has loaded the Space - emulation is recorded as EMULATED, not as a device | |
| - [Phase 01]: Product findings surfaced by the closing plan (boot's hard AudioContext dependency; sub-frame moras under slow rendering) are deferred with evidence rather than fixed inside 01-10, because they are pre-existing behaviour needing product code plus a redeploy | |
| - [Phase 01]: The AudioContext is unlocked by a resume() CALL issued synchronously on the gesture's own call stack, at both ends of the path (turn-loop.js entry points for programmatic callers, host.js handlers for gestures that never reach the loop); the facade forwards stagePort methods synchronously, so the first statement of an entry point is inside the tap and anything after an await is not | |
| - [Phase 01]: playBuffer never emits speech-start on a stopped clock: a bounded, activation-aware resume() (2 s if the document was never activated, 8 s once it has - a fixed 2 s produced a false "audio is blocked" error on the loaded fleet), then 'error' where: 'audio' and a throw; Chromium keeps a blocked resume() promise pending forever, so the bound is what makes the refusal reachable | |
| - [Phase 01]: Strict-autoplay tests read the page through CDP Runtime.evaluate with userGesture:false and tap with raw mouse input, because Playwright's evaluate / wait_for_function / page.click activate the document (measured) and the sticky policy then forgives a late resume; the precondition asserts unactivated AND 'suspended', and the proof is relative (running BEFORE the server answers; a resume() synchronous in a trusted-input listener), never "audio played" | |
| - [Phase 01]: --autoplay-policy=user-gesture-required is not a strict top-level profile (measured: the context is constructed running; it gates cross-origin frames only); it is the policy scripts/embed_autoplay_probe.py uses to reproduce the Hub-embed / Chrome-on-Android silence against the live Space | |
| - [Phase 01]: The gap-closure proof is mutation-tested and the reproduction kept as a runnable tool with its transcript in docs/evidence, so the verifier and the owner can re-run the exact scenario against the pushed revision | |
| - [Phase 02]: JLPT id counts corrected on the pinned tag: 7,748 unique jmdict_seq ids (research said 7,747; all 7,748 resolve in JMdict 3.6.2) and 505 = ids in more than one ROW, of which 447 are on more than one level and 71 are duplicated within one level under two readings - 02-03/02-05 must use 7,748 and 447 | |
| - [Phase 02]: jmdict-eng's top-level version is the bare 3.6.2 (the +20260831182826 build stamp is only in the tag/asset name); build_jmdict asserts 3.6.2 + dictDate 2026-08-31 + 218,672 words, pins the tgz SHA-256, and records the release tag in meta.version | |
| - [Phase 02]: Data builds are byte-idempotent by construction: gzip written with mtime=0 and no embedded name, no volatile numbers (load seconds) in READMEs, expected counts pinned in the script so a mismatch exits 1 with nothing written | |
| - [Phase 02]: Unit joining: A0 (a breaker unit accepts nothing) is the first check inside attach_rule() on the previous unit's head, before A1-A6; A7 (prefix attaches forward) lives in the walk because it alone looks at the next morpheme. A2's tail is the corrected form (助動詞, or 動詞/形容詞 in 未然形/連用形) - the 形容詞 branch makes 美味しくない one unit | |
| - [Phase 02]: Sudachi MORPHEME facts are observed and pinned; unit SPLITS are the D-05 contract and never re-pinned from output. Measured POS facts recorded: で in 三人で is 助動詞 (lemma だ), 行か is 動詞,非自立可能, 三 is OOV with normalized_form '3' so 三人's level_key is '3' - all re-derive to the plan's splits | |
| - [Phase 02]: ruby.is_kanji covers the supplementary planes (U+20000-U+2FA1F, e.g. 𠮷) plus 々; kanji_runs returns only non-kana runs that contain a kanji, so the kanji-axis gate never evaluates a digit/Latin run. jaconv.normalize is never applied to readings (folds ー and small kana) | |
| - [Phase 02]: Human-judged readings live in ONE committed table data/nlp/reading_overrides.json with a mandatory why; apply_overrides is idempotent because the replaced reading no longer equals reading_from_sudachi. JPN-01 left Pending - bound to 02-10's rows | |
| - [Phase 02]: JMdict lookup ranks headword precedence (level_key kanji > level_key kana > lemma kanji > lemma kana) right after the reading match: against the real JLPT table the plan's rank returned 煎る for いる/居る because five listed common entries tied on (reading, on-list, common) and file order picked the earliest; grading by level would still tie 居る with 要る | |
| - [Phase 02]: Two upstream lexicon quirks recorded as known edges for 02-05's fixtures, not patched: n5.csv's なる row is JMdict 2138260 (archaic copula) while the verb 成る 1375610 is on the N3 list, and いい has its own JMdict 3.6.2 entry 2820690 so reading-first sends it there, not to 良い 1605820 | |
| - [Phase 02]: Research Open Question 2 closed: the OPUS-MT CT2 int8 model is committed via LFS in-repo (77.3 MB model.bin + two .spm, pinned to Hub revision 0770961a, regenerated by scripts/convert_mt.py), not published as an owner Hub model repo; no download on the boot path, the Hub-repo + snapshot_download variant is recorded in data/mt/README.md as the fallback | |
| - [Phase 02]: The CTranslate2 converter writes shared_vocabulary.json/config.json with the platform newline; convert_mt.py normalises them to LF because the repo has core.autocrlf=input and hashing a CRLF working copy would fail test_mt_files_match_readme on every fresh clone (the Space). Research's 233 B / 1,269,578 B were Windows-CRLF sizes; committed bytes are 223 B / 1,208,862 B | |
| - [Phase 02]: data/mt/README.md carries run measurements (read-back ms, conversion s) and stays idempotent because the script leaves it untouched when every recorded hash matches; the conversion proves sentencepiece-only pieces == MarianTokenizer on 7 sentences before recording, so the Space imports neither transformers nor torch | |
| - [Phase 02]: JPN-04 left Pending after 02-04: 02-VALIDATION.md binds it to 02-10's deployed test_translation_reveal + docs/LATENCY.md § Translation row; translate() local baseline 0.35 s warm-up, 22-286 ms per contract sentence at intra_threads=2, MT_INTRA_THREADS is the contention lever | |
| - [Phase 02]: analyze(text) is the ONE canonical text -> tokens function: records built from TOKEN_KEYS (14 keys, fixed order), level from the level_key entry and gloss from the lemma entry, non-tappable units carry jlpt None / gloss [] / ruby [[surface, None]]; warm analysis 0.14-0.19 ms per 36-mora sentence | |
| - [Phase 02]: Three plan-assumed word levels corrected to what the pinned lists say and pinned by JMdict id, not patched: 日本語 1464530 is N1+ (on no list; kanji all N5 - the two-axis example), ない/なかった resolve to 無い 1529520 which no list carries (N1+), こんにちは 1289400 is N3; 人気 read ひとけ is the separate entry 1367020 (N1+). A level-alias table is the honest fix if the product wants otherwise (02-10 review), never a weaker lookup | |
| - [Phase 02]: levels.is_kanji IS ruby.is_kanji (02-03's parallel-wave copy removed): one predicate decides both ruby and the kanji-axis gate, so supplementary-plane kanji (𠮷) reach the axis as unlisted instead of being skipped | |
| - [Phase 02]: The fixed sample sentence set (SC-1) follows the golden_timeline pattern: 28 sentences / 115 records generated by the real pipeline, 18 EXPECTED facts checked before writing (exit 1, nothing written otherwise), frozen LF byte-for-byte with data pins read from the environment; JPN-01 still bound to 02-10's rows | |
| - [Phase 02]: warm_language() is a lock around an lru_cache'd body: lru_cache alone lets two concurrent first callers load the compact JMdict twice (CPython calls the function outside any lock), proven by a two-thread test; every server function and turn's analyze stage call it first so a request during Blocks.load waits for the singletons instead of racing a second +300 MB load | |
| - [Phase 02]: Language warm-up measured: local bare process tokenizer 0.04 s / JMdict 2.4 s / translator 0.34 s, RSS +467 MB vs research +410 (02-03's +312 MB JMdict accounts for it); inside app.py +558 MB because warm_synthesizer shares the Blocks.load window; rss_mb() reads /proc on the Space, the Windows working set via ctypes locally (the GetCurrentProcess pseudo-handle must be declared HANDLE) | |
| - [Phase 02]: Bridge helpers are requireBridge() + checkResult(), not one callBridge(name, payload): every turn-loop call site keeps its literal one-payload bridge.x({...}) for the seam grep and dispatchTurn keeps 'await bridge.' on the line the thinking-before-first-await guard reads; transports changed by 0 lines | |
| - [Phase 02]: A full-suite timing failure (strict-autoplay tap-to-running 11.5 s > Phase 1's 10 s sanity ceiling) was A/B'd against the pre-plan commit in a throwaway worktree on the same machine (6.2 s baseline vs 6.1 s HEAD inline) and a stamped probe showed the load warm-up finishing before the first frame: not a regression; the machine-speed ceiling is logged to deferred-items, the load-bearing relative assertion held every run | |
| - [Phase 02]: sudachipy.Tokenizer is NOT shareable across threads - it takes a mutable PyO3 borrow for the whole tokenize() call and raises RuntimeError: Already borrowed at every thread arriving while it is busy (5 of 8 measured). The Dictionary stays process-wide (the mmapped ~110 MB); the Tokenizer is now threading.local, built once per thread under a lock. 02-07 made it reachable: the host analyses the learner's line while the turn's analyse stage runs, and the lost analysis shipped a directive with tokens: [] and an avatar line with no furigana. Guard is mutation-tested. | |
| - [Phase 02]: The <link rel=stylesheet> inside the gr.HTML component value DOES load on the Gradio page (styleSheets carries /gradio_api/file=avatar/transcript.css; .said computes line-height 29.4px = 2.1em) - the inline <style> fallback the plan held in reserve did not ship. The rendered-height proof measures the .turn BLOCK, not the inline .said, because an inline box does not grow when an annotation sits above it | |
| - [Phase 02]: getDebug().furigana does not exist at ready - host.js binds after boot from the shared template - so every furigana read waits for window.Avatar.__debug.furigana; reading the selects in that window made a passing D-14 reload look like a failure | |
| - [Phase 02]: The learner line's analyze round trip measures 14.6-22.7 s in-browser against 0.02-0.19 ms of server work (Gradio serving it around a 12-19 s headless synthesis), so the plan's 5 s budget is recorded and printed, never asserted; the product consequence is logged to deferred-items with two levers | |
| - [Phase 02]: Deployed rows are rehearsed through scripts/rehearse_deployed.py (free port, DISABLE_GPU=1, /config polled, app stopped, pytest's exit code) - the reusable local-deployed harness for the rest of Phase 2; 02-07's three furigana rows are verified locally and await the owner's push, and JPN-02 stays Pending for 02-10 | |
| - [Phase 02]: The word lookup is a PURE CLIENT ACTION and it is asserted as one: a request_counter wraps the taps under both transports and on the deployed page and records [] every time. The glosses ride inside the token record the line already holds, which is why the iframe transport needed no new bridge call for JPN-03 | |
| - [Phase 02]: A popover clamped into the visible column can land ON TOP of the word it is anchored to, and the next tap is then swallowed in silence (the capture-phase dismiss returns early because the target is inside the card, and pointerup finds no .tok). Placement now picks the side of the word with more room and caps the card to that room, so it scrolls internally instead of growing over its anchor | |
| - [Phase 02]: insideTranscript can only be a fact if the transcript has edges: #transcript-text was an auto-height block, so ANY card anchored to the last line hung below it and the plan's phone must-have was unreachable on the real page. It is now height: 40vh / min-height: 14em / overflow-y: auto (335 px on a Pixel 7), which also makes host.js's long-standing el.scrollTop = el.scrollHeight do something for the first time. A 196 px min-height was tried first and was still too short - the card covered the line above its anchor | |
| - [Phase 02]: 02-09: the EN reveal is split at the seam - the renderer owns the control and its five states and cannot translate; host.js owns the per-line session Map, the single avatar.translate() call site and the error state. That split is what makes 'a re-show costs zero requests' checkable with a request counter (recorded [] under inline, iframe and the deployed page) rather than a claim about where a Map lives. | |
| - [Phase 02]: 02-09: docs/LATENCY.md now has two machine authors plus three human sections, so neither harness re-renders it - the warm-turn one CARRIES what it does not own (CARRIED_SECTIONS) and the translation one writes through upsert_section(); the replace-in-place branch that 02-10's re-measurement will take is unit-tested in the quick loop because no local rehearsal can reach it. | |
| - [Phase 02]: 02-09: two plan acceptance checks were unrunnable or hollow and were fixed rather than satisfied - 'no Map( in transcript.js' (the renderer has held two structural Maps since 02-07, so the assertion is an exact count of two) and grep 'avatar.translate(' in host.js (it was matching a COMMENT because the call was chained across lines; the call is now on one line). | |
| - [Phase 02]: 02-11: the three Phase 2 credit strings are defined ONCE as constants in ui/blocks.py and interpolated into both CREDITS_HTML and ABOUT_MD rather than retyped, so the footer and the About panel cannot drift; test_credits_visible loops over every key of LICENSES.md's credits block across footer, About and GET /config | |
| - [Phase 02]: 02-11: the JMdict credit is a PERSISTENT footer line, not a link behind a click - the EDRDG licence (read in full) requires the acknowledgement on each screen display for a WWW server showing words from the files AND separately on an About-style screen, cumulatively | |
| - [Phase 02]: 02-11: SudachiDict's Apache-2.0 rests on the repo's own LICENSE-2.0.txt, LEGAL and README, NOT on the GitHub API, which returns license:null for that repository; the bundled UniDic (BSD-3-Clause) and NEologd are named in the ledger | |
| - [Phase 02]: 02-11: the OPUS-MT row uses the Hub organisation's own fullname (Helsinki-NLP Research Group, University of Helsinki); data/mt/NOTICE keeps the longer historical name and was left untouched as the artefact convert_mt.py shipped | |
| - [Phase 02]: 02-10: EVERY Phase 2 number previously labelled "deployed" was a local DISABLE_GPU=1 rehearsal. The real ones are now in docs/LATENCY.md and docs/HOSTING.md against revision 7eb3e94, and from this plan on "deployed" means read from the running container on a named SHA. | |
| - [Phase 02]: 02-10: JPN-01..04 marked complete, each on its own deployed row passing on 7eb3e94 - the first requirement marks in the phase, deliberately withheld by all ten earlier plans. AVTR-01 stays UNMARKED: scripts/embed_autoplay_probe.py proves the 01-11 audio fix is live on the Space, but a laptop Chromium imitating Chrome-on-Android's autoplay policy is not a phone. | |
| - [Phase 02]: 02-10: the language warm-up is NOT on a visitor's clock. warm.total_s is 14.0 s on the Space (translator 11.65 s of it) but it is paid at Blocks.load; a tap issued ~1.2 s after `ready` reaches speech in 7,983 / 8,672 ms, and languageInfo() - which blocks on warm_language() - answers in 250 ms. 02-06's deferred question is closed; only the cold-container first visit remains open. | |
| - [Phase 02]: 02-10: the CTranslate2 translator loads 34x slower on the Space than locally (11.65 s vs 0.34 s) while tokenizer and JMdict match their local costs - 83 % of the app's boot, and invisible in every local measurement. | |
| - [Phase 02]: 02-10: MT_INTRA_THREADS=1 is now a MEASURED recommendation, not a prediction - mid-turn translate() p50 is 2.80x idle (772 vs 276 ms) and the model's own share goes 34 -> 303 ms. Deliberately not implemented in 02-10, because a code change would invalidate the revision the whole verification record is written against. | |
| - [Phase 02]: 02-10: warm RSS delta on the Space is 630.6 MB against the researched 410 MB (1.54x, under the 2x escalation threshold) and above the local 467 MB on the same code - container overhead, not a duplicated load (warm.errors is empty and jmdict_entries is counted once). | |
| - [Phase 02]: 02-10: a deployed row's "zero requests" assertion is scoped to the data path. test_lookup_popover was failing on a webfont - .lk-surface carries the only font-weight:600 rule in the codebase, so the first card painted fetches SourceSansPro-SemiBold.woff2. Font files are separated and named; anything else, including any /gradio_api/ call, still fails the row. | |
| ### Pending Todos | |
| [From .planning/todos/pending/ — ideas captured during sessions] | |
| None yet. | |
| ### Blockers/Concerns | |
| - **Hosting model — RESOLVED (01-01), and harder than research predicted.** Chosen: `zerogpu-free`. The Space is live at `WolfDavid/japanese-learning-avatar`, public, `requested: zero-a10g`, `gcTimeout` 172800. Account eligibility re-verified live: `isPro: false`, created 2023-11-27, 8 prior Gradio Spaces all `cpu-basic`, 0 ZeroGPU. **The finding that matters:** creating the Space on `cpu-basic` was rejected with **HTTP 402 Payment Required** ("hosting Gradio and Docker Spaces on free cpu-basic requires a PRO subscription"); the identical request on `zero-a10g` succeeded free. **ZeroGPU is therefore the only free path for this project — a later phase cannot fall back to `cpu-basic` without a $9/mo PRO purchase.** **1 of 2 free ZeroGPU slots is consumed; the remaining slot is the last free Gradio Space this account can create without PRO** — budget it against the other HF-profile projects. All evidence in `docs/HOSTING.md`. | |
| - **Asset licensing — VOICEVOX half RESOLVED (01-04), VRM half PROVISIONALLY resolved (01-01) with a gate still armed.** The VOICEVOX branch is closed: all three licence layers were read in full, third-party web synthesis is explicitly contemplated (software cl.3 attaches a flow-down obligation rather than prohibiting it), embedded `.vvm` redistribution is permitted (model cl.2), and the rights holder declines qualification questions as policy (SSS LLC 免責条項 2) — so there is no human step available or required. Compliance is a visible `VOICEVOX:<character>` credit plus a terms notice wherever audio is obtainable; all facts recorded in `docs/VOICEVOX-SETUP.md`. **Voice switched 2026-09-12 (quick task 260912-eg4):** the character is now 雨晴はう (style 10, same `.vvm`), credit `VOICEVOX:雨晴はう`, rights holder Amehare Project, terms https://amehau.com/?page_id=225 read from the primary source (browser UA required; the page 403s bots). Character-name credit is optional under those terms, the VOICEVOX credit mandatory; R18 use of the voice is asked to be avoided; voice-changer models trained on the output are prohibited; and, the character being a nurse, deliberately false or misleading medical claims are prohibited - relevant once Phase 3 lets an LLM speak. Verbatim clauses are in `LICENSES.md § 3. Character`. **VRM half:** `avatar/assets/tutor.vrm` is `VRM1_Constraint_Twist_Sample` by pixiv Inc. (VRM 1.0), whose *embedded* `VRMC_vrm.meta` grants `allowRedistribution`, `avatarPermission: everyone`, `modification: allowModificationRedistribution`, `commercialUsage: corporation`, `creditNotation: unnecessary` — read from the file, not a README, because VRM Public License 1.0 is a per-file template rather than a fixed grant. It is legally clean and, **as of 01-10 (2026-09-06), FINAL**: the sourcing gate was re-raised at the phase-exit checkpoint and the owner kept it (a bespoke VRoid Studio character stays a polish option, not a gate). Facts and the decision are in `docs/ASSETS.md § Sourcing decision`; LICENSES.md's eight-asset checklist was reviewed at the same time with no gaps. Note for anyone tempted: `AvatarSample_A/B/C` are **NOT** CC0. | |
| - **Real-device mobile verification is STILL open (AVTR-01) — but the fix is now provably live and the push has happened.** The owner pushed on 2026-09-12; the Space runs `7eb3e94`. On that revision `scripts/embed_autoplay_probe.py` prints `RESULT: speech-end arrived - audible` with both `AudioContext.resume()` calls `inDispatch: true, sync: true` from a `suspended` state (on `b4d182b` the single resume arrived 8.5 s after the tap and was refused), and the deployed `test_first_tap_is_audible` is green. **AVTR-01 is deliberately NOT marked**: both are a laptop Chromium imitating Chrome-on-Android's autoplay policy, and neither can speak to iOS Safari, a real speaker or the silent switch. To close it: on the phone, silent switch OFF, tap Say hello once and listen; record device / OS / browser / audible / renders / push-to-talk in `docs/LATENCY.md § Mobile` as a row NOT labelled EMULATED, then `requirements mark-complete AVTR-01`. The questions are written out in `docs/LANGUAGE.md § Phone verification` row 5. | |
| - **The avatar boot hard-depends on Web Audio** (`avatar/avatar.js:58` constructs `AudioContext` before the stage mounts): a browser without it shows a blank canvas and "the avatar failed to load". Not expected on real iOS/Android, but a one-line-class fragility on the portfolio path; deferred-items.md. | |
| - **Per-turn GPU-second cost is unmeasured.** The entire free-tier UX budget depends on this number — measure in Phase 1/3, do not assume. | |
| - **REQUIREMENTS.md originally miscounted v1 as 31.** Actual count is 33; traceability and coverage corrected during roadmap creation. | |
| - **Phase 1 plan frontmatter over-claims shared requirements — do not blind-run `requirements mark-complete`.** DPLY-01 appears in the `requirements:` field of plans 01-01, 01-02, 01-05, 01-09 and 01-10, but it is only *satisfied* by 01-05 (`test_space_reachable`). Executing 01-02 marked it Complete; this was reverted, and DPLY-01 is correctly `Pending`. The same over-claiming applies to AVTR-01/02, VOIC-02/03/04/05 and DPLY-04 across plans 01-09 and 01-10. Verify a requirement's actual acceptance test has run before checking it off. **Held on 01-01:** that plan's frontmatter also lists DPLY-01 and DPLY-04; `requirements mark-complete` was deliberately **not** run for either. Both verified still `Pending` after execution. DPLY-01 was satisfied by 01-05's `test_space_reachable` against the live Space on 2026-09-05 and is now Complete; DPLY-04 needs `LICENSES.md` plus `test_vrm_meta_matches_licenses` and `test_credits_visible` from 01-09/01-10. | |
| - CORRECTION for plan 01-06 (not a blocker, a spec fix): 01-RESEARCH.md's frame quantisation is wrong for speedScale != 1.0. Correct form is round(round(length*93.75)/speed), NOT round(length/speed*93.75) - verified 24/24 vs 8/24 across 6 speeds x 4 sentences, worst error 5 frames (53ms). Also: pauseLength/pauseLengthScale do not exist in voicevox_core 0.17.0, and the slow/long duration ratio is 1.341085, not exactly 1/0.75 (re-measured as **1.334669** after the 2026-09-12 switch to 雨晴はう; `test_speed_scale` is pinned to that value at the same 1e-6 tolerance). See docs/VOICEVOX-SETUP.md section 'Frame quantisation'. | |
| - CORRECTION for plans 01-08/01-09/01-10 (not a blocker, a spec fix): 01-RESEARCH.md 'Browser ASR (VOIC-02)' is wrong on two measured points. (1) dtype q8/int8/quantized cannot create an ONNX Runtime session on the WASM backend for any candidate whisper model - use q4, the only quantisation measured working on both tiers. (2) The prescribed 'catch the WebGPU failure and re-instantiate on WASM' fallback does not fall back: a failed pipeline(device:'webgpu') poisons the ORT backend registry for the whole page, so avatar/asr.js probes requestAdapter() first. Also: WebGPU is only available in HEADED Chromium here - headless exposes navigator.gpu and returns a null adapter. See docs/ASR-TIERS.md. | |
| ### Quick Tasks Completed | |
| | # | Description | Date | Commit | Directory | | |
| |---|-------------|------|--------|-----------| | |
| | 260912-eg4 | Switch avatar voice to 雨晴はう (VOICEVOX style 10): speaker + credit surfaces (`94dba69`), regenerated fixtures / golden timeline / demo audio (`831334e`), licence ledger + docs with the character's terms verbatim (`03804fb`) | 2026-09-12 | 70e033b | [260912-eg4-switch-avatar-voice-to-amehare-hau-voice](./quick/260912-eg4-switch-avatar-voice-to-amehare-hau-voice/) | | |
| ## Session Continuity | |
| Last activity: 2026-09-12 - Completed quick task 260912-eg4: switch avatar voice to 雨晴はう (VOICEVOX style 10). Local `master` is 10 commits ahead of `space/main`; **the Space still runs `7eb3e94` with ずんだもん until the owner pushes**. On the push, the deployed rows that exercise the change are `test_credits_visible` (new credit string + amehau.com terms link), `test_ptt_turn` / `test_first_tap_is_audible` (new audio), `test_slower` and the latency harness (regenerated `synth_meta.json`). One local e2e row is red on this machine's strict-autoplay profile and is logged, not hidden: `test_audio_unlocks_inside_the_gesture` - the new voice's コ in こんにちは is 6 frames (64 ms) and the ~15 fps headless renderer drops it (Phase 1's "sub-frame moras under slow rendering" entry; re-run once, failed identically; levers in `deferred-items.md`). Phase 2's Task 3 phone rows remain the open human checkpoint. | |
| Last session: 2026-09-12T14:40:00.000Z | |
| Stopped at: 02-10 Tasks 1 and 2 COMPLETE; Task 3 (phone rows) is an OPEN human checkpoint - the only thing left in Phase 2. Owner pushed df2c231..7eb3e94 (55 commits); executor read back sha 7eb3e94d4bbe4feeafc9079a1aa33da35f05a010, stage RUNNING, hardware zero-a10g, DISABLE_GPU=1, and git log space/main..master = 0 after git fetch space. Deployed suite 23 passed / 0 failed on that revision (one failure first time round - test_lookup_popover caught a webfont inside its zero-requests window; scoped in 9e2b095 and the whole file re-run green). docs/LATENCY.md Translation written from the Space: idle p50 276 / p95 357 ms (N=10), mid-turn p50 772 ms (N=3), click->shown p50 1344 ms (N=3), server translate_ms 34 ms idle / 303 ms mid-turn. Warm turns regenerated with the analyze stage (p50 1 ms, N=30); 7257 and 8031 ms on two runs of the SAME revision vs Phase 1's 7107 - no regression claimed. docs/HOSTING.md Phase 2 container record: cpu_cores 16, memory 104.0G, RSS 1334.5 after warm, warm total 14.0 s (translator 11.65 s = 34x the local cost), rss_delta 630.6 vs 410 expected (1.54x, under the 2x flag), four LFS objects proven resolved. JPN-01..04 MARKED on their deployed rows; AVTR-01 deliberately NOT marked (embed probe says the 01-11 audio fix is live - RESULT: speech-end arrived - audible - but a laptop Chromium is not a phone). 02-06's deferred first-turn question ANSWERED: no penalty, the warm-up is paid at Blocks.load (tap 1.2 s after ready -> speech in 7,983 / 8,672 ms; languageInfo() answers in 250 ms). 02-VALIDATION.md: every row filled, no pending markers, nyquist_compliant true, wave_0_complete true, status EXECUTING (signed-off is Task 3's). Local: full suite 399 passed / 0 failed / 0 skipped, quick loop 32.24 s, ruff clean 58 files. Nothing pushed by the executor. Next: the owner's four phone answers. | |
| Resume file: None | |
| **Roadmap-level open decision:** the latency trigger is STILL fired, and Phase 2 is now measured on the Space rather than predicted. Warm-turn p50 **7,257 ms** on `7eb3e94` (Phase 1 read 7,107 ms on `c911d74`; a second N=30 run on the same Phase 2 revision read 8,031 ms, so the revisions are indistinguishable within the run-to-run spread). **~90 % of it is still VOICEVOX synthesis** — the CPU translation model Phase 2 added costs the turn nothing: the new `analyze` stage is p50 **1 ms** over N=30, and `translate()` is only called when a learner taps EN. So the lever is unchanged and unrelated to the language core: synthesis. Consider `/gsd:insert-phase` for streaming / synthesis relocation before Phase 3 lengthens turns with an LLM. Full record: `docs/LATENCY.md`. | |