# Turn latency (VOIC-05) **Measured:** 2026-09-06T04:32:30Z **Space:** WolfDavid/japanese-learning-avatar **Revision:** `c911d74b378981e6ba0d0111eb570ed3b20927db` **URL measured:** https://wolfdavid-japanese-learning-avatar.hf.space **Hardware:** zero-a10g (stage `RUNNING`) **DISABLE_GPU:** 1 **Browser:** Chromium 151.0.7922.34 (Playwright, headless) **Client:** Windows 11, public internet, one visitor **Method:** N=30 warm text turns after 1 throwaway warm-up (こんにちは), 5 replay turns, five sentences of 5–36 moras in a fixed cycle (`tests/e2e/test_latency_harness.py`). Measured from the page's own `performance.mark` entries; server stages from `__debug.lastStageTimings`. p95 computed by **nearest-rank** (the ceil(0.95·N)-th smallest observed value; p50 likewise), so no percentile is an interpolated value that was never observed. Generated by the harness; the **Cold start**, **Mobile** and **Lip-sync** sections are the owner's and are carried forward verbatim when the harness re-runs. ## Warm turns (N=30) | Stage | p50 (ms) | p95 (ms) | min | max | |---|---|---|---|---| | dispatch -> speech-start (THE NUMBER) | 7107 | 10713 | 4589 | 10858 | | dispatch -> response | 6829 | 10443 | 4326 | 10592 | | server: audio_query | 1 | 2 | 1 | 2 | | server: synthesis | 6422 | 10118 | 4134 | 10120 | | server: timeline | 0 | 1 | 0 | 1 | | server: encode | 0 | 1 | 0 | 1 | | server: total | 6424 | 10121 | 4136 | 10124 | Per sentence (dispatch -> speech-start, ms; each sentence was spoken 6 times): | # | Sentence | Moras | p50 | min | max | synthesis p50 | |---|---|---|---|---|---|---| | 1 | こんにちは | 5 | 6417 | 4589 | 6706 | 5860 | | 2 | はじめまして、よろしくお願いします。 | 17 (hand) | 8142 | 5572 | 10677 | 7510 | | 3 | 今日はいい天気ですから、公園を散歩してから、買い物に行きました。 | 36 | 10207 | 9634 | 10858 | 9653 | | 4 | 駅はどこですか。 | 8 (hand) | 6828 | 6674 | 7113 | 6224 | | 5 | 日本語を勉強しています。 | 14 (hand) | 6365 | 4872 | 8437 | 5835 | ## Replay (N=5, client-side, zero network) | Stage | p50 (ms) | p95 (ms) | min | max | |---|---|---|---|---| | dispatch -> speech-start | 0 | 0 | 0 | 0 | Replay p50 0 ms vs warm-turn p50 7107 ms: the cache is free, the round trip is the cost. ## Cold start | Measurement | Value | How obtained | |---|---|---| | Space PAUSED -> first HTTP 200 serving the Gradio app | **18.0 s** after `restart_space` — stage `BUILDING` at +0.2 s, `APP_STARTING` at +2.2 s, `RUNNING` at +18.6 s | Hub API: `HfApi.pause_space()` (stage `PAUSED` 0.1 s later; the URL answered **503**, 51 bytes, while paused), then `HfApi.restart_space(factory_reboot=False)`; `GET /` polled every 2 s until a 200 whose body names gradio. **Method: pause + restart via the API**, not a factory rebuild | | Page load -> avatar `ready` -> first rendered frame, first visit after the restart | ready **+38.1 s**, first rendered frame **+39.6 s** from the restart (**19.4 s / 20.8 s after navigation**, which began at +18.7 s) | headless Chromium 151 navigated the moment the first 200 arrived; `AVATAR_READY`, then `breathValue !== 0`; `threeInstanceCount` 1, `armDown` -0.932 / -0.932, `mountCount` 1, transport `inline` | | First turn after boot (typed こんにちは -> speech-start) | **7.2 s after Enter** (+46.9 s from the restart); `lastTurnMs` **6600**, `synthesis_ms` 6076; the second turn in the same session: 6368 / 5796 | the first typed turn on the freshly restarted container. The synthesiser is loaded at app start, so the first turn pays **no** model load — the plan's "synthesiser load dominates" prediction did not materialise; the first turn is the same CPU synthesis as every warm turn | | Page load -> avatar `ready` (client-side, no backend), Space already RUNNING | harness run: ready 4.2 s, first rendered frame 5.7 s; lip-sync run (headed, GPU): ready 6.7 s; mobile-emulation runs 9.9–12.4 s (SwiftShader at 2.6–3x DPR) | `wait_for_avatar_ready`; plan 01-05 recorded 4.1–8.3 s warm with tutor.vrm 3.0–6.8 s of it, 01-09 4.0–16.8 s; Space rebuild-to-RUNNING 106 s (01-05); warm wake 0.2–2.7 s | Revision live during the cold-start test: **`b4d182ba9d4ccf77c9910375c9842ddf29496a22`** (read from the Hub before the pause, during the restart and after: unchanged). Measured **2026-09-06T12:09:04Z**; performed by the execution agent through the Hub API on the owner's instruction ("yes to all five"), not from the Space Settings page; the Space was confirmed `RUNNING` afterwards and left that way. Raw record with every stage transition and HTTP poll: `docs/evidence/2026-09-06-cold-start.json`. What this is and is not: a **pause/restart reuses the built image**, which is why it reaches RUNNING in 18 s where the plan 01-05 rebuild took 106 s. The 48-hour `SLEEPING` -> wake path (what a first visitor gets after two idle days) was not forced and remains unmeasured as such; per `01-RESEARCH.md` "cold = SLEEPING or manually paused/restarted", the restart is the sanctioned proxy. The recruiter-facing sum on this path is **~40 s to a moving avatar and ~47 s to the first spoken reply** from a paused container, of which the browser's own VRM load (19–21 s here over the public internet, incl. the 10.3 MiB `tutor.vrm`) is the larger half. ## Mobile **EMULATED — real-device verification still open.** No real phone was available; these rows are Playwright device descriptors driven by the execution agent on the owner's instruction, and they say nothing about a real mobile GPU, thermals or the mobile microphone-permission flow. AVTR-01's manual row therefore stays **open** and the requirement is **not** marked on this evidence. Space revision `b4d182ba9d4ccf77c9910375c9842ddf29496a22`, 2026-09-06T12:32:52Z; raw record `docs/evidence/2026-09-06-mobile-emulation.json`; screenshots `docs/evidence/2026-09-06-mobile-emulated-*.png`. | Device (EMULATED) | OS | Browser | URL | Renders? | Observed smoothness | Push-to-talk | Turn latency | Page -> ready / first frame | |---|---|---|---|---|---|---|---|---| | Playwright `Pixel 7` descriptor (412x839 CSS px, DPR 2.625, touch, mobile UA) | Android 14 UA string | Chromium 151.0.7922.34 headless, WebGL 2 via **SwiftShader** (no GPU) | direct `hf.space` | **yes** — `ready`, first frame, `threeInstanceCount` 1, `mountCount` 1, `armDown` -0.932 / -0.932, canvas 324x480 CSS (648x960 buffer) | **7.7 fps** over 5 s (rAF = getDebug samples) — a slideshow, but that is the software rasteriser on this laptop, not a phone GPU | **not tested** (mobile mic permission needs a device) | typed こんにちは -> speech-start **8.0 s** after Enter; `lastTurnMs` 6612, `synthesis_ms` 5694; mouth peaks aa/ih/oh 1.0; audio 1.056 s | 4.5 s / 6.0 s | | same `Pixel 7` descriptor | Android 14 UA | same Chromium | **huggingface.co Space page** (inside the Hub's `iframe.space-iframe`, `?__theme=system`) | **yes** — same numbers read inside the iframe: 1 instance, mount 1, arms posed, "Running on ZERO" badge above the avatar | **7.3 fps** (SwiftShader) | not tested | speech-start **8.8 s** after Enter; `lastTurnMs` 7354, `synthesis_ms` 6012; turn completed | 5.7 s / 8.0 s | | Playwright `iPhone 14` descriptor (390x664, DPR 3, touch, iOS UA) | iOS UA string | **WebKit 26.5, Playwright's Windows build** — WebGL 2 present ("Apple GPU"), **no Web Audio API** (`AudioContext` and `webkitAudioContext` both `undefined`) | direct and Hub page | **NO — not a device result.** Boot aborts before the stage mounts: `avatar boot failed: TypeError: undefined is not a constructor (evaluating 'new AudioCtor()')`, status "the avatar failed to load - see the browser console", blank canvas (`…iphone14-webkit-bootfail.png`) | — | — | — | — | What the emulation does and does not establish: - **Established on the Android/Chromium profile:** the responsive layout stacks the canvas above the controls at a phone width; the VRM, the idle life and a full typed turn (server round trip, decode, lip-synced playback) work at mobile viewport, DPR and user agent, both on the direct URL and embedded in the Hub page. The Hub page swaps its iframe once while resolving the Space (a frame found too early detaches), and the Hub's own script logs one `getBoundingClientRect` error at that moment — neither is the app's. - **Not established:** frame rate on a phone (SwiftShader's 7–8 fps is this laptop's CPU rasterising a 648x960 buffer; a real Adreno/Mali/Apple GPU is a different measurement), warmth, battery, the mobile mic-permission and push-to-talk flow, and anything about Safari: **the iPhone rows are a host limitation, not an iOS result.** Real iOS Safari 14.5+ ships `AudioContext`, so the boot is expected to succeed there — but it is expected, not observed. - **Finding for the product (deferred, `deferred-items.md`):** the boot constructs the `AudioContext` before the stage mounts (`avatar/avatar.js:58-59`), so a browser without Web Audio gets a blank canvas even though rendering needs no audio. Constructing it lazily on the first turn or mic press would let the avatar render everywhere and only degrade sound. **To close the row:** open `https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar` on a real phone; record device, OS version, browser version, renders / smooth-or-slideshow, warmth, whether the mic permission and push-to-talk worked, and a rough turn time; then mark AVTR-01. **Audio on the phone (01-HUMAN-UAT gap 1, fixed in plan 01-11 — retest after the redeploy).** The owner's first phone visit (2026-09-06, revision `b4d182b`) rendered but was silent: the only `AudioContext.resume()` on that build ran inside `playBuffer`, seconds after the tap, and a gesture-gated browser refuses it. Reproduced in Chromium with the Space inside a cross-origin frame under `--autoplay-policy=user-gesture-required` (Chrome-on-Android's policy, and the Hub page's embed): the resume was refused 8.5 s after the tap and the page stayed on "thinking…" (`docs/evidence/2026-09-06-autoplay-embed-b4d182b.txt`). Plan 01-11 resumes the context inside every tap. After the push: 1. On the phone, open the Hub page, make sure the **ring/silent switch is OFF** (on iOS it mutes Web Audio entirely and looks identical to this bug) and the media volume is up. 2. Tap **Say hello** once, before anything else. Expect the greeting to be audible within ~10 s and the status line to go "thinking…" → "speaking…" → "ready". 3. Then a typed turn and a Replay. Both should be audible without another unlock. 4. If it stays on "thinking…", or the status reads "audio is blocked by the browser - tap Say hello or Send once, then try again", record device / OS / browser and the status text here. From a laptop, `python scripts/embed_autoplay_probe.py` re-runs the cross-origin embed scenario against the live Space and prints whether `speech-end` arrived. ## Lip-sync verification (AVTR-02) **Reviewed from captured frames and viseme telemetry by the execution agent, not by the owner** (the owner asked for the manual rows to be performed on their behalf; this section is that review and says so). Space revision `b4d182ba9d4ccf77c9910375c9842ddf29496a22`, 2026-09-06T12:16:28Z, **headed** Chromium 151 (real GPU, so the render rate is a visitor's, not SwiftShader's), viewport 800x600 at device-scale 2. **Sentence (113 chars, 211 timeline events, decoded audio 20.565 s):** こんにちは。今日はいい天気ですね。昨日、駅の近くの喫茶店で、友達とコーヒーを飲みました。ちょっと高かったですが、とても美味しかったです。明日は、きっと図書館で日本語を勉強します。それから、うちの近くの公園を、ゆっくり歩きます。 — chosen to carry ん (こんにちは, 天気, 飲み, 勉強, 公園), っ (ちょっと, きっと, ゆっくり), 、and 。pauses, です/ました/ます endings (devoiced vowels) and several voiced う (うち, ゆっくり, 歩きます). Submit -> speech-start **22.06 s** (`synthesis_ms` 20,700 for 20.6 s of audio: the Space's CPU synthesises long text at roughly **1.0x real time**). Speech-start -> speech-end 20.66 s wall for a 20.565 s buffer. **Recording:** `docs/evidence/2026-09-06-lipsync-20s.webm` (23 s, the utterance window of the page recording, VP8 800x600, 1.98 MB; the 57 s original was mostly the avatar idling while the server synthesised). **Face crops during the utterance:** `docs/evidence/2026-09-06-lipsync-frames-01..10.png`, each targeted at a long timeline event and labelled with the mouth read immediately before and after the capture (`docs/evidence/2026-09-06-lipsync-telemetry.json` § frames). **Reference sheet** of the five presets held at weight 1.0 on the same VRM by the same `vrm-stage.js` (local static server, same canvas size): `docs/evidence/2026-09-06-viseme-ref-{closed,aa,ih,ou,ee,oh}.png`. **Telemetry:** 2,063 samples of `currentVisemes` + the player's audio clock at ~100 samples/s (mean interval 10.0 ms) from speech-start to speech-end, plus the server's timeline for the same utterance — all in the telemetry JSON. | Frame | Target (timeline) | Mouth read after the capture | Reads as | |---|---|---|---| | 01 | い @0.50 s (96 ms) | ih 0.20 (fading in) | mostly closed, opening | | 02 | あ @0.85 s (160 ms) | aa 0.69 | open oval — あ | | 03 | い @1.79 s (96 ms) | ih 0.88 (aa 0.11 residual) | open, on the あ->い cross-fade | | 04 | え @2.69 s (171 ms) | ee 0.92 | flat, wide, teeth visible — え | | 05 | お @3.50 s (160 ms) | oh 1.00 | tall rounded — お | | 06 | あ @8.09 s (160 ms) | aa 1.00 (captured 0.9 s late, at 9.09 s — a later あ; the screenshot call stalled) | open oval — あ | | 07 | う @17.10 s (117 ms) | ou 0.60 | small, pursed — う | | 08 | pause @18.79 s (459 ms, before それから) | all 0 (oh 0.19 decaying at the read) | shut | | 09 | う @19.32 s (107 ms) | ou 0.87 | small, pursed — う | | 10 | devoiced う @20.42 s (ます, weight 0.5) | ou 0.21 | nearly shut, slight purse | **The four judgements** 1. **Do あ / い / う visibly differ? — YES.** Each preset is driven independently to full weight: stage peak-hold over the utterance `aa` 1.00, `ih` 1.00, `ou` 0.92, `ee` 1.00, `oh` 1.00 (`ou` tops out at 0.92 on a 117 ms う and at 0.66–0.69 on 53 ms ones — the 50 ms cross-fade in `lipsync.js` is the limit, not the timeline). On this VRM あ is a tall open oval (ref-aa, frames 02/06), い a flat wide opening with teeth (ref-ih), う a small pursed mouth (ref-ou, frames 07/09) — three shapes no one would confuse, so this is not amplitude flapping. **Two honest caveats.** (a) On this model `ih`≈`ee` and `aa`≈`oh` are visually near-identical pairs (reference sheet): the five presets read as three groups. That is a property of the sample character, recorded in `docs/ASSETS.md`, not of the pipeline. (b) Both deployed い captures (01, 03) landed on a cross-fade — い moras in this sentence are 96 ms, shorter than the ~110 ms a screenshot round trip takes — so the clean い is the reference frame; the telemetry shows `ih` at 1.00 with `aa` 0 on the 21 full-weight い/し/ち moras. 2. **Does the mouth close on ん, っ and pauses? — YES.** The timeline has 113 `closed` events (N, cl, pau, and consonant onsets). Every pause ≥ 139 ms (the 、/。 rests, 245–459 ms; 12 events) reached a sampled mouth sum of **exactly 0**. The mora-length closures — ん and っ, 85–107 ms, 17 events — reached a **median 0.02** of open (min 0, max 0.44): shut to the eye, the residual being the 50 ms cross-fade's tail on the shortest ones. The 32–75 ms consonant onsets inside moras dip to a median 0.20 and reopen, which is how speech looks. Frame 08 is the 459 ms pause: shut. 3. **Is the END of the sentence as well synced as the start? — YES, no drift.** The timeline ends at **20.565333 s** and the decoded `AudioBuffer` lasts **20.565333 s** — a **0.000 ms** difference over 211 events. Per-event agreement (the dominant sampled viseme at 60 % through every vowel event ≥ 90 ms) is **19/19 in the first third, 19/19 in the second, 22/22 in the last** — zero mismatches. The last vowel (the devoiced う of ます, 20.363–20.469 s) released on time and the last sample with the mouth open is at clock 20.550 s, 15 ms before speech-end; the first vowel event is at 0.192 s and the first open sample at 0.198 s. Clock discipline held: the player reads `audioCtx.currentTime - scheduledStart`, and the last third looks exactly like the first. 4. **Does the mouth freeze on です / した (devoiced vowels)? — NO, it moves at half weight as designed.** Eight devoiced moras (weight 0.5 in the timeline: す ×5 incl. です/します/ます, し/ち ×3 in ました/美味しかった/ちょっと) opened to sampled peaks of **0.35–0.45** against the 0.5 target (attack-limited on 53–107 ms moras) and were never held open (0 past 0.75) and never frozen shut (0 at 0.00 while a vowel was due). Frame 10 is the final devoiced ます: nearly shut with a slight purse, which is the intended look. The uppercase-vowel path is intact. **Verdict of this review: all four criteria hold on the deployed Space; no drift, no frozen mouth, no amplitude flapping.** Escalation per the plan is therefore not triggered. The judgement is the execution agent's from the evidence above; the owner has the recording and the frames to overrule it. **Findings recorded, not fixed (Phase 1 ships as is):** (a) under **headless SwiftShader at ~15 fps** — a first, discarded run — a lone 53 ms う was skipped entirely (peak-hold `ou` 0.5, from a devoiced mora only): a mora shorter than one render frame can be dropped when the browser renders below ~19 fps. Real GPUs render at 60+; the emulated-mobile rows above (6.7 fps under SwiftShader) say what a genuinely slow device would look like, which is one more reason the real-device row stays open. (b) The 50 ms cross-fade caps how far a 53 ms mora opens (~0.66) — a tunable, and a design choice this review did not second-guess. ## Interpretation - **p50 dispatch -> speech-start is 7107 ms** (over the ~1500 ms VOIC-05 perceived-response target; OVER the 2500 ms re-plan trigger). > **RE-PLAN TRIGGER FIRED.** p50 7107 ms exceeds 2500 ms. Per `01-RESEARCH.md` Open Question 6 - *"treat p50 > ~2.5 s as a re-plan trigger, not a phase failure"* - this is a decision for the roadmap, not a failing test: the criterion for Phase 1 is *"p50/p95 measured and recorded"*, which this file is. - **What dominates:** VOICEVOX synthesis on the Space's CPU - server `synthesis_ms` p50 6422 ms is 90 % of the p50 turn; `audio_query` (1.4 ms), `timeline` (0.36 ms) and `encode` (0.17 ms) are noise. Network + Gradio round trip above the server total: 405 ms at p50; browser decode + schedule after the response: 278 ms at p50. - **Cheapest improvement (recorded, NOT implemented in Phase 1):** the same sentence synthesises 3–6x faster on a developer laptop than on the Space (plan 01-09), so the container's CPU share is the lever - check `cpu_num_threads` on the Space's `Synthesizer` and whether the ZeroGPU container's CPU allocation is the ceiling; beyond that, VOICEVOX CORE 0.17's streaming synthesis would let speech start before the whole utterance is rendered, and per-sentence synthesis of long replies would bound the first-sample latency by the first sentence rather than the whole reply. ## Notes - Phase 1 has no LLM. These numbers are NOT predictive of Phase 3, which adds one. - Space sleeps after 48 h (gcTimeout 172800), so a cold start is the default first-visit experience. - Numbers live in git because Space disk is ephemeral. Continuous telemetry is deliberately deferred to Phase 4/6, which owns the privacy story. - Headless Chromium renders through SwiftShader on the client; the client-side decode + schedule share above is therefore an upper bound on what a GPU-rendered visitor sees. ## Raw samples (dispatch -> speech-start, ms, in run order) | Turn | Sentence # | dispatch->speech | dispatch->response | synthesis | lastTurnMs | |---|---|---|---|---|---| | 1 | 1 | 6538 | 6262 | 5997 | 6538 | | 2 | 2 | 9639 | 9361 | 8902 | 9639 | | 3 | 3 | 9661 | 9390 | 9165 | 9661 | | 4 | 4 | 6674 | 6401 | 6117 | 6674 | | 5 | 5 | 7722 | 7441 | 7100 | 7722 | | 6 | 1 | 4589 | 4326 | 4134 | 4589 | | 7 | 2 | 10677 | 10406 | 10099 | 10677 | | 8 | 3 | 10713 | 10443 | 10120 | 10713 | | 9 | 4 | 7113 | 6836 | 6503 | 7113 | | 10 | 5 | 6365 | 6086 | 5787 | 6365 | | 11 | 1 | 6417 | 6144 | 5860 | 6417 | | 12 | 2 | 8740 | 8460 | 8201 | 8740 | | 13 | 3 | 10858 | 10592 | 10118 | 10858 | | 14 | 4 | 7107 | 6829 | 6422 | 7107 | | 15 | 5 | 8437 | 8157 | 7936 | 8437 | | 16 | 1 | 6662 | 6390 | 6097 | 6662 | | 17 | 2 | 7552 | 7282 | 6998 | 7553 | | 18 | 3 | 9634 | 9360 | 9083 | 9634 | | 19 | 4 | 7113 | 6847 | 6646 | 7113 | | 20 | 5 | 6635 | 6364 | 6089 | 6635 | | 21 | 1 | 6706 | 6434 | 6081 | 6706 | | 22 | 2 | 8142 | 7868 | 7510 | 8142 | | 23 | 3 | 10552 | 10277 | 9963 | 10552 | | 24 | 4 | 6828 | 6552 | 6127 | 6828 | | 25 | 5 | 4872 | 4601 | 4284 | 4872 | | 26 | 1 | 6188 | 5912 | 5713 | 6188 | | 27 | 2 | 5572 | 5297 | 4900 | 5572 | | 28 | 3 | 10207 | 9924 | 9653 | 10207 | | 29 | 4 | 6775 | 6505 | 6224 | 6775 | | 30 | 5 | 6337 | 6068 | 5835 | 6337 | Replays (dispatch -> speech-start, ms): 0, 0, 0, 0, 0