WolfDavid's picture
test(01-11): prove the tap unlocks audio under a strict autoplay policy at three layers
73b5215
|
Raw History Blame
21.4 kB

Turn latency (VOIC-05)

Measured: 2026-09-06T04:32:30Z
Space: WolfDavid/japanese-learning-avatar Revision: c911d74b378981e6ba0d0111eb570ed3b20927db URL measured: https://wolfdavid-japanese-learning-avatar.hf.space
Hardware: zero-a10g (stage RUNNING) DISABLE_GPU: 1
Browser: Chromium 151.0.7922.34 (Playwright, headless) Client: Windows 11, public internet, one visitor
Method: N=30 warm text turns after 1 throwaway warm-up (ใ“ใ‚“ใซใกใฏ), 5 replay turns, five sentences of 5โ€“36 moras in a fixed cycle (tests/e2e/test_latency_harness.py). Measured from the page's own performance.mark entries; server stages from __debug.lastStageTimings.
p95 computed by nearest-rank (the ceil(0.95ยทN)-th smallest observed value; p50 likewise), so no percentile is an interpolated value that was never observed.

Generated by the harness; the Cold start, Mobile and Lip-sync sections are the owner's and are carried forward verbatim when the harness re-runs.

Warm turns (N=30)

Stage p50 (ms) p95 (ms) min max
dispatch -> speech-start (THE NUMBER) 7107 10713 4589 10858
dispatch -> response 6829 10443 4326 10592
server: audio_query 1 2 1 2
server: synthesis 6422 10118 4134 10120
server: timeline 0 1 0 1
server: encode 0 1 0 1
server: total 6424 10121 4136 10124

Per sentence (dispatch -> speech-start, ms; each sentence was spoken 6 times):

# Sentence Moras p50 min max synthesis p50
1 ใ“ใ‚“ใซใกใฏ 5 6417 4589 6706 5860
2 ใฏใ˜ใ‚ใพใ—ใฆใ€ใ‚ˆใ‚ใ—ใใŠ้ก˜ใ„ใ—ใพใ™ใ€‚ 17 (hand) 8142 5572 10677 7510
3 ไปŠๆ—ฅใฏใ„ใ„ๅคฉๆฐ—ใงใ™ใ‹ใ‚‰ใ€ๅ…ฌๅœ’ใ‚’ๆ•ฃๆญฉใ—ใฆใ‹ใ‚‰ใ€่ฒทใ„็‰ฉใซ่กŒใใพใ—ใŸใ€‚ 36 10207 9634 10858 9653
4 ้ง…ใฏใฉใ“ใงใ™ใ‹ใ€‚ 8 (hand) 6828 6674 7113 6224
5 ๆ—ฅๆœฌ่ชžใ‚’ๅ‹‰ๅผทใ—ใฆใ„ใพใ™ใ€‚ 14 (hand) 6365 4872 8437 5835

Replay (N=5, client-side, zero network)

Stage p50 (ms) p95 (ms) min max
dispatch -> speech-start 0 0 0 0

Replay p50 0 ms vs warm-turn p50 7107 ms: the cache is free, the round trip is the cost.

Cold start

Measurement Value How obtained
Space PAUSED -> first HTTP 200 serving the Gradio app 18.0 s after restart_space โ€” stage BUILDING at +0.2 s, APP_STARTING at +2.2 s, RUNNING at +18.6 s Hub API: HfApi.pause_space() (stage PAUSED 0.1 s later; the URL answered 503, 51 bytes, while paused), then HfApi.restart_space(factory_reboot=False); GET / polled every 2 s until a 200 whose body names gradio. Method: pause + restart via the API, not a factory rebuild
Page load -> avatar ready -> first rendered frame, first visit after the restart ready +38.1 s, first rendered frame +39.6 s from the restart (19.4 s / 20.8 s after navigation, which began at +18.7 s) headless Chromium 151 navigated the moment the first 200 arrived; AVATAR_READY, then breathValue !== 0; threeInstanceCount 1, armDown -0.932 / -0.932, mountCount 1, transport inline
First turn after boot (typed ใ“ใ‚“ใซใกใฏ -> speech-start) 7.2 s after Enter (+46.9 s from the restart); lastTurnMs 6600, synthesis_ms 6076; the second turn in the same session: 6368 / 5796 the first typed turn on the freshly restarted container. The synthesiser is loaded at app start, so the first turn pays no model load โ€” the plan's "synthesiser load dominates" prediction did not materialise; the first turn is the same CPU synthesis as every warm turn
Page load -> avatar ready (client-side, no backend), Space already RUNNING harness run: ready 4.2 s, first rendered frame 5.7 s; lip-sync run (headed, GPU): ready 6.7 s; mobile-emulation runs 9.9โ€“12.4 s (SwiftShader at 2.6โ€“3x DPR) wait_for_avatar_ready; plan 01-05 recorded 4.1โ€“8.3 s warm with tutor.vrm 3.0โ€“6.8 s of it, 01-09 4.0โ€“16.8 s; Space rebuild-to-RUNNING 106 s (01-05); warm wake 0.2โ€“2.7 s

Revision live during the cold-start test: b4d182ba9d4ccf77c9910375c9842ddf29496a22 (read from the Hub before the pause, during the restart and after: unchanged). Measured 2026-09-06T12:09:04Z; performed by the execution agent through the Hub API on the owner's instruction ("yes to all five"), not from the Space Settings page; the Space was confirmed RUNNING afterwards and left that way. Raw record with every stage transition and HTTP poll: docs/evidence/2026-09-06-cold-start.json.

What this is and is not: a pause/restart reuses the built image, which is why it reaches RUNNING in 18 s where the plan 01-05 rebuild took 106 s. The 48-hour SLEEPING -> wake path (what a first visitor gets after two idle days) was not forced and remains unmeasured as such; per 01-RESEARCH.md "cold = SLEEPING or manually paused/restarted", the restart is the sanctioned proxy. The recruiter-facing sum on this path is ~40 s to a moving avatar and ~47 s to the first spoken reply from a paused container, of which the browser's own VRM load (19โ€“21 s here over the public internet, incl. the 10.3 MiB tutor.vrm) is the larger half.

Mobile

EMULATED โ€” real-device verification still open. No real phone was available; these rows are Playwright device descriptors driven by the execution agent on the owner's instruction, and they say nothing about a real mobile GPU, thermals or the mobile microphone-permission flow. AVTR-01's manual row therefore stays open and the requirement is not marked on this evidence. Space revision b4d182ba9d4ccf77c9910375c9842ddf29496a22, 2026-09-06T12:32:52Z; raw record docs/evidence/2026-09-06-mobile-emulation.json; screenshots docs/evidence/2026-09-06-mobile-emulated-*.png.

Device (EMULATED) OS Browser URL Renders? Observed smoothness Push-to-talk Turn latency Page -> ready / first frame
Playwright Pixel 7 descriptor (412x839 CSS px, DPR 2.625, touch, mobile UA) Android 14 UA string Chromium 151.0.7922.34 headless, WebGL 2 via SwiftShader (no GPU) direct hf.space yes โ€” ready, first frame, threeInstanceCount 1, mountCount 1, armDown -0.932 / -0.932, canvas 324x480 CSS (648x960 buffer) 7.7 fps over 5 s (rAF = getDebug samples) โ€” a slideshow, but that is the software rasteriser on this laptop, not a phone GPU not tested (mobile mic permission needs a device) typed ใ“ใ‚“ใซใกใฏ -> speech-start 8.0 s after Enter; lastTurnMs 6612, synthesis_ms 5694; mouth peaks aa/ih/oh 1.0; audio 1.056 s 4.5 s / 6.0 s
same Pixel 7 descriptor Android 14 UA same Chromium huggingface.co Space page (inside the Hub's iframe.space-iframe, ?__theme=system) yes โ€” same numbers read inside the iframe: 1 instance, mount 1, arms posed, "Running on ZERO" badge above the avatar 7.3 fps (SwiftShader) not tested speech-start 8.8 s after Enter; lastTurnMs 7354, synthesis_ms 6012; turn completed 5.7 s / 8.0 s
Playwright iPhone 14 descriptor (390x664, DPR 3, touch, iOS UA) iOS UA string WebKit 26.5, Playwright's Windows build โ€” WebGL 2 present ("Apple GPU"), no Web Audio API (AudioContext and webkitAudioContext both undefined) direct and Hub page NO โ€” not a device result. Boot aborts before the stage mounts: avatar boot failed: TypeError: undefined is not a constructor (evaluating 'new AudioCtor()'), status "the avatar failed to load - see the browser console", blank canvas (โ€ฆiphone14-webkit-bootfail.png) โ€” โ€” โ€” โ€”

What the emulation does and does not establish:

  • Established on the Android/Chromium profile: the responsive layout stacks the canvas above the controls at a phone width; the VRM, the idle life and a full typed turn (server round trip, decode, lip-synced playback) work at mobile viewport, DPR and user agent, both on the direct URL and embedded in the Hub page. The Hub page swaps its iframe once while resolving the Space (a frame found too early detaches), and the Hub's own script logs one getBoundingClientRect error at that moment โ€” neither is the app's.
  • Not established: frame rate on a phone (SwiftShader's 7โ€“8 fps is this laptop's CPU rasterising a 648x960 buffer; a real Adreno/Mali/Apple GPU is a different measurement), warmth, battery, the mobile mic-permission and push-to-talk flow, and anything about Safari: the iPhone rows are a host limitation, not an iOS result. Real iOS Safari 14.5+ ships AudioContext, so the boot is expected to succeed there โ€” but it is expected, not observed.
  • Finding for the product (deferred, deferred-items.md): the boot constructs the AudioContext before the stage mounts (avatar/avatar.js:58-59), so a browser without Web Audio gets a blank canvas even though rendering needs no audio. Constructing it lazily on the first turn or mic press would let the avatar render everywhere and only degrade sound.

To close the row: open https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar on a real phone; record device, OS version, browser version, renders / smooth-or-slideshow, warmth, whether the mic permission and push-to-talk worked, and a rough turn time; then mark AVTR-01.

Audio on the phone (01-HUMAN-UAT gap 1, fixed in plan 01-11 โ€” retest after the redeploy). The owner's first phone visit (2026-09-06, revision b4d182b) rendered but was silent: the only AudioContext.resume() on that build ran inside playBuffer, seconds after the tap, and a gesture-gated browser refuses it. Reproduced in Chromium with the Space inside a cross-origin frame under --autoplay-policy=user-gesture-required (Chrome-on-Android's policy, and the Hub page's embed): the resume was refused 8.5 s after the tap and the page stayed on "thinkingโ€ฆ" (docs/evidence/2026-09-06-autoplay-embed-b4d182b.txt). Plan 01-11 resumes the context inside every tap. After the push:

  1. On the phone, open the Hub page, make sure the ring/silent switch is OFF (on iOS it mutes Web Audio entirely and looks identical to this bug) and the media volume is up.
  2. Tap Say hello once, before anything else. Expect the greeting to be audible within ~10 s and the status line to go "thinkingโ€ฆ" โ†’ "speakingโ€ฆ" โ†’ "ready".
  3. Then a typed turn and a Replay. Both should be audible without another unlock.
  4. If it stays on "thinkingโ€ฆ", or the status reads "audio is blocked by the browser - tap Say hello or Send once, then try again", record device / OS / browser and the status text here. From a laptop, python scripts/embed_autoplay_probe.py re-runs the cross-origin embed scenario against the live Space and prints whether speech-end arrived.

Lip-sync verification (AVTR-02)

Reviewed from captured frames and viseme telemetry by the execution agent, not by the owner (the owner asked for the manual rows to be performed on their behalf; this section is that review and says so). Space revision b4d182ba9d4ccf77c9910375c9842ddf29496a22, 2026-09-06T12:16:28Z, headed Chromium 151 (real GPU, so the render rate is a visitor's, not SwiftShader's), viewport 800x600 at device-scale 2.

Sentence (113 chars, 211 timeline events, decoded audio 20.565 s): ใ“ใ‚“ใซใกใฏใ€‚ไปŠๆ—ฅใฏใ„ใ„ๅคฉๆฐ—ใงใ™ใญใ€‚ๆ˜จๆ—ฅใ€้ง…ใฎ่ฟ‘ใใฎๅ–ซ่Œถๅบ—ใงใ€ๅ‹้”ใจใ‚ณใƒผใƒ’ใƒผใ‚’้ฃฒใฟใพใ—ใŸใ€‚ใกใ‚‡ใฃใจ้ซ˜ใ‹ใฃใŸใงใ™ใŒใ€ใจใฆใ‚‚็พŽๅ‘ณใ—ใ‹ใฃใŸใงใ™ใ€‚ๆ˜Žๆ—ฅใฏใ€ใใฃใจๅ›ณๆ›ธ้คจใงๆ—ฅๆœฌ่ชžใ‚’ๅ‹‰ๅผทใ—ใพใ™ใ€‚ใใ‚Œใ‹ใ‚‰ใ€ใ†ใกใฎ่ฟ‘ใใฎๅ…ฌๅœ’ใ‚’ใ€ใ‚†ใฃใใ‚Šๆญฉใใพใ™ใ€‚ โ€” chosen to carry ใ‚“ (ใ“ใ‚“ใซใกใฏ, ๅคฉๆฐ—, ้ฃฒใฟ, ๅ‹‰ๅผท, ๅ…ฌๅœ’), ใฃ (ใกใ‚‡ใฃใจ, ใใฃใจ, ใ‚†ใฃใใ‚Š), ใ€and ใ€‚pauses, ใงใ™/ใพใ—ใŸ/ใพใ™ endings (devoiced vowels) and several voiced ใ† (ใ†ใก, ใ‚†ใฃใใ‚Š, ๆญฉใใพใ™). Submit -> speech-start 22.06 s (synthesis_ms 20,700 for 20.6 s of audio: the Space's CPU synthesises long text at roughly 1.0x real time). Speech-start -> speech-end 20.66 s wall for a 20.565 s buffer.

Recording: docs/evidence/2026-09-06-lipsync-20s.webm (23 s, the utterance window of the page recording, VP8 800x600, 1.98 MB; the 57 s original was mostly the avatar idling while the server synthesised). Face crops during the utterance: docs/evidence/2026-09-06-lipsync-frames-01..10.png, each targeted at a long timeline event and labelled with the mouth read immediately before and after the capture (docs/evidence/2026-09-06-lipsync-telemetry.json ยง frames). Reference sheet of the five presets held at weight 1.0 on the same VRM by the same vrm-stage.js (local static server, same canvas size): docs/evidence/2026-09-06-viseme-ref-{closed,aa,ih,ou,ee,oh}.png. Telemetry: 2,063 samples of currentVisemes + the player's audio clock at ~100 samples/s (mean interval 10.0 ms) from speech-start to speech-end, plus the server's timeline for the same utterance โ€” all in the telemetry JSON.

Frame Target (timeline) Mouth read after the capture Reads as
01 ใ„ @0.50 s (96 ms) ih 0.20 (fading in) mostly closed, opening
02 ใ‚ @0.85 s (160 ms) aa 0.69 open oval โ€” ใ‚
03 ใ„ @1.79 s (96 ms) ih 0.88 (aa 0.11 residual) open, on the ใ‚->ใ„ cross-fade
04 ใˆ @2.69 s (171 ms) ee 0.92 flat, wide, teeth visible โ€” ใˆ
05 ใŠ @3.50 s (160 ms) oh 1.00 tall rounded โ€” ใŠ
06 ใ‚ @8.09 s (160 ms) aa 1.00 (captured 0.9 s late, at 9.09 s โ€” a later ใ‚; the screenshot call stalled) open oval โ€” ใ‚
07 ใ† @17.10 s (117 ms) ou 0.60 small, pursed โ€” ใ†
08 pause @18.79 s (459 ms, before ใใ‚Œใ‹ใ‚‰) all 0 (oh 0.19 decaying at the read) shut
09 ใ† @19.32 s (107 ms) ou 0.87 small, pursed โ€” ใ†
10 devoiced ใ† @20.42 s (ใพใ™, weight 0.5) ou 0.21 nearly shut, slight purse

The four judgements

  1. Do ใ‚ / ใ„ / ใ† visibly differ? โ€” YES. Each preset is driven independently to full weight: stage peak-hold over the utterance aa 1.00, ih 1.00, ou 0.92, ee 1.00, oh 1.00 (ou tops out at 0.92 on a 117 ms ใ† and at 0.66โ€“0.69 on 53 ms ones โ€” the 50 ms cross-fade in lipsync.js is the limit, not the timeline). On this VRM ใ‚ is a tall open oval (ref-aa, frames 02/06), ใ„ a flat wide opening with teeth (ref-ih), ใ† a small pursed mouth (ref-ou, frames 07/09) โ€” three shapes no one would confuse, so this is not amplitude flapping. Two honest caveats. (a) On this model ihโ‰ˆee and aaโ‰ˆoh are visually near-identical pairs (reference sheet): the five presets read as three groups. That is a property of the sample character, recorded in docs/ASSETS.md, not of the pipeline. (b) Both deployed ใ„ captures (01, 03) landed on a cross-fade โ€” ใ„ moras in this sentence are 96 ms, shorter than the ~110 ms a screenshot round trip takes โ€” so the clean ใ„ is the reference frame; the telemetry shows ih at 1.00 with aa 0 on the 21 full-weight ใ„/ใ—/ใก moras.
  2. Does the mouth close on ใ‚“, ใฃ and pauses? โ€” YES. The timeline has 113 closed events (N, cl, pau, and consonant onsets). Every pause โ‰ฅ 139 ms (the ใ€/ใ€‚ rests, 245โ€“459 ms; 12 events) reached a sampled mouth sum of exactly 0. The mora-length closures โ€” ใ‚“ and ใฃ, 85โ€“107 ms, 17 events โ€” reached a median 0.02 of open (min 0, max 0.44): shut to the eye, the residual being the 50 ms cross-fade's tail on the shortest ones. The 32โ€“75 ms consonant onsets inside moras dip to a median 0.20 and reopen, which is how speech looks. Frame 08 is the 459 ms pause: shut.
  3. Is the END of the sentence as well synced as the start? โ€” YES, no drift. The timeline ends at 20.565333 s and the decoded AudioBuffer lasts 20.565333 s โ€” a 0.000 ms difference over 211 events. Per-event agreement (the dominant sampled viseme at 60 % through every vowel event โ‰ฅ 90 ms) is 19/19 in the first third, 19/19 in the second, 22/22 in the last โ€” zero mismatches. The last vowel (the devoiced ใ† of ใพใ™, 20.363โ€“20.469 s) released on time and the last sample with the mouth open is at clock 20.550 s, 15 ms before speech-end; the first vowel event is at 0.192 s and the first open sample at 0.198 s. Clock discipline held: the player reads audioCtx.currentTime - scheduledStart, and the last third looks exactly like the first.
  4. Does the mouth freeze on ใงใ™ / ใ—ใŸ (devoiced vowels)? โ€” NO, it moves at half weight as designed. Eight devoiced moras (weight 0.5 in the timeline: ใ™ ร—5 incl. ใงใ™/ใ—ใพใ™/ใพใ™, ใ—/ใก ร—3 in ใพใ—ใŸ/็พŽๅ‘ณใ—ใ‹ใฃใŸ/ใกใ‚‡ใฃใจ) opened to sampled peaks of 0.35โ€“0.45 against the 0.5 target (attack-limited on 53โ€“107 ms moras) and were never held open (0 past 0.75) and never frozen shut (0 at 0.00 while a vowel was due). Frame 10 is the final devoiced ใพใ™: nearly shut with a slight purse, which is the intended look. The uppercase-vowel path is intact.

Verdict of this review: all four criteria hold on the deployed Space; no drift, no frozen mouth, no amplitude flapping. Escalation per the plan is therefore not triggered. The judgement is the execution agent's from the evidence above; the owner has the recording and the frames to overrule it.

Findings recorded, not fixed (Phase 1 ships as is): (a) under headless SwiftShader at ~15 fps โ€” a first, discarded run โ€” a lone 53 ms ใ† was skipped entirely (peak-hold ou 0.5, from a devoiced mora only): a mora shorter than one render frame can be dropped when the browser renders below 19 fps. Real GPUs render at 60+; the emulated-mobile rows above (6.7 fps under SwiftShader) say what a genuinely slow device would look like, which is one more reason the real-device row stays open. (b) The 50 ms cross-fade caps how far a 53 ms mora opens (0.66) โ€” a tunable, and a design choice this review did not second-guess.

Interpretation

  • p50 dispatch -> speech-start is 7107 ms (over the ~1500 ms VOIC-05 perceived-response target; OVER the 2500 ms re-plan trigger).

RE-PLAN TRIGGER FIRED. p50 7107 ms exceeds 2500 ms. Per 01-RESEARCH.md Open Question 6 - "treat p50 > ~2.5 s as a re-plan trigger, not a phase failure" - this is a decision for the roadmap, not a failing test: the criterion for Phase 1 is "p50/p95 measured and recorded", which this file is.

  • What dominates: VOICEVOX synthesis on the Space's CPU - server synthesis_ms p50 6422 ms is 90 % of the p50 turn; audio_query (1.4 ms), timeline (0.36 ms) and encode (0.17 ms) are noise. Network + Gradio round trip above the server total: 405 ms at p50; browser decode + schedule after the response: 278 ms at p50.
  • Cheapest improvement (recorded, NOT implemented in Phase 1): the same sentence synthesises 3โ€“6x faster on a developer laptop than on the Space (plan 01-09), so the container's CPU share is the lever - check cpu_num_threads on the Space's Synthesizer and whether the ZeroGPU container's CPU allocation is the ceiling; beyond that, VOICEVOX CORE 0.17's streaming synthesis would let speech start before the whole utterance is rendered, and per-sentence synthesis of long replies would bound the first-sample latency by the first sentence rather than the whole reply.

Notes

  • Phase 1 has no LLM. These numbers are NOT predictive of Phase 3, which adds one.
  • Space sleeps after 48 h (gcTimeout 172800), so a cold start is the default first-visit experience.
  • Numbers live in git because Space disk is ephemeral. Continuous telemetry is deliberately deferred to Phase 4/6, which owns the privacy story.
  • Headless Chromium renders through SwiftShader on the client; the client-side decode + schedule share above is therefore an upper bound on what a GPU-rendered visitor sees.

Raw samples (dispatch -> speech-start, ms, in run order)

Turn Sentence # dispatch->speech dispatch->response synthesis lastTurnMs
1 1 6538 6262 5997 6538
2 2 9639 9361 8902 9639
3 3 9661 9390 9165 9661
4 4 6674 6401 6117 6674
5 5 7722 7441 7100 7722
6 1 4589 4326 4134 4589
7 2 10677 10406 10099 10677
8 3 10713 10443 10120 10713
9 4 7113 6836 6503 7113
10 5 6365 6086 5787 6365
11 1 6417 6144 5860 6417
12 2 8740 8460 8201 8740
13 3 10858 10592 10118 10858
14 4 7107 6829 6422 7107
15 5 8437 8157 7936 8437
16 1 6662 6390 6097 6662
17 2 7552 7282 6998 7553
18 3 9634 9360 9083 9634
19 4 7113 6847 6646 7113
20 5 6635 6364 6089 6635
21 1 6706 6434 6081 6706
22 2 8142 7868 7510 8142
23 3 10552 10277 9963 10552
24 4 6828 6552 6127 6828
25 5 4872 4601 4284 4872
26 1 6188 5912 5713 6188
27 2 5572 5297 4900 5572
28 3 10207 9924 9653 10207
29 4 6775 6505 6224 6775
30 5 6337 6068 5835 6337

Replays (dispatch -> speech-start, ms): 0, 0, 0, 0, 0