Spaces:
Running on Zero
Download docs/LATENCY.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 21.4 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/docs/LATENCY.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/docs/LATENCY.md
-
curl -L -o LATENCY.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/docs/LATENCY.md
Turn latency (VOIC-05)
Measured: 2026-09-06T04:32:30Z
Space: WolfDavid/japanese-learning-avatar Revision: c911d74b378981e6ba0d0111eb570ed3b20927db URL measured: https://wolfdavid-japanese-learning-avatar.hf.space
Hardware: zero-a10g (stage RUNNING) DISABLE_GPU: 1
Browser: Chromium 151.0.7922.34 (Playwright, headless) Client: Windows 11, public internet, one visitor
Method: N=30 warm text turns after 1 throwaway warm-up (ใใใซใกใฏ), 5 replay turns, five sentences of 5โ36 moras in a fixed cycle (tests/e2e/test_latency_harness.py). Measured from the page's own performance.mark entries; server stages from __debug.lastStageTimings.
p95 computed by nearest-rank (the ceil(0.95ยทN)-th smallest observed value; p50 likewise), so no percentile is an interpolated value that was never observed.
Generated by the harness; the Cold start, Mobile and Lip-sync sections are the owner's and are carried forward verbatim when the harness re-runs.
Warm turns (N=30)
| Stage | p50 (ms) | p95 (ms) | min | max |
|---|---|---|---|---|
| dispatch -> speech-start (THE NUMBER) | 7107 | 10713 | 4589 | 10858 |
| dispatch -> response | 6829 | 10443 | 4326 | 10592 |
| server: audio_query | 1 | 2 | 1 | 2 |
| server: synthesis | 6422 | 10118 | 4134 | 10120 |
| server: timeline | 0 | 1 | 0 | 1 |
| server: encode | 0 | 1 | 0 | 1 |
| server: total | 6424 | 10121 | 4136 | 10124 |
Per sentence (dispatch -> speech-start, ms; each sentence was spoken 6 times):
| # | Sentence | Moras | p50 | min | max | synthesis p50 |
|---|---|---|---|---|---|---|
| 1 | ใใใซใกใฏ | 5 | 6417 | 4589 | 6706 | 5860 |
| 2 | ใฏใใใพใใฆใใใใใใ้กใใใพใใ | 17 (hand) | 8142 | 5572 | 10677 | 7510 |
| 3 | ไปๆฅใฏใใๅคฉๆฐใงใใใใๅ ฌๅใๆฃๆญฉใใฆใใใ่ฒทใ็ฉใซ่กใใพใใใ | 36 | 10207 | 9634 | 10858 | 9653 |
| 4 | ้ง ใฏใฉใใงใใใ | 8 (hand) | 6828 | 6674 | 7113 | 6224 |
| 5 | ๆฅๆฌ่ชใๅๅผทใใฆใใพใใ | 14 (hand) | 6365 | 4872 | 8437 | 5835 |
Replay (N=5, client-side, zero network)
| Stage | p50 (ms) | p95 (ms) | min | max |
|---|---|---|---|---|
| dispatch -> speech-start | 0 | 0 | 0 | 0 |
Replay p50 0 ms vs warm-turn p50 7107 ms: the cache is free, the round trip is the cost.
Cold start
| Measurement | Value | How obtained |
|---|---|---|
| Space PAUSED -> first HTTP 200 serving the Gradio app | 18.0 s after restart_space โ stage BUILDING at +0.2 s, APP_STARTING at +2.2 s, RUNNING at +18.6 s |
Hub API: HfApi.pause_space() (stage PAUSED 0.1 s later; the URL answered 503, 51 bytes, while paused), then HfApi.restart_space(factory_reboot=False); GET / polled every 2 s until a 200 whose body names gradio. Method: pause + restart via the API, not a factory rebuild |
Page load -> avatar ready -> first rendered frame, first visit after the restart |
ready +38.1 s, first rendered frame +39.6 s from the restart (19.4 s / 20.8 s after navigation, which began at +18.7 s) | headless Chromium 151 navigated the moment the first 200 arrived; AVATAR_READY, then breathValue !== 0; threeInstanceCount 1, armDown -0.932 / -0.932, mountCount 1, transport inline |
| First turn after boot (typed ใใใซใกใฏ -> speech-start) | 7.2 s after Enter (+46.9 s from the restart); lastTurnMs 6600, synthesis_ms 6076; the second turn in the same session: 6368 / 5796 |
the first typed turn on the freshly restarted container. The synthesiser is loaded at app start, so the first turn pays no model load โ the plan's "synthesiser load dominates" prediction did not materialise; the first turn is the same CPU synthesis as every warm turn |
Page load -> avatar ready (client-side, no backend), Space already RUNNING |
harness run: ready 4.2 s, first rendered frame 5.7 s; lip-sync run (headed, GPU): ready 6.7 s; mobile-emulation runs 9.9โ12.4 s (SwiftShader at 2.6โ3x DPR) | wait_for_avatar_ready; plan 01-05 recorded 4.1โ8.3 s warm with tutor.vrm 3.0โ6.8 s of it, 01-09 4.0โ16.8 s; Space rebuild-to-RUNNING 106 s (01-05); warm wake 0.2โ2.7 s |
Revision live during the cold-start test: b4d182ba9d4ccf77c9910375c9842ddf29496a22 (read from the Hub before the pause, during the restart and after: unchanged). Measured 2026-09-06T12:09:04Z; performed by the execution agent through the Hub API on the owner's instruction ("yes to all five"), not from the Space Settings page; the Space was confirmed RUNNING afterwards and left that way. Raw record with every stage transition and HTTP poll: docs/evidence/2026-09-06-cold-start.json.
What this is and is not: a pause/restart reuses the built image, which is why it reaches RUNNING in 18 s where the plan 01-05 rebuild took 106 s. The 48-hour SLEEPING -> wake path (what a first visitor gets after two idle days) was not forced and remains unmeasured as such; per 01-RESEARCH.md "cold = SLEEPING or manually paused/restarted", the restart is the sanctioned proxy. The recruiter-facing sum on this path is ~40 s to a moving avatar and ~47 s to the first spoken reply from a paused container, of which the browser's own VRM load (19โ21 s here over the public internet, incl. the 10.3 MiB tutor.vrm) is the larger half.
Mobile
EMULATED โ real-device verification still open. No real phone was available; these rows are Playwright device descriptors driven by the execution agent on the owner's instruction, and they say nothing about a real mobile GPU, thermals or the mobile microphone-permission flow. AVTR-01's manual row therefore stays open and the requirement is not marked on this evidence. Space revision b4d182ba9d4ccf77c9910375c9842ddf29496a22, 2026-09-06T12:32:52Z; raw record docs/evidence/2026-09-06-mobile-emulation.json; screenshots docs/evidence/2026-09-06-mobile-emulated-*.png.
| Device (EMULATED) | OS | Browser | URL | Renders? | Observed smoothness | Push-to-talk | Turn latency | Page -> ready / first frame |
|---|---|---|---|---|---|---|---|---|
Playwright Pixel 7 descriptor (412x839 CSS px, DPR 2.625, touch, mobile UA) |
Android 14 UA string | Chromium 151.0.7922.34 headless, WebGL 2 via SwiftShader (no GPU) | direct hf.space |
yes โ ready, first frame, threeInstanceCount 1, mountCount 1, armDown -0.932 / -0.932, canvas 324x480 CSS (648x960 buffer) |
7.7 fps over 5 s (rAF = getDebug samples) โ a slideshow, but that is the software rasteriser on this laptop, not a phone GPU | not tested (mobile mic permission needs a device) | typed ใใใซใกใฏ -> speech-start 8.0 s after Enter; lastTurnMs 6612, synthesis_ms 5694; mouth peaks aa/ih/oh 1.0; audio 1.056 s |
4.5 s / 6.0 s |
same Pixel 7 descriptor |
Android 14 UA | same Chromium | huggingface.co Space page (inside the Hub's iframe.space-iframe, ?__theme=system) |
yes โ same numbers read inside the iframe: 1 instance, mount 1, arms posed, "Running on ZERO" badge above the avatar | 7.3 fps (SwiftShader) | not tested | speech-start 8.8 s after Enter; lastTurnMs 7354, synthesis_ms 6012; turn completed |
5.7 s / 8.0 s |
Playwright iPhone 14 descriptor (390x664, DPR 3, touch, iOS UA) |
iOS UA string | WebKit 26.5, Playwright's Windows build โ WebGL 2 present ("Apple GPU"), no Web Audio API (AudioContext and webkitAudioContext both undefined) |
direct and Hub page | NO โ not a device result. Boot aborts before the stage mounts: avatar boot failed: TypeError: undefined is not a constructor (evaluating 'new AudioCtor()'), status "the avatar failed to load - see the browser console", blank canvas (โฆiphone14-webkit-bootfail.png) |
โ | โ | โ | โ |
What the emulation does and does not establish:
- Established on the Android/Chromium profile: the responsive layout stacks the canvas above the controls at a phone width; the VRM, the idle life and a full typed turn (server round trip, decode, lip-synced playback) work at mobile viewport, DPR and user agent, both on the direct URL and embedded in the Hub page. The Hub page swaps its iframe once while resolving the Space (a frame found too early detaches), and the Hub's own script logs one
getBoundingClientRecterror at that moment โ neither is the app's. - Not established: frame rate on a phone (SwiftShader's 7โ8 fps is this laptop's CPU rasterising a 648x960 buffer; a real Adreno/Mali/Apple GPU is a different measurement), warmth, battery, the mobile mic-permission and push-to-talk flow, and anything about Safari: the iPhone rows are a host limitation, not an iOS result. Real iOS Safari 14.5+ ships
AudioContext, so the boot is expected to succeed there โ but it is expected, not observed. - Finding for the product (deferred,
deferred-items.md): the boot constructs theAudioContextbefore the stage mounts (avatar/avatar.js:58-59), so a browser without Web Audio gets a blank canvas even though rendering needs no audio. Constructing it lazily on the first turn or mic press would let the avatar render everywhere and only degrade sound.
To close the row: open https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar on a real phone; record device, OS version, browser version, renders / smooth-or-slideshow, warmth, whether the mic permission and push-to-talk worked, and a rough turn time; then mark AVTR-01.
Audio on the phone (01-HUMAN-UAT gap 1, fixed in plan 01-11 โ retest after the redeploy). The owner's first phone visit (2026-09-06, revision b4d182b) rendered but was silent: the only AudioContext.resume() on that build ran inside playBuffer, seconds after the tap, and a gesture-gated browser refuses it. Reproduced in Chromium with the Space inside a cross-origin frame under --autoplay-policy=user-gesture-required (Chrome-on-Android's policy, and the Hub page's embed): the resume was refused 8.5 s after the tap and the page stayed on "thinkingโฆ" (docs/evidence/2026-09-06-autoplay-embed-b4d182b.txt). Plan 01-11 resumes the context inside every tap. After the push:
- On the phone, open the Hub page, make sure the ring/silent switch is OFF (on iOS it mutes Web Audio entirely and looks identical to this bug) and the media volume is up.
- Tap Say hello once, before anything else. Expect the greeting to be audible within ~10 s and the status line to go "thinkingโฆ" โ "speakingโฆ" โ "ready".
- Then a typed turn and a Replay. Both should be audible without another unlock.
- If it stays on "thinkingโฆ", or the status reads "audio is blocked by the browser - tap Say hello or Send once, then try again", record device / OS / browser and the status text here. From a laptop,
python scripts/embed_autoplay_probe.pyre-runs the cross-origin embed scenario against the live Space and prints whetherspeech-endarrived.
Lip-sync verification (AVTR-02)
Reviewed from captured frames and viseme telemetry by the execution agent, not by the owner (the owner asked for the manual rows to be performed on their behalf; this section is that review and says so). Space revision b4d182ba9d4ccf77c9910375c9842ddf29496a22, 2026-09-06T12:16:28Z, headed Chromium 151 (real GPU, so the render rate is a visitor's, not SwiftShader's), viewport 800x600 at device-scale 2.
Sentence (113 chars, 211 timeline events, decoded audio 20.565 s):
ใใใซใกใฏใไปๆฅใฏใใๅคฉๆฐใงใใญใๆจๆฅใ้ง
ใฎ่ฟใใฎๅซ่ถๅบใงใๅ้ใจใณใผใใผใ้ฃฒใฟใพใใใใกใใฃใจ้ซใใฃใใงใใใใจใฆใ็พๅณใใใฃใใงใใๆๆฅใฏใใใฃใจๅณๆธ้คจใงๆฅๆฌ่ชใๅๅผทใใพใใใใใใใใใกใฎ่ฟใใฎๅ
ฌๅใใใใฃใใๆญฉใใพใใ
โ chosen to carry ใ (ใใใซใกใฏ, ๅคฉๆฐ, ้ฃฒใฟ, ๅๅผท, ๅ
ฌๅ), ใฃ (ใกใใฃใจ, ใใฃใจ, ใใฃใใ), ใand ใpauses, ใงใ/ใพใใ/ใพใ endings (devoiced vowels) and several voiced ใ (ใใก, ใใฃใใ, ๆญฉใใพใ). Submit -> speech-start 22.06 s (synthesis_ms 20,700 for 20.6 s of audio: the Space's CPU synthesises long text at roughly 1.0x real time). Speech-start -> speech-end 20.66 s wall for a 20.565 s buffer.
Recording: docs/evidence/2026-09-06-lipsync-20s.webm (23 s, the utterance window of the page recording, VP8 800x600, 1.98 MB; the 57 s original was mostly the avatar idling while the server synthesised). Face crops during the utterance: docs/evidence/2026-09-06-lipsync-frames-01..10.png, each targeted at a long timeline event and labelled with the mouth read immediately before and after the capture (docs/evidence/2026-09-06-lipsync-telemetry.json ยง frames). Reference sheet of the five presets held at weight 1.0 on the same VRM by the same vrm-stage.js (local static server, same canvas size): docs/evidence/2026-09-06-viseme-ref-{closed,aa,ih,ou,ee,oh}.png. Telemetry: 2,063 samples of currentVisemes + the player's audio clock at ~100 samples/s (mean interval 10.0 ms) from speech-start to speech-end, plus the server's timeline for the same utterance โ all in the telemetry JSON.
| Frame | Target (timeline) | Mouth read after the capture | Reads as |
|---|---|---|---|
| 01 | ใ @0.50 s (96 ms) | ih 0.20 (fading in) | mostly closed, opening |
| 02 | ใ @0.85 s (160 ms) | aa 0.69 | open oval โ ใ |
| 03 | ใ @1.79 s (96 ms) | ih 0.88 (aa 0.11 residual) | open, on the ใ->ใ cross-fade |
| 04 | ใ @2.69 s (171 ms) | ee 0.92 | flat, wide, teeth visible โ ใ |
| 05 | ใ @3.50 s (160 ms) | oh 1.00 | tall rounded โ ใ |
| 06 | ใ @8.09 s (160 ms) | aa 1.00 (captured 0.9 s late, at 9.09 s โ a later ใ; the screenshot call stalled) | open oval โ ใ |
| 07 | ใ @17.10 s (117 ms) | ou 0.60 | small, pursed โ ใ |
| 08 | pause @18.79 s (459 ms, before ใใใใ) | all 0 (oh 0.19 decaying at the read) | shut |
| 09 | ใ @19.32 s (107 ms) | ou 0.87 | small, pursed โ ใ |
| 10 | devoiced ใ @20.42 s (ใพใ, weight 0.5) | ou 0.21 | nearly shut, slight purse |
The four judgements
- Do ใ / ใ / ใ visibly differ? โ YES. Each preset is driven independently to full weight: stage peak-hold over the utterance
aa1.00,ih1.00,ou0.92,ee1.00,oh1.00 (outops out at 0.92 on a 117 ms ใ and at 0.66โ0.69 on 53 ms ones โ the 50 ms cross-fade inlipsync.jsis the limit, not the timeline). On this VRM ใ is a tall open oval (ref-aa, frames 02/06), ใ a flat wide opening with teeth (ref-ih), ใ a small pursed mouth (ref-ou, frames 07/09) โ three shapes no one would confuse, so this is not amplitude flapping. Two honest caveats. (a) On this modelihโeeandaaโohare visually near-identical pairs (reference sheet): the five presets read as three groups. That is a property of the sample character, recorded indocs/ASSETS.md, not of the pipeline. (b) Both deployed ใ captures (01, 03) landed on a cross-fade โ ใ moras in this sentence are 96 ms, shorter than the ~110 ms a screenshot round trip takes โ so the clean ใ is the reference frame; the telemetry showsihat 1.00 withaa0 on the 21 full-weight ใ/ใ/ใก moras. - Does the mouth close on ใ, ใฃ and pauses? โ YES. The timeline has 113
closedevents (N, cl, pau, and consonant onsets). Every pause โฅ 139 ms (the ใ/ใ rests, 245โ459 ms; 12 events) reached a sampled mouth sum of exactly 0. The mora-length closures โ ใ and ใฃ, 85โ107 ms, 17 events โ reached a median 0.02 of open (min 0, max 0.44): shut to the eye, the residual being the 50 ms cross-fade's tail on the shortest ones. The 32โ75 ms consonant onsets inside moras dip to a median 0.20 and reopen, which is how speech looks. Frame 08 is the 459 ms pause: shut. - Is the END of the sentence as well synced as the start? โ YES, no drift. The timeline ends at 20.565333 s and the decoded
AudioBufferlasts 20.565333 s โ a 0.000 ms difference over 211 events. Per-event agreement (the dominant sampled viseme at 60 % through every vowel event โฅ 90 ms) is 19/19 in the first third, 19/19 in the second, 22/22 in the last โ zero mismatches. The last vowel (the devoiced ใ of ใพใ, 20.363โ20.469 s) released on time and the last sample with the mouth open is at clock 20.550 s, 15 ms before speech-end; the first vowel event is at 0.192 s and the first open sample at 0.198 s. Clock discipline held: the player readsaudioCtx.currentTime - scheduledStart, and the last third looks exactly like the first. - Does the mouth freeze on ใงใ / ใใ (devoiced vowels)? โ NO, it moves at half weight as designed. Eight devoiced moras (weight 0.5 in the timeline: ใ ร5 incl. ใงใ/ใใพใ/ใพใ, ใ/ใก ร3 in ใพใใ/็พๅณใใใฃใ/ใกใใฃใจ) opened to sampled peaks of 0.35โ0.45 against the 0.5 target (attack-limited on 53โ107 ms moras) and were never held open (0 past 0.75) and never frozen shut (0 at 0.00 while a vowel was due). Frame 10 is the final devoiced ใพใ: nearly shut with a slight purse, which is the intended look. The uppercase-vowel path is intact.
Verdict of this review: all four criteria hold on the deployed Space; no drift, no frozen mouth, no amplitude flapping. Escalation per the plan is therefore not triggered. The judgement is the execution agent's from the evidence above; the owner has the recording and the frames to overrule it.
Findings recorded, not fixed (Phase 1 ships as is): (a) under headless SwiftShader at ~15 fps โ a first, discarded run โ a lone 53 ms ใ was skipped entirely (peak-hold ou 0.5, from a devoiced mora only): a mora shorter than one render frame can be dropped when the browser renders below 19 fps. Real GPUs render at 60+; the emulated-mobile rows above (6.7 fps under SwiftShader) say what a genuinely slow device would look like, which is one more reason the real-device row stays open. (b) The 50 ms cross-fade caps how far a 53 ms mora opens (0.66) โ a tunable, and a design choice this review did not second-guess.
Interpretation
- p50 dispatch -> speech-start is 7107 ms (over the ~1500 ms VOIC-05 perceived-response target; OVER the 2500 ms re-plan trigger).
RE-PLAN TRIGGER FIRED. p50 7107 ms exceeds 2500 ms. Per
01-RESEARCH.mdOpen Question 6 - "treat p50 > ~2.5 s as a re-plan trigger, not a phase failure" - this is a decision for the roadmap, not a failing test: the criterion for Phase 1 is "p50/p95 measured and recorded", which this file is.
- What dominates: VOICEVOX synthesis on the Space's CPU - server
synthesis_msp50 6422 ms is 90 % of the p50 turn;audio_query(1.4 ms),timeline(0.36 ms) andencode(0.17 ms) are noise. Network + Gradio round trip above the server total: 405 ms at p50; browser decode + schedule after the response: 278 ms at p50. - Cheapest improvement (recorded, NOT implemented in Phase 1): the same sentence synthesises 3โ6x faster on a developer laptop than on the Space (plan 01-09), so the container's CPU share is the lever - check
cpu_num_threadson the Space'sSynthesizerand whether the ZeroGPU container's CPU allocation is the ceiling; beyond that, VOICEVOX CORE 0.17's streaming synthesis would let speech start before the whole utterance is rendered, and per-sentence synthesis of long replies would bound the first-sample latency by the first sentence rather than the whole reply.
Notes
- Phase 1 has no LLM. These numbers are NOT predictive of Phase 3, which adds one.
- Space sleeps after 48 h (gcTimeout 172800), so a cold start is the default first-visit experience.
- Numbers live in git because Space disk is ephemeral. Continuous telemetry is deliberately deferred to Phase 4/6, which owns the privacy story.
- Headless Chromium renders through SwiftShader on the client; the client-side decode + schedule share above is therefore an upper bound on what a GPU-rendered visitor sees.
Raw samples (dispatch -> speech-start, ms, in run order)
| Turn | Sentence # | dispatch->speech | dispatch->response | synthesis | lastTurnMs |
|---|---|---|---|---|---|
| 1 | 1 | 6538 | 6262 | 5997 | 6538 |
| 2 | 2 | 9639 | 9361 | 8902 | 9639 |
| 3 | 3 | 9661 | 9390 | 9165 | 9661 |
| 4 | 4 | 6674 | 6401 | 6117 | 6674 |
| 5 | 5 | 7722 | 7441 | 7100 | 7722 |
| 6 | 1 | 4589 | 4326 | 4134 | 4589 |
| 7 | 2 | 10677 | 10406 | 10099 | 10677 |
| 8 | 3 | 10713 | 10443 | 10120 | 10713 |
| 9 | 4 | 7113 | 6836 | 6503 | 7113 |
| 10 | 5 | 6365 | 6086 | 5787 | 6365 |
| 11 | 1 | 6417 | 6144 | 5860 | 6417 |
| 12 | 2 | 8740 | 8460 | 8201 | 8740 |
| 13 | 3 | 10858 | 10592 | 10118 | 10858 |
| 14 | 4 | 7107 | 6829 | 6422 | 7107 |
| 15 | 5 | 8437 | 8157 | 7936 | 8437 |
| 16 | 1 | 6662 | 6390 | 6097 | 6662 |
| 17 | 2 | 7552 | 7282 | 6998 | 7553 |
| 18 | 3 | 9634 | 9360 | 9083 | 9634 |
| 19 | 4 | 7113 | 6847 | 6646 | 7113 |
| 20 | 5 | 6635 | 6364 | 6089 | 6635 |
| 21 | 1 | 6706 | 6434 | 6081 | 6706 |
| 22 | 2 | 8142 | 7868 | 7510 | 8142 |
| 23 | 3 | 10552 | 10277 | 9963 | 10552 |
| 24 | 4 | 6828 | 6552 | 6127 | 6828 |
| 25 | 5 | 4872 | 4601 | 4284 | 4872 |
| 26 | 1 | 6188 | 5912 | 5713 | 6188 |
| 27 | 2 | 5572 | 5297 | 4900 | 5572 |
| 28 | 3 | 10207 | 9924 | 9653 | 10207 |
| 29 | 4 | 6775 | 6505 | 6224 | 6775 |
| 30 | 5 | 6337 | 6068 | 5835 | 6337 |
Replays (dispatch -> speech-start, ms): 0, 0, 0, 0, 0