Spaces:
Running on Zero
Download docs/ASR-TIERS.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 11.3 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/docs/ASR-TIERS.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/docs/ASR-TIERS.md
-
curl -L -o ASR-TIERS.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/docs/ASR-TIERS.md
ASR tiers and model choice (VOIC-02)
Every number on this page was measured on this project's own audio, by
tests/fixtures/make_asr_ab_results.py driving avatar/asr-harness.html in a real
browser. The raw data is tests/fixtures/asr_ab_results.json. Nothing here is quoted
from a model card.
| Measured | 2026-08-27 |
| Browser | Chromium 151.0.0.0 (Playwright), Windows NT 10.0 Win64 |
| GPU adapter | intel xe-lpg (WebGPU available, headed Chromium only) |
| Runtime | https://esm.sh/@huggingface/transformers@4.2.0 |
| Clips | speech_ja.wav, speech_ja_long.wav, speech_ja_slow.wav β VOICEVOX γγγ γγ, plan 01-04 |
Tier design
| Tier | Path | GPU quota | Availability |
|---|---|---|---|
| A | @huggingface/transformers 4.2.0, device: 'webgpu' |
zero | Chrome 113+ desktop, Chrome 121+ Android 12+, Safari macOS 26 / iOS 26, Firefox 141+ |
| B | same library, device: 'wasm' |
zero | everywhere WASM+SIMD works; several times slower |
| C | type instead of speak (VOIC-04) | zero | universal β the true floor |
| D | hosted whisper-large-v3-turbo on accelerated hardware | burns visitor quota | NOT built, and must not be. It would consume the exact resource SC-4 tests the loop's survival without |
tests/test_transport_seam.py::test_no_remote_asr_path_exists fails the build if tier D
ever appears in avatar/.
Gate thresholds (measured against the committed fixtures)
Read straight off the WAV files, independent of any microphone: 20 ms frames, peak frame RMS / median frame RMS. Reproduce with window.__measureClip(url) in the harness.
| Fixture | duration | RMS | peak/median frame RMS | gate verdict | condition that fired |
|---|---|---|---|---|---|
silence_30s.wav |
30.000 s | 0.00000 | β (median is 0) | REJECT | rms-floor |
cafe_noise_30s.wav |
30.000 s | 0.05760 | 1.961 | REJECT | envelope-modulation |
speech_ja.wav |
1.056 s | 0.07153 | 10.711 | ACCEPT | β |
speech_ja_long.wav |
5.504 s | 0.05515 | 7.509 | ACCEPT | β |
speech_ja_slow.wav |
7.381 s | 0.04918 | 6.368 | ACCEPT | β |
Thresholds: 300 ms duration floor, 0.01 RMS floor, 2.5 modulation floor.
The 2.5 figure is the whole reason the third condition exists. The cafΓ© fixture is written at β24.8 dBFS β roughly six times the RMS floor β so an RMS-only gate passes it comfortably. It is rejected because steady broadband noise is not modulated: 1.961 against speech's 6.4β10.7. The threshold sits with ~28% headroom below the noise and a 2.5Γ margin below the least-modulated speech clip.
The same gate, measured through a real microphone
Chromium fed the fixture as its capture device
(--use-file-for-fake-audio-capture), which is the path a learner actually exercises.
| Fixture | hold | browser audio processing | duration | RMS | modulation | verdict |
|---|---|---|---|---|---|---|
silence_30s.wav |
3000 ms | on | 2731 ms | 0.00000 | β | REJECT rms-floor |
cafe_noise_30s.wav |
3000 ms | on | 2731 ms | 0.00547 | 2.344 | REJECT rms-floor |
speech_ja.wav |
1100 ms | on | 768 ms | 0.15451 | 6.694 | ACCEPT |
silence_30s.wav |
3000 ms | off | 2731 ms | 0.00000 | β | REJECT rms-floor |
cafe_noise_30s.wav |
3000 ms | off | 2304 ms | 0.04409 | 1.360 | REJECT envelope-modulation |
speech_ja.wav |
1400 ms | off | 1195 ms | 0.06717 | 15.012 | ACCEPT |
speech_ja.wav |
150 ms | off | 0 ms | β | β | REJECT duration-floor |
Chromium's WebRTC noise suppression is doing part of the gate's job for us β it drops
the cafΓ© fixture from 0.0577 to 0.0055 RMS, below the floor, so the modulation condition
is never consulted. That is a good outcome in production and a useless one in a test,
because it would leave the third condition unexercised while the suite looked green.
tests/e2e/test_asr_standalone.py therefore drives the harness with ?processing=off and
verifies the gate in its pessimistic configuration β the raw microphone every browser
without WebRTC processing hands us anyway. Production keeps all three processing flags on.
The hallucination blocklist
Applied to the transcript, not the audio, and only under 1.5 s. Seven known Japanese
subtitle-boilerplate strings, listed in avatar/mic.js.
Deliberately not blocklisted: the bare polite form γγγγγ¨γγγγγΎγγγ. It is an
ordinary thing a learner says out loud, and swallowing a real utterance is a worse failure
than echoing one hallucination.
tests/test_transport_seam.py::test_blocklist_does_not_swallow_ordinary_japanese exists
specifically so a later "helpful" edit cannot add it back.
Model A/B
CER is against the exact strings plan 01-04 handed to VOICEVOX, so the ground truth is
known rather than transcribed. "CER no-punct" strips γγ and spaces from both sides,
because punctuation is a rendering choice rather than a mishearing. "infer" is the median
of the three clips. "first load" is the Content-Length of the two ONNX files the runtime
fetches.
| Model | dtype | device asked | tier used | first load | cold load | warm load | infer (p50) | CER | CER no-punct |
|---|---|---|---|---|---|---|---|---|---|
whisper-base |
q4 | wasm | wasm | 135.8 MB | 4064 ms | 1617 ms | 3636 ms | 0.062 | 0.023 |
whisper-base |
q4 | webgpu | webgpu | 135.8 MB | 4415 ms | 2152 ms | 779 ms | 0.062 | 0.023 |
whisper-small |
q4 | wasm | wasm | 285.5 MB | 7355 ms | 3908 ms | 15485 ms | 0.083 | 0.023 |
whisper-small |
q4 | webgpu | webgpu | 285.5 MB | 12846 ms | 4480 ms | 1934 ms | 0.083 | 0.023 |
whisper-large-v3-turbo |
q4f16 | webgpu | webgpu | 537.4 MB | 29413 ms | 4055 ms | 2858 ms | 0.063 | 0.000 |
whisper-large-v3-turbo |
q4 | wasm | wasm | 723.9 MB | 20692 ms | 4113 ms | 68200 ms | 0.063 | 0.000 |
whisper-base |
q8 | wasm | SESSION FAILED | 73.3 MB | 3471 ms | β | β | β | β |
whisper-base |
q8 | webgpu | webgpu | 73.3 MB | 3474 ms | 2021 ms | 4679 ms | 0.021 | 0.000 |
whisper-small |
q8 | wasm | SESSION FAILED | 237.5 MB | 5867 ms | β | β | β | β |
Transcripts, per clip, in tests/fixtures/asr_ab_results.json. The single error
whisper-base makes on this material is writing δ»ζ₯γ―θ―γε€©ζ° where VOICEVOX was given
δ»ζ₯γ―γγε€©ζ° β the same word, the same reading, a different orthography. It is not a
mishearing, which is why the punctuation-insensitive CER is 0.023 rather than 0.062.
The finding that changed the default: q8 does not work on the WASM tier
01-RESEARCH.md recommends onnx-community/whisper-base with dtype: 'q8'. Measured,
that configuration cannot create an inference session on the WASM backend at all:
Can't create a session. ERROR_CODE: 1, ERROR_MESSAGE: qdq_actions.cc:137
TransposeDQWeightsForMatMulNBits Missing required scale:
model.decoder.embed_tokens.weight_merged_0_scale
for node: model.decoder.embed_tokens.weight_transposed_DequantizeLinear
The same error appears for whisper-base, whisper-small and whisper-large-v3-turbo,
and for the int8 and quantized aliases. Every one of them loads fine on WebGPU,
which is precisely how this would have shipped: green on a developer machine with a GPU,
dead on every browser without one β the tier the fallback exists to serve. q4 is the only
quantisation measured working on both tiers, so q4 is the default.
The other finding: the documented fallback does not fall back
01-RESEARCH.md prescribes "feature-detect navigator.gpu, request webgpu when present,
catch initialisation failure, re-instantiate on WASM." Measured, that does not work.
When navigator.gpu exists but requestAdapter() resolves to null β a headless browser,
a blocklisted GPU, a browser started without GPU access β calling
pipeline(..., { device: 'webgpu' }) throws and poisons the ONNX Runtime Web backend
registry for the rest of the page. The subsequent pipeline(..., { device: 'wasm' }) for
the same model then fails with the identical WebGPU error:
no available backend found. ERR: [webgpu] Error: Failed to get GPU adapter.
So avatar/asr.js probes for an actual adapter first (probeWebGpu()) and only issues
the WebGPU call when one exists. The doomed call is never made, so the registry is never
poisoned. The try/catch remains as a second line of defence for failures a probe cannot
predict, such as an out-of-memory adapter.
tests/test_transport_seam.py::test_asr_webgpu_fallback_is_real_not_aspirational fails if
requestAdapter ever disappears from avatar/asr.js.
Default chosen
onnx-community/whisper-base at q4, on WebGPU when an adapter is available and on
WASM when it is not.
- It is the only measured configuration that runs on both tiers.
q8is smaller (73.3 MB) and more accurate (CER 0.021), and it is disqualified outright because it cannot create a session on WASM. - 135.8 MB first load, then ~1.6β2.2 s to warm from the browser cache.
- 779 ms median inference on WebGPU, 3636 ms on WASM. Both leave room inside VOIC-05's
perceived-response budget;
whisper-smallat 15.5 s on WASM does not. - CER 0.023 ignoring punctuation, and its one error is an orthography choice for the same word rather than a misheard one.
whisper-small is rejected on the numbers: 2.1Γ the bytes of whisper-base, an
identical punctuation-insensitive CER of 0.023, and 4.3Γ the WASM inference time. It is
dominated on every axis that matters.
"Better accuracy" toggle: onnx-community/whisper-large-v3-turbo at q4f16, WebGPU
only. CER 0.000 on this material and 2858 ms median inference β but a 537.4 MB first
load, and 68 s per clip if it ever ran on WASM, which is why it must never be offered
without an adapter. Both entries live in MODELS in avatar/asr.js; plan 01-08 owns the
control that switches between them.
Honest caveats
- These clips are VOICEVOX-synthesised speech: cleaner, more regular and better articulated than a learner speaking into a laptop microphone in a room with a fan. Every CER above is a best case and overstates real-world accuracy. Real-speech confirmation is a manual verification in plan 01-10.
- Phase 1's bar is "their transcript appears", not "the transcript is right". If the chosen default turns out to be embarrassing on real speech, that is a Phase 2/3 input, not a Phase 1 failure.
- Three clips and two distinct sentences is a small sample. It is enough to rule
whisper-smallout and to disqualifyq8, and it is not enough to claim a CER figure with confidence intervals. - Load and inference times are from one Windows laptop with an
intel xe-lpgiGPU on one network. A visitor on a phone will see different numbers, particularly on WASM. - WebGPU numbers required headed Chromium. Headless Chromium exposes
navigator.gpuand then returns a null adapter, so a headless re-run of the A/B records every candidate as WASM. That is true, and useless as a comparison. - Tier D (hosted ASR on accelerated hardware) is deliberately not built: it would consume visitor GPU quota and contradict SC-4.