Spaces:
Running on Zero
Running on Zero
|
Download docs/ASR-TIERS.md from WolfDavid/japanese-learning-avatar: direct link, hf CLI and curl.
- Browser
- Download file 11.3 kB
-
https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/docs/ASR-TIERS.md
- Command line
-
hf download hf://spaces/WolfDavid/japanese-learning-avatar@706670a3ece0683b9770f0b7db36e4daa955f91e/docs/ASR-TIERS.md
-
curl -L -o ASR-TIERS.md https://huggingface.co/spaces/WolfDavid/japanese-learning-avatar/resolve/706670a3ece0683b9770f0b7db36e4daa955f91e/docs/ASR-TIERS.md
11.3 kB
| # ASR tiers and model choice (VOIC-02) | |
| Every number on this page was measured on this project's own audio, by | |
| `tests/fixtures/make_asr_ab_results.py` driving `avatar/asr-harness.html` in a real | |
| browser. The raw data is `tests/fixtures/asr_ab_results.json`. Nothing here is quoted | |
| from a model card. | |
| | | | | |
| |---|---| | |
| | Measured | 2026-08-27 | | |
| | Browser | Chromium 151.0.0.0 (Playwright), Windows NT 10.0 Win64 | | |
| | GPU adapter | `intel xe-lpg` (WebGPU available, **headed** Chromium only) | | |
| | Runtime | `https://esm.sh/@huggingface/transformers@4.2.0` | | |
| | Clips | `speech_ja.wav`, `speech_ja_long.wav`, `speech_ja_slow.wav` β VOICEVOX γγγ γγ, plan 01-04 | | |
| --- | |
| ## Tier design | |
| | Tier | Path | GPU quota | Availability | | |
| |---|---|---|---| | |
| | **A** | `@huggingface/transformers` 4.2.0, `device: 'webgpu'` | **zero** | Chrome 113+ desktop, Chrome 121+ Android 12+, Safari macOS 26 / iOS 26, Firefox 141+ | | |
| | **B** | same library, `device: 'wasm'` | **zero** | everywhere WASM+SIMD works; several times slower | | |
| | **C** | type instead of speak (VOIC-04) | zero | universal β the true floor | | |
| | **D** | hosted whisper-large-v3-turbo on accelerated hardware | **burns visitor quota** | **NOT built, and must not be.** It would consume the exact resource SC-4 tests the loop's survival without | | |
| `tests/test_transport_seam.py::test_no_remote_asr_path_exists` fails the build if tier D | |
| ever appears in `avatar/`. | |
| --- | |
| ## Gate thresholds (measured against the committed fixtures) | |
| Read straight off the WAV files, independent of any microphone: 20 ms frames, `peak frame | |
| RMS / median frame RMS`. Reproduce with `window.__measureClip(url)` in the harness. | |
| | Fixture | duration | RMS | peak/median frame RMS | gate verdict | condition that fired | | |
| |---|---|---|---|---|---| | |
| | `silence_30s.wav` | 30.000 s | 0.00000 | β (median is 0) | **REJECT** | `rms-floor` | | |
| | `cafe_noise_30s.wav` | 30.000 s | 0.05760 | **1.961** | **REJECT** | `envelope-modulation` | | |
| | `speech_ja.wav` | 1.056 s | 0.07153 | **10.711** | **ACCEPT** | β | | |
| | `speech_ja_long.wav` | 5.504 s | 0.05515 | 7.509 | **ACCEPT** | β | | |
| | `speech_ja_slow.wav` | 7.381 s | 0.04918 | 6.368 | **ACCEPT** | β | | |
| Thresholds: **300 ms** duration floor, **0.01** RMS floor, **2.5** modulation floor. | |
| The 2.5 figure is the whole reason the third condition exists. The cafΓ© fixture is written | |
| at β24.8 dBFS β roughly **six times** the RMS floor β so an RMS-only gate passes it | |
| comfortably. It is rejected because steady broadband noise is not modulated: 1.961 against | |
| speech's 6.4β10.7. The threshold sits with ~28% headroom below the noise and a 2.5Γ | |
| margin below the least-modulated speech clip. | |
| ### The same gate, measured through a real microphone | |
| Chromium fed the fixture as its capture device | |
| (`--use-file-for-fake-audio-capture`), which is the path a learner actually exercises. | |
| | Fixture | hold | browser audio processing | duration | RMS | modulation | verdict | | |
| |---|---|---|---|---|---|---| | |
| | `silence_30s.wav` | 3000 ms | on | 2731 ms | 0.00000 | β | REJECT `rms-floor` | | |
| | `cafe_noise_30s.wav` | 3000 ms | on | 2731 ms | **0.00547** | 2.344 | REJECT `rms-floor` | | |
| | `speech_ja.wav` | 1100 ms | on | 768 ms | 0.15451 | 6.694 | ACCEPT | | |
| | `silence_30s.wav` | 3000 ms | **off** | 2731 ms | 0.00000 | β | REJECT `rms-floor` | | |
| | `cafe_noise_30s.wav` | 3000 ms | **off** | 2304 ms | 0.04409 | **1.360** | REJECT `envelope-modulation` | | |
| | `speech_ja.wav` | 1400 ms | **off** | 1195 ms | 0.06717 | 15.012 | ACCEPT | | |
| | `speech_ja.wav` | 150 ms | **off** | 0 ms | β | β | REJECT `duration-floor` | | |
| **Chromium's WebRTC noise suppression is doing part of the gate's job for us** β it drops | |
| the cafΓ© fixture from 0.0577 to 0.0055 RMS, below the floor, so the modulation condition | |
| is never consulted. That is a good outcome in production and a useless one in a test, | |
| because it would leave the third condition unexercised while the suite looked green. | |
| `tests/e2e/test_asr_standalone.py` therefore drives the harness with `?processing=off` and | |
| verifies the gate in its **pessimistic** configuration β the raw microphone every browser | |
| without WebRTC processing hands us anyway. Production keeps all three processing flags on. | |
| ### The hallucination blocklist | |
| Applied to the transcript, not the audio, and only under **1.5 s**. Seven known Japanese | |
| subtitle-boilerplate strings, listed in `avatar/mic.js`. | |
| Deliberately **not** blocklisted: the bare polite form γγγγγ¨γγγγγΎγγγ. It is an | |
| ordinary thing a learner says out loud, and swallowing a real utterance is a worse failure | |
| than echoing one hallucination. | |
| `tests/test_transport_seam.py::test_blocklist_does_not_swallow_ordinary_japanese` exists | |
| specifically so a later "helpful" edit cannot add it back. | |
| --- | |
| ## Model A/B | |
| CER is against the exact strings plan 01-04 handed to VOICEVOX, so the ground truth is | |
| known rather than transcribed. "CER no-punct" strips γγ and spaces from both sides, | |
| because punctuation is a rendering choice rather than a mishearing. "infer" is the median | |
| of the three clips. "first load" is the `Content-Length` of the two ONNX files the runtime | |
| fetches. | |
| | Model | dtype | device asked | tier used | first load | cold load | warm load | infer (p50) | CER | CER no-punct | | |
| |---|---|---|---|---|---|---|---|---|---| | |
| | `whisper-base` | q4 | wasm | wasm | 135.8 MB | 4064 ms | 1617 ms | 3636 ms | 0.062 | **0.023** | | |
| | `whisper-base` | q4 | webgpu | webgpu | 135.8 MB | 4415 ms | 2152 ms | **779 ms** | 0.062 | **0.023** | | |
| | `whisper-small` | q4 | wasm | wasm | 285.5 MB | 7355 ms | 3908 ms | 15485 ms | 0.083 | 0.023 | | |
| | `whisper-small` | q4 | webgpu | webgpu | 285.5 MB | 12846 ms | 4480 ms | 1934 ms | 0.083 | 0.023 | | |
| | `whisper-large-v3-turbo` | q4f16 | webgpu | webgpu | 537.4 MB | 29413 ms | 4055 ms | 2858 ms | 0.063 | **0.000** | | |
| | `whisper-large-v3-turbo` | q4 | wasm | wasm | 723.9 MB | 20692 ms | 4113 ms | **68200 ms** | 0.063 | 0.000 | | |
| | `whisper-base` | **q8** | wasm | **SESSION FAILED** | 73.3 MB | 3471 ms | β | β | β | β | | |
| | `whisper-base` | q8 | webgpu | webgpu | 73.3 MB | 3474 ms | 2021 ms | 4679 ms | 0.021 | 0.000 | | |
| | `whisper-small` | **q8** | wasm | **SESSION FAILED** | 237.5 MB | 5867 ms | β | β | β | β | | |
| Transcripts, per clip, in `tests/fixtures/asr_ab_results.json`. The single error | |
| `whisper-base` makes on this material is writing δ»ζ₯γ―**θ―γ**ε€©ζ° where VOICEVOX was given | |
| δ»ζ₯γ―**γγ**ε€©ζ° β the same word, the same reading, a different orthography. It is not a | |
| mishearing, which is why the punctuation-insensitive CER is 0.023 rather than 0.062. | |
| ### The finding that changed the default: `q8` does not work on the WASM tier | |
| `01-RESEARCH.md` recommends `onnx-community/whisper-base` with `dtype: 'q8'`. Measured, | |
| that configuration **cannot create an inference session on the WASM backend at all**: | |
| ``` | |
| Can't create a session. ERROR_CODE: 1, ERROR_MESSAGE: qdq_actions.cc:137 | |
| TransposeDQWeightsForMatMulNBits Missing required scale: | |
| model.decoder.embed_tokens.weight_merged_0_scale | |
| for node: model.decoder.embed_tokens.weight_transposed_DequantizeLinear | |
| ``` | |
| The same error appears for `whisper-base`, `whisper-small` and `whisper-large-v3-turbo`, | |
| and for the `int8` and `quantized` aliases. Every one of them loads fine on **WebGPU**, | |
| which is precisely how this would have shipped: green on a developer machine with a GPU, | |
| dead on every browser without one β the tier the fallback exists to serve. `q4` is the only | |
| quantisation measured working on **both** tiers, so `q4` is the default. | |
| ### The other finding: the documented fallback does not fall back | |
| `01-RESEARCH.md` prescribes "feature-detect `navigator.gpu`, request `webgpu` when present, | |
| catch initialisation failure, re-instantiate on WASM." Measured, that does not work. | |
| When `navigator.gpu` exists but `requestAdapter()` resolves to `null` β a headless browser, | |
| a blocklisted GPU, a browser started without GPU access β calling | |
| `pipeline(..., { device: 'webgpu' })` throws **and poisons the ONNX Runtime Web backend | |
| registry for the rest of the page**. The subsequent `pipeline(..., { device: 'wasm' })` for | |
| the same model then fails with the identical WebGPU error: | |
| ``` | |
| no available backend found. ERR: [webgpu] Error: Failed to get GPU adapter. | |
| ``` | |
| So `avatar/asr.js` **probes for an actual adapter first** (`probeWebGpu()`) and only issues | |
| the WebGPU call when one exists. The doomed call is never made, so the registry is never | |
| poisoned. The `try`/`catch` remains as a second line of defence for failures a probe cannot | |
| predict, such as an out-of-memory adapter. | |
| `tests/test_transport_seam.py::test_asr_webgpu_fallback_is_real_not_aspirational` fails if | |
| `requestAdapter` ever disappears from `avatar/asr.js`. | |
| --- | |
| ## Default chosen | |
| **`onnx-community/whisper-base` at `q4`**, on WebGPU when an adapter is available and on | |
| WASM when it is not. | |
| - It is the **only** measured configuration that runs on **both** tiers. `q8` is smaller | |
| (73.3 MB) and more accurate (CER 0.021), and it is disqualified outright because it | |
| cannot create a session on WASM. | |
| - 135.8 MB first load, then ~1.6β2.2 s to warm from the browser cache. | |
| - 779 ms median inference on WebGPU, 3636 ms on WASM. Both leave room inside VOIC-05's | |
| perceived-response budget; `whisper-small` at 15.5 s on WASM does not. | |
| - CER 0.023 ignoring punctuation, and its one error is an orthography choice for the same | |
| word rather than a misheard one. | |
| **`whisper-small` is rejected** on the numbers: 2.1Γ the bytes of `whisper-base`, an | |
| *identical* punctuation-insensitive CER of 0.023, and 4.3Γ the WASM inference time. It is | |
| dominated on every axis that matters. | |
| **"Better accuracy" toggle: `onnx-community/whisper-large-v3-turbo` at `q4f16`, WebGPU | |
| only.** CER 0.000 on this material and 2858 ms median inference β but a 537.4 MB first | |
| load, and 68 s per clip if it ever ran on WASM, which is why it must never be offered | |
| without an adapter. Both entries live in `MODELS` in `avatar/asr.js`; plan 01-08 owns the | |
| control that switches between them. | |
| --- | |
| ## Honest caveats | |
| - These clips are **VOICEVOX-synthesised speech**: cleaner, more regular and better | |
| articulated than a learner speaking into a laptop microphone in a room with a fan. Every | |
| CER above is a **best case** and overstates real-world accuracy. Real-speech confirmation | |
| is a manual verification in plan 01-10. | |
| - Phase 1's bar is "their transcript appears", not "the transcript is right". If the chosen | |
| default turns out to be embarrassing on real speech, that is a Phase 2/3 input, not a | |
| Phase 1 failure. | |
| - Three clips and two distinct sentences is a **small sample**. It is enough to rule | |
| `whisper-small` out and to disqualify `q8`, and it is not enough to claim a CER figure | |
| with confidence intervals. | |
| - Load and inference times are from one Windows laptop with an `intel xe-lpg` iGPU on one | |
| network. A visitor on a phone will see different numbers, particularly on WASM. | |
| - WebGPU numbers required **headed** Chromium. Headless Chromium exposes `navigator.gpu` | |
| and then returns a null adapter, so a headless re-run of the A/B records every candidate | |
| as WASM. That is true, and useless as a comparison. | |
| - Tier D (hosted ASR on accelerated hardware) is deliberately not built: it would consume | |
| visitor GPU quota and contradict SC-4. | |