WolfDavid's picture
feat(01-07): add tiered browser ASR and wire push-to-talk into the shared turn loop
b4ca58d
|
Raw History Blame
11.3 kB
# ASR tiers and model choice (VOIC-02)
Every number on this page was measured on this project's own audio, by
`tests/fixtures/make_asr_ab_results.py` driving `avatar/asr-harness.html` in a real
browser. The raw data is `tests/fixtures/asr_ab_results.json`. Nothing here is quoted
from a model card.
| | |
|---|---|
| Measured | 2026-08-27 |
| Browser | Chromium 151.0.0.0 (Playwright), Windows NT 10.0 Win64 |
| GPU adapter | `intel xe-lpg` (WebGPU available, **headed** Chromium only) |
| Runtime | `https://esm.sh/@huggingface/transformers@4.2.0` |
| Clips | `speech_ja.wav`, `speech_ja_long.wav`, `speech_ja_slow.wav` β€” VOICEVOX γšγ‚“γ γ‚‚γ‚“, plan 01-04 |
---
## Tier design
| Tier | Path | GPU quota | Availability |
|---|---|---|---|
| **A** | `@huggingface/transformers` 4.2.0, `device: 'webgpu'` | **zero** | Chrome 113+ desktop, Chrome 121+ Android 12+, Safari macOS 26 / iOS 26, Firefox 141+ |
| **B** | same library, `device: 'wasm'` | **zero** | everywhere WASM+SIMD works; several times slower |
| **C** | type instead of speak (VOIC-04) | zero | universal β€” the true floor |
| **D** | hosted whisper-large-v3-turbo on accelerated hardware | **burns visitor quota** | **NOT built, and must not be.** It would consume the exact resource SC-4 tests the loop's survival without |
`tests/test_transport_seam.py::test_no_remote_asr_path_exists` fails the build if tier D
ever appears in `avatar/`.
---
## Gate thresholds (measured against the committed fixtures)
Read straight off the WAV files, independent of any microphone: 20 ms frames, `peak frame
RMS / median frame RMS`. Reproduce with `window.__measureClip(url)` in the harness.
| Fixture | duration | RMS | peak/median frame RMS | gate verdict | condition that fired |
|---|---|---|---|---|---|
| `silence_30s.wav` | 30.000 s | 0.00000 | ∞ (median is 0) | **REJECT** | `rms-floor` |
| `cafe_noise_30s.wav` | 30.000 s | 0.05760 | **1.961** | **REJECT** | `envelope-modulation` |
| `speech_ja.wav` | 1.056 s | 0.07153 | **10.711** | **ACCEPT** | β€” |
| `speech_ja_long.wav` | 5.504 s | 0.05515 | 7.509 | **ACCEPT** | β€” |
| `speech_ja_slow.wav` | 7.381 s | 0.04918 | 6.368 | **ACCEPT** | β€” |
Thresholds: **300 ms** duration floor, **0.01** RMS floor, **2.5** modulation floor.
The 2.5 figure is the whole reason the third condition exists. The cafΓ© fixture is written
at βˆ’24.8 dBFS β€” roughly **six times** the RMS floor β€” so an RMS-only gate passes it
comfortably. It is rejected because steady broadband noise is not modulated: 1.961 against
speech's 6.4–10.7. The threshold sits with ~28% headroom below the noise and a 2.5Γ—
margin below the least-modulated speech clip.
### The same gate, measured through a real microphone
Chromium fed the fixture as its capture device
(`--use-file-for-fake-audio-capture`), which is the path a learner actually exercises.
| Fixture | hold | browser audio processing | duration | RMS | modulation | verdict |
|---|---|---|---|---|---|---|
| `silence_30s.wav` | 3000 ms | on | 2731 ms | 0.00000 | ∞ | REJECT `rms-floor` |
| `cafe_noise_30s.wav` | 3000 ms | on | 2731 ms | **0.00547** | 2.344 | REJECT `rms-floor` |
| `speech_ja.wav` | 1100 ms | on | 768 ms | 0.15451 | 6.694 | ACCEPT |
| `silence_30s.wav` | 3000 ms | **off** | 2731 ms | 0.00000 | ∞ | REJECT `rms-floor` |
| `cafe_noise_30s.wav` | 3000 ms | **off** | 2304 ms | 0.04409 | **1.360** | REJECT `envelope-modulation` |
| `speech_ja.wav` | 1400 ms | **off** | 1195 ms | 0.06717 | 15.012 | ACCEPT |
| `speech_ja.wav` | 150 ms | **off** | 0 ms | β€” | β€” | REJECT `duration-floor` |
**Chromium's WebRTC noise suppression is doing part of the gate's job for us** β€” it drops
the cafΓ© fixture from 0.0577 to 0.0055 RMS, below the floor, so the modulation condition
is never consulted. That is a good outcome in production and a useless one in a test,
because it would leave the third condition unexercised while the suite looked green.
`tests/e2e/test_asr_standalone.py` therefore drives the harness with `?processing=off` and
verifies the gate in its **pessimistic** configuration β€” the raw microphone every browser
without WebRTC processing hands us anyway. Production keeps all three processing flags on.
### The hallucination blocklist
Applied to the transcript, not the audio, and only under **1.5 s**. Seven known Japanese
subtitle-boilerplate strings, listed in `avatar/mic.js`.
Deliberately **not** blocklisted: the bare polite form γ€Œγ‚γ‚ŠγŒγ¨γ†γ”γ–γ„γΎγ—γŸγ€. It is an
ordinary thing a learner says out loud, and swallowing a real utterance is a worse failure
than echoing one hallucination.
`tests/test_transport_seam.py::test_blocklist_does_not_swallow_ordinary_japanese` exists
specifically so a later "helpful" edit cannot add it back.
---
## Model A/B
CER is against the exact strings plan 01-04 handed to VOICEVOX, so the ground truth is
known rather than transcribed. "CER no-punct" strips 、。 and spaces from both sides,
because punctuation is a rendering choice rather than a mishearing. "infer" is the median
of the three clips. "first load" is the `Content-Length` of the two ONNX files the runtime
fetches.
| Model | dtype | device asked | tier used | first load | cold load | warm load | infer (p50) | CER | CER no-punct |
|---|---|---|---|---|---|---|---|---|---|
| `whisper-base` | q4 | wasm | wasm | 135.8 MB | 4064 ms | 1617 ms | 3636 ms | 0.062 | **0.023** |
| `whisper-base` | q4 | webgpu | webgpu | 135.8 MB | 4415 ms | 2152 ms | **779 ms** | 0.062 | **0.023** |
| `whisper-small` | q4 | wasm | wasm | 285.5 MB | 7355 ms | 3908 ms | 15485 ms | 0.083 | 0.023 |
| `whisper-small` | q4 | webgpu | webgpu | 285.5 MB | 12846 ms | 4480 ms | 1934 ms | 0.083 | 0.023 |
| `whisper-large-v3-turbo` | q4f16 | webgpu | webgpu | 537.4 MB | 29413 ms | 4055 ms | 2858 ms | 0.063 | **0.000** |
| `whisper-large-v3-turbo` | q4 | wasm | wasm | 723.9 MB | 20692 ms | 4113 ms | **68200 ms** | 0.063 | 0.000 |
| `whisper-base` | **q8** | wasm | **SESSION FAILED** | 73.3 MB | 3471 ms | β€” | β€” | β€” | β€” |
| `whisper-base` | q8 | webgpu | webgpu | 73.3 MB | 3474 ms | 2021 ms | 4679 ms | 0.021 | 0.000 |
| `whisper-small` | **q8** | wasm | **SESSION FAILED** | 237.5 MB | 5867 ms | β€” | β€” | β€” | β€” |
Transcripts, per clip, in `tests/fixtures/asr_ab_results.json`. The single error
`whisper-base` makes on this material is writing 今ζ—₯は**良い**倩気 where VOICEVOX was given
今ζ—₯は**いい**倩気 β€” the same word, the same reading, a different orthography. It is not a
mishearing, which is why the punctuation-insensitive CER is 0.023 rather than 0.062.
### The finding that changed the default: `q8` does not work on the WASM tier
`01-RESEARCH.md` recommends `onnx-community/whisper-base` with `dtype: 'q8'`. Measured,
that configuration **cannot create an inference session on the WASM backend at all**:
```
Can't create a session. ERROR_CODE: 1, ERROR_MESSAGE: qdq_actions.cc:137
TransposeDQWeightsForMatMulNBits Missing required scale:
model.decoder.embed_tokens.weight_merged_0_scale
for node: model.decoder.embed_tokens.weight_transposed_DequantizeLinear
```
The same error appears for `whisper-base`, `whisper-small` and `whisper-large-v3-turbo`,
and for the `int8` and `quantized` aliases. Every one of them loads fine on **WebGPU**,
which is precisely how this would have shipped: green on a developer machine with a GPU,
dead on every browser without one β€” the tier the fallback exists to serve. `q4` is the only
quantisation measured working on **both** tiers, so `q4` is the default.
### The other finding: the documented fallback does not fall back
`01-RESEARCH.md` prescribes "feature-detect `navigator.gpu`, request `webgpu` when present,
catch initialisation failure, re-instantiate on WASM." Measured, that does not work.
When `navigator.gpu` exists but `requestAdapter()` resolves to `null` β€” a headless browser,
a blocklisted GPU, a browser started without GPU access β€” calling
`pipeline(..., { device: 'webgpu' })` throws **and poisons the ONNX Runtime Web backend
registry for the rest of the page**. The subsequent `pipeline(..., { device: 'wasm' })` for
the same model then fails with the identical WebGPU error:
```
no available backend found. ERR: [webgpu] Error: Failed to get GPU adapter.
```
So `avatar/asr.js` **probes for an actual adapter first** (`probeWebGpu()`) and only issues
the WebGPU call when one exists. The doomed call is never made, so the registry is never
poisoned. The `try`/`catch` remains as a second line of defence for failures a probe cannot
predict, such as an out-of-memory adapter.
`tests/test_transport_seam.py::test_asr_webgpu_fallback_is_real_not_aspirational` fails if
`requestAdapter` ever disappears from `avatar/asr.js`.
---
## Default chosen
**`onnx-community/whisper-base` at `q4`**, on WebGPU when an adapter is available and on
WASM when it is not.
- It is the **only** measured configuration that runs on **both** tiers. `q8` is smaller
(73.3 MB) and more accurate (CER 0.021), and it is disqualified outright because it
cannot create a session on WASM.
- 135.8 MB first load, then ~1.6–2.2 s to warm from the browser cache.
- 779 ms median inference on WebGPU, 3636 ms on WASM. Both leave room inside VOIC-05's
perceived-response budget; `whisper-small` at 15.5 s on WASM does not.
- CER 0.023 ignoring punctuation, and its one error is an orthography choice for the same
word rather than a misheard one.
**`whisper-small` is rejected** on the numbers: 2.1Γ— the bytes of `whisper-base`, an
*identical* punctuation-insensitive CER of 0.023, and 4.3Γ— the WASM inference time. It is
dominated on every axis that matters.
**"Better accuracy" toggle: `onnx-community/whisper-large-v3-turbo` at `q4f16`, WebGPU
only.** CER 0.000 on this material and 2858 ms median inference β€” but a 537.4 MB first
load, and 68 s per clip if it ever ran on WASM, which is why it must never be offered
without an adapter. Both entries live in `MODELS` in `avatar/asr.js`; plan 01-08 owns the
control that switches between them.
---
## Honest caveats
- These clips are **VOICEVOX-synthesised speech**: cleaner, more regular and better
articulated than a learner speaking into a laptop microphone in a room with a fan. Every
CER above is a **best case** and overstates real-world accuracy. Real-speech confirmation
is a manual verification in plan 01-10.
- Phase 1's bar is "their transcript appears", not "the transcript is right". If the chosen
default turns out to be embarrassing on real speech, that is a Phase 2/3 input, not a
Phase 1 failure.
- Three clips and two distinct sentences is a **small sample**. It is enough to rule
`whisper-small` out and to disqualify `q8`, and it is not enough to claim a CER figure
with confidence intervals.
- Load and inference times are from one Windows laptop with an `intel xe-lpg` iGPU on one
network. A visitor on a phone will see different numbers, particularly on WASM.
- WebGPU numbers required **headed** Chromium. Headless Chromium exposes `navigator.gpu`
and then returns a null adapter, so a headless re-run of the A/B records every candidate
as WASM. That is true, and useless as a comparison.
- Tier D (hosted ASR on accelerated hardware) is deliberately not built: it would consume
visitor GPU quota and contradict SC-4.