Auto-AVSR lip reader (LRS3, visual-only) β int8 ONNX, 203 MB
The encoder + CTC head of Auto-AVSR's LRS3_V_WER19.1 visual speech recognition model, exported to
ONNX and dynamically quantized to int8. It reads lips from 88Γ88 grayscale mouth crops and runs on
CPU in onnxruntime or in the browser with onnxruntime-web (WASM). Built for a StormHacks 2026
assistive silent-speech app (mouth words at a webcam β text β voice).
Files
| File | Size | sha256 |
|---|---|---|
lipread_ctc.int8.onnx |
203.4 MB | da02d72ee472eb3d7d50f7dd309bfe81973acb4cac8bb770dd86aa97872f195b |
tokens.json |
78 KB | 97f0f20cfebadaf51f7cdb9e7a26835fa0fbd6a956701434e899a4f9934d979c |
quantization.json |
β | recipe, op counts, sanity checks |
Inputs, outputs, decoding
- Input
video: float32[1, 1, T, 88, 88](N, C, T, H, W), 25 fps, T β€ 500 (20 s; the position table was trimmed to 500 frames). Pixels/255, then(x β 0.421) / 0.165. - Output
log_probs: float32[T, 5049], CTC log-probabilities over the SentencePiece unigram-5000 vocabulary intokens.json. - Greedy CTC: argmax per frame β collapse repeats β drop blank (index 0) and
<eos>(last index) β join pieces ββmarks a word start. The model emits UPPERCASE text.
Preprocessing (must match training)
25 fps β MediaPipe BlazeFace (short-range) keypoints: right eye, left eye, nose tip, mouth centre β
fill missed frames by linear interpolation β Β±6-frame smoothing β similarity transform onto the mean
face (256Γ256 reference) β 96Γ96 mouth patch β grayscale β centre-crop 88Γ88 β normalise as above.
Reference implementation: auto_avsr's preparation/ (MediaPipe detector + video_process.py).
Accuracy and speed
Greedy CTC, no language model.
| Eval | int8 (this) | fp32 export |
|---|---|---|
LRS3 test, first 100 clips (mattymchen/lrs3-test, pre-made crops) |
28.5% WER | 28.6% WER |
| 20 raw face videos through the crop pipeline above | 25.4% WER | 26.2% WER |
- Mean frame-level argmax agreement with fp32: 0.993. Over 300 LRS3 clips the WER difference is β0.31 points (95% CI β0.85 β¦ +0.22).
- The original checkpoint's published 19.1% WER uses beam search + an RNN language model on the full LRS3 test set; this file is the encoder + CTC head only.
- onnxruntime CPU: 1.2β1.3Γ faster than the fp32 export. onnxruntime-web 1.30 (WASM) in Chromium on a laptop: ~1.05 s for a 2.8 s utterance, ~3 s cold start.
- Accuracy drops sharply when the camera delivers fewer than 25 fps (offline: 38% β 48% β 60% WER at 25 β 15 β 10 fps on 30 clips), so capture at a steady 25β30 fps.
Usage (Python)
import json
import numpy as np
import onnxruntime as ort
sess = ort.InferenceSession("lipread_ctc.int8.onnx")
tokens = json.load(open("tokens.json")) # 5049 pieces
# crops: uint8 [T, 88, 88] mouth crops at 25 fps (see Preprocessing), T <= 500
x = (crops.astype(np.float32) / 255.0 - 0.421) / 0.165
log_probs = sess.run(["log_probs"], {"video": x[None, None]})[0] # [T, 5049]
ids, prev = [], None
for i in log_probs.argmax(-1):
if i != prev and i not in (0, len(tokens) - 1): # blank = 0, <eos> = last
ids.append(int(i))
prev = i
text = "".join(tokens[i] for i in ids).replace("β", " ").strip()
In the browser, onnxruntime-web with executionProviders: ["wasm"] is the tested path. WebGPU also
loads it on some adapters but was slower in our tests.
Quantization recipe (dyn-pw8-rn16)
Dynamic int8, per-channel MatMulInteger for all 134 MatMuls, including the 24 pointwise Conv1d
layers rewritten as MatMul. The ResNet and 3D front-end convolutions are stored as fp16 and cast at
load. Position table trimmed from 9,999 to 999 rows. Opset 17, standard ONNX ops only (no contrib
ops). Built from the fp32 export with quantize_onnx.py --variant dyn-pw8-rn16, and checked by a
regression suite: WER gates against fp32 plus an exact-output lock.
Licence and attribution
- Derived from Auto-AVSR
LRS3_V_WER19.1: Ma et al., Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels, ICASSP 2023 (https://github.com/mpc001/auto_avsr). Weights obtained from theAmanvir/LRS3_V_WER19.1mirror used by Chaplin (https://github.com/amanvirparhar/chaplin). - Auto-AVSR's code is Apache-2.0. Its pretrained models carry the terms of their training data (LRS3: BBC / TED), which allow research and non-commercial use only. This quantized copy inherits those terms: do not use it commercially.
Model tree for eschmechel/auto-avsr-lrs3-vsr-int8-onnx
Base model
Amanvir/LRS3_V_WER19.1