Auto-AVSR lip reader (LRS3, visual-only) β€” int8 ONNX, 203 MB

The encoder + CTC head of Auto-AVSR's LRS3_V_WER19.1 visual speech recognition model, exported to ONNX and dynamically quantized to int8. It reads lips from 88Γ—88 grayscale mouth crops and runs on CPU in onnxruntime or in the browser with onnxruntime-web (WASM). Built for a StormHacks 2026 assistive silent-speech app (mouth words at a webcam β†’ text β†’ voice).

Files

File Size sha256
lipread_ctc.int8.onnx 203.4 MB da02d72ee472eb3d7d50f7dd309bfe81973acb4cac8bb770dd86aa97872f195b
tokens.json 78 KB 97f0f20cfebadaf51f7cdb9e7a26835fa0fbd6a956701434e899a4f9934d979c
quantization.json β€” recipe, op counts, sanity checks

Inputs, outputs, decoding

  • Input video: float32 [1, 1, T, 88, 88] (N, C, T, H, W), 25 fps, T ≀ 500 (20 s; the position table was trimmed to 500 frames). Pixels /255, then (x βˆ’ 0.421) / 0.165.
  • Output log_probs: float32 [T, 5049], CTC log-probabilities over the SentencePiece unigram-5000 vocabulary in tokens.json.
  • Greedy CTC: argmax per frame β†’ collapse repeats β†’ drop blank (index 0) and <eos> (last index) β†’ join pieces β†’ ▁ marks a word start. The model emits UPPERCASE text.

Preprocessing (must match training)

25 fps β†’ MediaPipe BlazeFace (short-range) keypoints: right eye, left eye, nose tip, mouth centre β†’ fill missed frames by linear interpolation β†’ Β±6-frame smoothing β†’ similarity transform onto the mean face (256Γ—256 reference) β†’ 96Γ—96 mouth patch β†’ grayscale β†’ centre-crop 88Γ—88 β†’ normalise as above. Reference implementation: auto_avsr's preparation/ (MediaPipe detector + video_process.py).

Accuracy and speed

Greedy CTC, no language model.

Eval int8 (this) fp32 export
LRS3 test, first 100 clips (mattymchen/lrs3-test, pre-made crops) 28.5% WER 28.6% WER
20 raw face videos through the crop pipeline above 25.4% WER 26.2% WER
  • Mean frame-level argmax agreement with fp32: 0.993. Over 300 LRS3 clips the WER difference is βˆ’0.31 points (95% CI βˆ’0.85 … +0.22).
  • The original checkpoint's published 19.1% WER uses beam search + an RNN language model on the full LRS3 test set; this file is the encoder + CTC head only.
  • onnxruntime CPU: 1.2–1.3Γ— faster than the fp32 export. onnxruntime-web 1.30 (WASM) in Chromium on a laptop: ~1.05 s for a 2.8 s utterance, ~3 s cold start.
  • Accuracy drops sharply when the camera delivers fewer than 25 fps (offline: 38% β†’ 48% β†’ 60% WER at 25 β†’ 15 β†’ 10 fps on 30 clips), so capture at a steady 25–30 fps.

Usage (Python)

import json
import numpy as np
import onnxruntime as ort

sess = ort.InferenceSession("lipread_ctc.int8.onnx")
tokens = json.load(open("tokens.json"))  # 5049 pieces

# crops: uint8 [T, 88, 88] mouth crops at 25 fps (see Preprocessing), T <= 500
x = (crops.astype(np.float32) / 255.0 - 0.421) / 0.165
log_probs = sess.run(["log_probs"], {"video": x[None, None]})[0]  # [T, 5049]

ids, prev = [], None
for i in log_probs.argmax(-1):
    if i != prev and i not in (0, len(tokens) - 1):  # blank = 0, <eos> = last
        ids.append(int(i))
    prev = i
text = "".join(tokens[i] for i in ids).replace("▁", " ").strip()

In the browser, onnxruntime-web with executionProviders: ["wasm"] is the tested path. WebGPU also loads it on some adapters but was slower in our tests.

Quantization recipe (dyn-pw8-rn16)

Dynamic int8, per-channel MatMulInteger for all 134 MatMuls, including the 24 pointwise Conv1d layers rewritten as MatMul. The ResNet and 3D front-end convolutions are stored as fp16 and cast at load. Position table trimmed from 9,999 to 999 rows. Opset 17, standard ONNX ops only (no contrib ops). Built from the fp32 export with quantize_onnx.py --variant dyn-pw8-rn16, and checked by a regression suite: WER gates against fp32 plus an exact-output lock.

Licence and attribution

  • Derived from Auto-AVSR LRS3_V_WER19.1: Ma et al., Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels, ICASSP 2023 (https://github.com/mpc001/auto_avsr). Weights obtained from the Amanvir/LRS3_V_WER19.1 mirror used by Chaplin (https://github.com/amanvirparhar/chaplin).
  • Auto-AVSR's code is Apache-2.0. Its pretrained models carry the terms of their training data (LRS3: BBC / TED), which allow research and non-commercial use only. This quantized copy inherits those terms: do not use it commercially.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for eschmechel/auto-avsr-lrs3-vsr-int8-onnx

Quantized
(1)
this model