dllm-vsr / README.md
jh-y's picture
Update model card (EMNLP 2026 camera-ready) and add length-predictor outputs (K_pred±5 pools for LRS3 test / WildVSR)
7a9cd42 verified
|
Raw History Blame Contribute Delete
2.72 kB
metadata
license: mit
library_name: pytorch
tags:
  - visual-speech-recognition
  - lip-reading
  - discrete-diffusion
  - diffusion-language-model
  - lrs3
  - arxiv:2605.28456
datasets:
  - lrs3

Diffusion Large Language Models for Visual Speech Recognition

Paper checkpoints for DLLM-VSR — adapting the Dream-7B discrete-diffusion LLM to Visual Speech Recognition (VSR) on LRS3.

Contents

Path Description Size
usr2/dream_stage2/ USR 2.0 + Dream-7B stage 2 (LoRA + adapter) 117 MB
usr2/len_pred/ Length predictor for USR 2.0 features 8.2 MB
avhubert/dream_stage2/ AV-HuBERT + Dream-7B stage 2 102 MB
avhubert/len_pred/ Length predictor for AV-HuBERT features 8.0 MB

Each dream_stage2/ holds trainable_model.safetensors (LoRA adapters + visual-feature projector). Each len_pred/ holds trainable_model.pt (small Transformer over visual features) together with its precomputed outputs len_pred_test.jsonl (LRS3 test) and len_pred_wildvsr.jsonl (WildVSR) — the K_pred ± 5 candidate pools used in the paper, so the evaluation scripts can skip length-predictor inference.

Note: Visual encoder weights (USR 2.0 Huge, AV-HuBERT Large) are not redistributed here. Download them from the original repos:

Results (WER, %)

All entries are trained on LRS3 (433h) only.

LRS3 test:

Decoding USR 2.0 AV-HuBERT
Implicit-length decoding 20.5 23.1
Length-guided candidate decoding (paper main) 19.4 21.9
Oracle-length decoding (upper bound) 17.7 20.2

WildVSR (out-of-domain, same LRS3-trained models):

Decoding USR 2.0 AV-HuBERT
Length-guided candidate decoding 39.9 46.1

Usage

huggingface-cli download jh-y/dllm-vsr --local-dir ckpt

Then follow the code repo's README for environment setup, preprocessing (auto-avsr pipeline), and inference scripts.

Citation

@article{yeo2026dllmvsr,
  title={Diffusion Large Language Models for Visual Speech Recognition},
  author={Yeo, Jeong Hun and Kim, Chae Won and Rha, Hyeongseop and Ro, Yong Man},
  journal={arXiv preprint arXiv:2605.28456},
  year={2026}
}