File size: 2,715 Bytes
b6574cb
 
 
 
 
 
 
7a9cd42
b6574cb
7a9cd42
b6574cb
 
 
 
398dab7
b6574cb
 
 
398dab7
7a9cd42
b6574cb
 
 
 
 
 
 
 
 
 
 
7a9cd42
b6574cb
 
 
7a9cd42
b6574cb
7a9cd42
b6574cb
 
 
7a9cd42
 
 
 
 
 
 
 
 
 
b6574cb
 
7a9cd42
b6574cb
 
 
 
 
 
 
7a9cd42
b6574cb
 
 
 
 
 
 
398dab7
b6574cb
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
---
license: mit
library_name: pytorch
tags:
- visual-speech-recognition
- lip-reading
- discrete-diffusion
- diffusion-language-model
- lrs3
- arxiv:2605.28456
datasets:
- lrs3
---

# Diffusion Large Language Models for Visual Speech Recognition

Paper checkpoints for **DLLM-VSR** — adapting the Dream-7B discrete-diffusion LLM to Visual Speech Recognition (VSR) on LRS3.

- Paper: [arxiv.org/abs/2605.28456](http://arxiv.org/abs/2605.28456)
- Code: [github.com/JeongHun0716/dllm-vsr](https://github.com/JeongHun0716/dllm-vsr)
- Authors: Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro

## Contents

| Path | Description | Size |
|---|---|---|
| `usr2/dream_stage2/`     | USR 2.0 + Dream-7B stage 2 (LoRA + adapter) | 117 MB |
| `usr2/len_pred/`         | Length predictor for USR 2.0 features        | 8.2 MB |
| `avhubert/dream_stage2/` | AV-HuBERT + Dream-7B stage 2                 | 102 MB |
| `avhubert/len_pred/`     | Length predictor for AV-HuBERT features      | 8.0 MB |

Each `dream_stage2/` holds `trainable_model.safetensors` (LoRA adapters + visual-feature projector). Each `len_pred/` holds `trainable_model.pt` (small Transformer over visual features) together with its precomputed outputs `len_pred_test.jsonl` (LRS3 test) and `len_pred_wildvsr.jsonl` (WildVSR) — the K_pred ± 5 candidate pools used in the paper, so the evaluation scripts can skip length-predictor inference.

**Note**: Visual encoder weights (USR 2.0 Huge, AV-HuBERT Large) are **not** redistributed here. Download them from the original repos:
- AV-HuBERT: https://github.com/facebookresearch/av_hubert
- USR 2.0:   https://github.com/ahaliassos/usr2 (Huge, LRS2+LRS3+Vox2+AVS pretrain checkpoint)

## Results (WER, %)

All entries are trained on **LRS3 (433h)** only.

LRS3 test:

| Decoding | USR 2.0 | AV-HuBERT |
|---|:---:|:---:|
| Implicit-length decoding                      | 20.5 | 23.1 |
| Length-guided candidate decoding (paper main) | **19.4** | **21.9** |
| Oracle-length decoding (upper bound)          | 17.7 | 20.2 |

WildVSR (out-of-domain, same LRS3-trained models):

| Decoding | USR 2.0 | AV-HuBERT |
|---|:---:|:---:|
| Length-guided candidate decoding | **39.9** | **46.1** |

## Usage

```bash
huggingface-cli download jh-y/dllm-vsr --local-dir ckpt
```

Then follow the [code repo's README](https://github.com/JeongHun0716/dllm-vsr) for environment setup, preprocessing (auto-avsr pipeline), and inference scripts.

## Citation

```bibtex
@article{yeo2026dllmvsr,
  title={Diffusion Large Language Models for Visual Speech Recognition},
  author={Yeo, Jeong Hun and Kim, Chae Won and Rha, Hyeongseop and Ro, Yong Man},
  journal={arXiv preprint arXiv:2605.28456},
  year={2026}
}
```