jh-y commited on
Commit
b6574cb
·
verified ·
1 Parent(s): 0f3f563

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +63 -0
README.md ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: pytorch
4
+ tags:
5
+ - visual-speech-recognition
6
+ - lip-reading
7
+ - discrete-diffusion
8
+ - lrs3
9
+ datasets:
10
+ - lrs3
11
+ ---
12
+
13
+ # DLLM-VSR: Diffusion Large Language Models for Visual Speech Recognition
14
+
15
+ Paper checkpoints for **DLLM-VSR** — adapting the Dream-7B discrete-diffusion LLM to Visual Speech Recognition (VSR) on LRS3.
16
+
17
+ - Paper: [arxiv.org/abs/XXXX.XXXXX](http://arxiv.org/abs/XXXX.XXXXX)
18
+ - Code: [github.com/jh-y/dllm-vsr](https://github.com/jh-y/dllm-vsr)
19
+ - Authors: Jeong Hun Yeo, Chae Won Kim, Hyeongseop Rha, Yong Man Ro
20
+
21
+ ## Contents
22
+
23
+ | Path | Description | Size |
24
+ |---|---|---|
25
+ | `usr2/dream_stage2/` | USR 2.0 + Dream-7B stage 2 (LoRA + adapter) | 117 MB |
26
+ | `usr2/len_pred/` | Length predictor for USR 2.0 features | 8.2 MB |
27
+ | `avhubert/dream_stage2/` | AV-HuBERT + Dream-7B stage 2 | 102 MB |
28
+ | `avhubert/len_pred/` | Length predictor for AV-HuBERT features | 8.0 MB |
29
+
30
+ Each `dream_stage2/` holds `trainable_model.safetensors` (LoRA adapters + visual-feature projector). Each `len_pred/` holds `trainable_model.pt` (small Transformer over visual features).
31
+
32
+ **Note**: Visual encoder weights (USR 2.0 Huge, AV-HuBERT Large) are **not** redistributed here. Download them from the original repos:
33
+ - AV-HuBERT: https://github.com/facebookresearch/av_hubert
34
+ - USR 2.0: https://github.com/ahaliassos/usr2
35
+
36
+ ## Results on LRS3 test (WER, %)
37
+
38
+ All entries are trained on **LRS3 (433h)** only.
39
+
40
+ | Decoding | USR 2.0 | AV-HuBERT |
41
+ |---|:---:|:---:|
42
+ | Direct | 20.5 | 23.1 |
43
+ | Length-guided candidate decoding (paper main) | **19.5** | **21.9** |
44
+ | Oracle-length (upper-bound reference) | 17.7 | 20.2 |
45
+
46
+ ## Usage
47
+
48
+ ```bash
49
+ huggingface-cli download jh-y/dllm-vsr --local-dir ckpt
50
+ ```
51
+
52
+ Then follow the [code repo's README](https://github.com/jh-y/dllm-vsr) for environment setup, preprocessing (auto-avsr pipeline), and inference scripts.
53
+
54
+ ## Citation
55
+
56
+ ```bibtex
57
+ @article{yeo2026dllmvsr,
58
+ title={Diffusion Large Language Models for Visual Speech Recognition},
59
+ author={Yeo, Jeong Hun and Kim, Chae Won and Rha, Hyeongseop and Ro, Yong Man},
60
+ journal={arXiv preprint arXiv:XXXX.XXXXX},
61
+ year={2026}
62
+ }
63
+ ```