rag-gesture-weights / README.md
m-hamza-mughal's picture
Upload README.md with huggingface_hub
ad15b7f verified
|
Raw
History Blame Contribute Delete
3.67 kB
---
license: cc-by-nc-4.0
language:
- en
tags:
- gesture-generation
- co-speech-gesture
- motion-synthesis
- diffusion
- retrieval-augmented-generation
- smplx
- beat2
library_name: pytorch
---
# RAG-Gesture: Retrieval-Augmented Co-Speech Gesture Generation
Pretrained weights for **RAG-Gesture**, a retrieval-augmented latent
diffusion framework for generating semantically-aware 3D co-speech gestures
on the BEAT2 (BEATX) dataset.
**Paper:** *Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis* (CVPR 2025).
**Code:** https://github.com/m-hamza-mughal/RAG-Gesture
## What's in this repo
| Path | Description |
| --------------------- | -------------------------------------------------------------------------------------------- |
| `vae/` | Four part-specific Transformer VAEs (upper / hands / face / lower+translation) used as the motion encoder/decoder. Each chunks 15 frames at 15 fps into one 512-d latent. |
| `diffusion/` | Trained RAG-Gesture diffusion transformer checkpoint (`base_beatx_len150fps15_finalweights/`). |
| `assets_deps/` | SMPL-X model files required by the dataloader and decoder for converting axis-angle poses to joints. |
## How the model works
RAG-Gesture denoises a body-part-factored latent sequence conditioned on
text (BERT), audio (wav2vec2), and a speaker embedding. At inference,
retrieval strategies — **discourse-relation** & **gesture-type** — pull semantically-matched motion clips from a database
of training samples. The retrieved latents are injected into the diffusion
process through *DDIM inversion + per-step insertion guidance*, biasing the
denoiser toward the retrieved semantics without sacrificing the base model's
fluency.
Key design choices:
- **4-VAE body-part tokenization** (upper / hands / face / lower+translation)
with a `[sep]` token between parts, all concatenated along the time axis.
- **Multi-conditional CFG** mixing text-only and unconditional predictions
with a coarse-to-fine scaling schedule.
- **Insertion guidance** that runs a small inner gradient loop on the
current `x_t` against the per-timestep inverted retrieval latent.
## How to use
```bash
# 1. Clone the code
git clone https://github.com/m-hamza-mughal/RAG-Gesture
cd RAG-Gesture
pip install -r requirements.txt
# 2. Pull these weights + SMPL-X deps
python tools/download_weights.py
# 3. Pull the BEAT2 dataset + RAG-Gesture annotations
python tools/download_annotations.py
# 4. Run RAG-guided inference (discourse retrieval)
PYTHONPATH=".":$PYTHONPATH python tools/visualize.py \
experiments/diffusion/base_beatx_len150fps15_finalweights/basegesture_len150_beat.py \
experiments/diffusion/base_beatx_len150fps15_finalweights/epoch_64.pth \
--retrieval_method discourse_guidance_test \
--use_retrieval --use_inversion --use_insertion_guidance \
--guidance_iters decreasing_till_25
```
See the [GitHub README](https://github.com/m-hamza-mughal/RAG-Gesture) for
training, evaluation, long-form synthesis, and ablation commands.
## Citation
```
@InProceedings{mughal2024raggesture,
title = {Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis},
author = {M. Hamza Mughal and Rishabh Dabral and Merel C. J. Scholman and Vera Demberg and Christian Theobalt},
booktitle = {Computer Vision and Pattern Recognition (CVPR)},
year = {2025}
}
```
## Acknowledgements
Built on [EMAGE / PantoMatrix](https://pantomatrix.github.io/EMAGE/) and
[ReMoDiffuse](https://github.com/mingyuan-zhang/ReMoDiffuse).