Text-to-Video
ViRDM / README.md
cr8br0ze's picture nielsr's picture
nielsr HF Staff
Add pipeline tag and project/code links (#1)
5af481c
|
Raw History Blame Contribute Delete
1.19 kB
---
license: other
pipeline_tag: text-to-video
---
# ViRDM models and training assets
This repository contains the trained causal four-step ViRDM generators and the
compact frozen artifacts required to reproduce training.
- `checkpoints/virdm_causal4_dynamic_step20.pt`: the primary
20-update model with dynamics regularization (`5e-4`).
- `checkpoints/virdm_causal4_nodynamic_step20.pt`: the matched model
without dynamics regularization.
- `prompt_data/data.mdb`: 6,505 captions in the original ordered
`zhuhz22/Causal-Forcing-data` training split; no video or latent payloads.
- `references/reference_M4096.pt`: the fixed 4,096-landmark joint video-text
reference used by ViRDM. Its publication metadata contains stable upstream
identifiers rather than machine-local paths.
- `references/siglip2_text_fp32.npy`: frozen SigLIP2 features aligned row for
row with the reference and prompt table.
- `asset_receipt.json`: row contract, filenames, and SHA-256 checksums.
📄 [Paper: ViRDM](https://arxiv.org/abs/2609.28923)
Project page: [https://neu-vi.github.io/ViRDM/](https://neu-vi.github.io/ViRDM/)
Code: [https://github.com/neu-vi/ViRDM](https://github.com/neu-vi/ViRDM)