Text-to-Video
ViRDM / README.md
cr8br0ze's picture nielsr's picture
nielsr HF Staff
Add pipeline tag and project/code links (#1)
5af481c
|
Raw History Blame Contribute Delete
1.19 kB
metadata
license: other
pipeline_tag: text-to-video

ViRDM models and training assets

This repository contains the trained causal four-step ViRDM generators and the compact frozen artifacts required to reproduce training.

  • checkpoints/virdm_causal4_dynamic_step20.pt: the primary 20-update model with dynamics regularization (5e-4).
  • checkpoints/virdm_causal4_nodynamic_step20.pt: the matched model without dynamics regularization.
  • prompt_data/data.mdb: 6,505 captions in the original ordered zhuhz22/Causal-Forcing-data training split; no video or latent payloads.
  • references/reference_M4096.pt: the fixed 4,096-landmark joint video-text reference used by ViRDM. Its publication metadata contains stable upstream identifiers rather than machine-local paths.
  • references/siglip2_text_fp32.npy: frozen SigLIP2 features aligned row for row with the reference and prompt table.
  • asset_receipt.json: row contract, filenames, and SHA-256 checksums.

📄 Paper: ViRDM

Project page: https://neu-vi.github.io/ViRDM/

Code: https://github.com/neu-vi/ViRDM