VAM-Cross two-camera MimicVideo Video2World LoRA
This public repository contains the trainable-only Video2World LoRA
checkpoint from iteration 200 of v2w_panda_widowx_level2_widowx_texture_2cam_hstack_from_widowx250_video_fused_f0cea76_lora_r256.
The source run has termination status completed;
the selected four-component checkpoint set was verified before selecting the
model weight.
Required initial backbone
- Kind:
fused_video2world_dit - Repository:
dreamdifferent/widowx250-video-fused - Revision:
f0cea76b62c5dd66b06b9f965932ddea32a7b546 - Path:
checkpoints/video_backbone/iter_000001060_fused.pt - Bytes: 3913057284
- SHA-256:
d0f24c049bee63b03d3b62747b240a2d1822ddd5f83a52fcd866a882e80122b1 - Source iteration: 1060
This is an adapter checkpoint, not a standalone fused model. Load the exact
initial Video2World backbone above first and then apply this LoRA checkpoint.
For fused_video2world_dit, the initial checkpoint already includes the
earlier WidowX/Bridge LoRA fusion; loading the original Bridge backbone instead
would be incorrect.
Supporting runtime artifacts
- MimicVideo commit:
e3355dbc93132b576c02f920a59b4fc18a4f5906 - Checkpoint bundle:
jonpai/mimic-video@f28339034831e3c2374be075e622e1ff38ebe0f8 - Video tokenizer:
video_backbone/tokenizer/tokenizer.pth - T5 text encoder directory:
text_encoder/t5-11b
Use the same MimicVideo code and config recorded in config.yaml,
vam_cross_video2world_config.json, and
vam_cross_video_lora_manifest.json.
Training data contract
- Dataset revision:
dreamdifferent/vam-cross-level2-panda-widowx-widowx-texture@994be7f8952807008c316db7943ec732bf70b978 - Episodes/frames: 256 / 54349
- Cameras:
observation.images.corner_cam, observation.images.front_cam - View layout:
hstackat 5 Hz - Tasks: 24 episode-conditioned instructions (listed in
vam_cross_video_lora_manifest.json)
The dataset itself is not included. Users must comply with the dataset's current access policy and the upstream MimicVideo, NVIDIA Cosmos, and base-checkpoint terms.
- Downloads last month
- 9