--- tags: - mimic-video - robotics - action-prediction --- # VAM-Cross MimicVideo World2Action decoder This public repository contains the World2Action decoder checkpoint from iteration 1800 of `w2a_so101_level4_widowx_texture_2cam_hstack_action_iter2374_videolora_iter400_widowx_teleop_recording_frame_v1`. The run stopped because of `completed`. The newest complete model, optimizer, scheduler, and trainer checkpoint set was verified before selecting the uploaded model weight. ## Required frozen inputs - MimicVideo commit: `e3355dbc93132b576c02f920a59b4fc18a4f5906` - Initial Video2World backbone: `dreamdifferent/widowx250-video-fused@f0cea76b62c5dd66b06b9f965932ddea32a7b546` - Initial action decoder: `dreamdifferent/vam-cross-target-widowx250-native-2cam-action-decoder@93750cccda01620e3c028477e4c49bc5c996a68d` - Frozen Video LoRA: `dreamdifferent/vam-cross-level4-so101-widowx-texture-video-lora-iter-400@d66032413191d12b18eb5c97be9f649db3ffcc26` ## Action and data contract - Dataset: `dreamdifferent/vam-cross-level4-so101-widowx-texture@4d2d4b0418eccc9f9398a2745ad2e4ed766a4ef6` - Episodes/frames: 151 / 54340 - Cameras: `observation.images.corner_cam, observation.images.front_cam` - Target: 15 achieved-EE/gripper actions at 5 Hz - Pose target: `relative_to_current_achieved_pose` in `widowx_reference_base/teleop_aligned_tool` - Rotation: `rotation_6d` The dataset and frozen inputs are not included. Use the pinned JSON and effective `config.yaml` included in this repository.