pengyue-polaron's picture
Make model card concise and move training details out of the overview
62ccdbd verified
|
Raw History Blame Contribute Delete
752 Bytes

Training

Setting Value
Optimizer steps 100
Hardware 2 × NVIDIA A100
Distributed strategy Full-parameter FSDP
Precision bfloat16
Effective global batch size 16
Optimizer Fused AdamW
Learning rate 1e-5 with 10-step warmup, then constant
Adam betas / weight decay (0.9, 0.95) / 0.1
Objective Video latent loss + action loss

Across the 100 optimizer steps, the mean video-latent loss was 0.161962 and the mean action loss was 0.020129.

The eight source action values are mapped to LingBot-VA action channels [0, 1, 2, 3, 4, 5, 6, 28]; channel 28 carries the gripper command. The exact normalization statistics and model settings are included in configs/va_a1_cfg.py.