pi05_mvtoken_22_27_04
pi0.5 SFT checkpoints fine-tuned on the MVTOKEN_22_27_04 real-robot dataset.
Checkpoints
Two checkpoints from the same run, laid out as in the training output directory:
global_step_2000/actor/model_state_dict/full_weights.pt # train loss ~0.014, 93 passes
global_step_4000/actor/model_state_dict/full_weights.pt # train loss ~0.008, 177 passes
norm_stats.json # shared by both
Evaluate both. The dataset is small (5,070 frames), so by step 4000 every frame has been
seen ~177 times. For comparison, the SweepIntoDustpan-v1_Real run reached the same loss
level after only ~15 passes — a lower training loss here does not imply a better policy.
norm_stats.json is the same file as in
aaroncaozj/pi05_norm_stats_collection
→ MVTOKEN_22_27_04/.
FSDP distributed-checkpoint shards (dcp_checkpoint/*.distcp, ~22 GB each) are not included;
they only serve to resume training and are not needed for inference.
Data
| episodes | 101 |
| frames | 5,070 |
| fps | 4 |
| duration / episode | ~12.5 s |
| state | 7-D end-effector pose + gripper width |
| actions | 7-D delta EE pose + binary gripper, ~2 cm per step |
| cameras | image (agentview) + wrist_image; back_image is an all-zero placeholder and is masked out per sample by FrankaEEInputs |
Training
Trained with RLinf (examples/sft/config/mvtoken_sft_openpi.yaml),
openpi data config pi05_mvtoken.
| base model | pi05_base |
| action_horizon | 4 |
| hardware | 7 x H200 NVL |
| micro / global batch | 32 / 224 (grad accum 1) |
| optimizer | AdamW, lr 2.5e-5, cosine, 1000 warmup |