pi0.5 โ€” anyh2r OpenArm + RH56F1, frozen SigLIP + stereo-consistent image aug (20k)

LeRobot-native pi05 fine-tuned on the 2026-09-13 anyh2r mixture (12 synthetic/IDM cells from human video + 4 real teleop cells).

Checkpoint: step 20,000 of a 60,000-step run.

What is different from the previous run

previous (...-wristik-grasp-mirror-20k) this run
vision encoder fine-tuned frozen (freeze_vision_encoder=true, 412.4M params)
image augmentation none photometric + affine, one draw replayed across the stereo pair
cell shares equal per cell โˆ sqrt(frames) (per-frame exposure spread 4.59x -> 2.14x)
mirror augmentation p=0.5 p=0.5 (unchanged)

Why

The previous checkpoint chose its motion mode from the image domain: holding the robot state fixed and swapping only the image source (real teleop frame vs generated frame) changed the predicted action amplitude by 1.35-1.93x. That made rollouts on real camera images reproduce only 27-65% of the trained motion, while in-distribution chunk reproduction was ~1.00.

Freezing the vision tower removes that dependence.

cross-domain probe (same state, image source swapped) previous 20k this 20k
ball 1.48x 1.16x
bottle 1.87x 0.91x
box 1.93x 1.00x
doll 1.35x 1.01x

Absolute amplitude is preserved: in-distribution chunk reproduction stays at 0.85-1.09 of ground truth, and predicted amplitude on real rollout frames is unchanged or higher.

Inputs

  • observation.images.base_0_rgb <- left ZED view, 288x512
  • observation.images.left_wrist_0_rgb <- right ZED view, 288x512
  • observation.state โ€” 28 dims: neck(2) | left_arm(7) | right_arm(7) | left_hand(6) | right_hand(6)
  • 20 fps, chunk 50 (2.5 s)
Downloads last month
19
Safetensors
Model size
4B params
Tensor type
F32
ยท
BF16
ยท
Video Preview
loading

Model tree for RyanL22/pi05-anyh2r-rh56f1-frozen-aug-20k

Finetuned
(776)
this model