pi05-anyh2r-rh56f1-wristik-grasp-mirror-20k

pi0.5 fine-tuned from lerobot/pi05_base on mixed human-to-robot data for a bimanual OpenArm platform with RH56F1 hands.

Adds left/right mirror augmentation (p=0.5) on top of the wrist-IK + grasp relabelling. The left arm moves substantially in only 3 of 16 cells, so half of every batch is mirrored -- actions, stereo images and the instruction together -- which raises the left arm's exposure from 29% to 53% of sampled windows without changing any cell's sampling share. Evaluated on held-in training frames, this lifts the left-arm reach from 0.81x to 0.98x of the ground-truth chunk on ball_in_cup, and 0.94x to 1.00x on air_fryer.

Training

Base lerobot/pi05_base
Framework LeRobot 0.6.1
Steps 20,000 (of a 30,000-step run)
Global batch 64 (16 x 4 H100)
Precision bfloat16, gradient checkpointing
Seed 1000
Final train loss 0.013 (grad norm 0.146, 4.65 epochs)

Data

A merged LeRobot v3.0 dataset of 855 episodes / 245,677 frames, drawn from 16 (source x category) cells with equal sampling probability per cell, so the model sees each category equally often despite a ~6x spread in cell size.

  • Synthetic (human-to-robot retargeted), 12 categories: depth IDM with wrist-IK and grasp refinement, 12 categories, mirrored 50% at load time
  • Real teleoperation, 4 categories, filtered to episodes whose measured content rate is

    = 19 Hz in both stereo views

Interface

Two cameras are mapped onto the pi0.5 slots:

observation.images.camera_ego_left  -> observation.images.base_0_rgb
observation.images.camera_ego_right -> observation.images.left_wrist_0_rgb
  • observation.state / action: 28 dims (the policy pads to its 32-dim width)
  • Action chunk: 50 steps at 20 fps, i.e. a 2.5 s horizon predicted per inference
  • Normalization: QUANTILES for state and action, IDENTITY for images -- the q01/q99 statistics are baked into the processor files in this repo
  • The neck joints (state dims 0-1) were held fixed during collection and are expected to stay fixed at inference

Load it the usual way:

from lerobot.policies.pi05 import PI05Policy
policy = PI05Policy.from_pretrained("RyanL22/pi05-anyh2r-rh56f1-wristik-grasp-mirror-20k")
Downloads last month
21
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for RyanL22/pi05-anyh2r-rh56f1-wristik-grasp-mirror-20k

Finetuned
(773)
this model