pi0.5 G1 — bring the can to the white table, πR² staircase (step 12000)

π₀.₅ fine-tuned from lerobot/pi05_base on whole-body Unitree G1 teleoperation data, trained with the πR² latency-adaptive staircase noise schedule (arXiv 2607.26055) instead of the ordinary shared-timestep flow-matching objective.

This is the only checkpoint here that can run --inference.type=pir2, which spends one denoising step per call and streams actions out continuously rather than planning a chunk at a time. The πR² engine refuses prefix-trained checkpoints, so the plain fine-tunes cannot be used with it.

base checkpoint lerobot/pi05_base
dataset nepyope/can_clean_final_30fps @ 1ffc09d981ff49a6793488d920ad5b20640c2618 — 105 episodes, 127,395 frames, 30 fps
robot unitree_g1, Damiao CAN grippers on both hands
cameras 3: ego_view, left_wrist, right_wrist at 480×640
task Bring the can to the white table
action dim 66 = 64-D SONIC motion token + 2 grippers
state dim 31 = 29 DOF + 2 grippers, padded to max_state_dim=32
chunk_size 50 — at 30 fps this is 1.67 s of motion
noise schedule staircase, rtc_training_max_delay=10, staircase_time_jitter=0.1, staircase_warmup_prob=0.2
trainable params 693M of 4.14B (train_expert_only=true, VLM frozen)
hardware 8×H100 80GB, one node, 30.7 GB per GPU
batch 16 per GPU × 8 = 128, no gradient accumulation
optimizer AdamW, LR 1e-4, weight decay 1e-4, betas (0.9, 0.95), grad clip 10.0, cosine decay to 1e-5 with 500 warm-up steps over 12000
throughput 1.45 s/step, 88 samples/s — 12000 steps in 4h57m (12.06 epochs)
code lerobot branch pir2-staircase @ 0a53c2f2e

What the staircase schedule does

An ordinary pi0.5 fine-tune samples one noise level per example and shares it across all 50 chunk positions. The staircase gives each position its own, ramping from clean at the front to pure noise at the back — which reflects the fact that near-future actions are already committed while distant ones are not. Concretely the per-position timestep is clamp((i - d) / (H - 2d), 0, 1): a clean front of d positions, a linear ramp, then a pure-noise tail of d.

d is sampled uniformly from {0..10} during training, so rtc_training_max_delay=10 means the weights have seen eleven ramp geometries. At inference d is the measured latency in frames — how many actions the robot consumes while one call is in flight, and equally how many each call emits. At 30 fps one unit of d is 33.3 ms, so this checkpoint covers per-call latencies up to 333 ms. That is deliberately generous: the target deployment runs over wifi and showed lag spikes, and d saturating at the cap means the robot drains the buffer faster than substeps refill it.

staircase_time_jitter=0.1 smears each position's timestep so the position index cannot become a perfect proxy for its noise level, and the 20% staircase_warmup_prob reverts to the ordinary shared-timestep objective so the weights can still denoise a chunk from pure noise — needed once per episode to initialize the buffer.

Loss

step 50 2k 4k 6k 8k 10k 12k
epoch 0.05 2.01 4.02 6.03 8.04 10.05 12.06
loss 1.239 0.096 0.081 0.076 0.070 0.063 0.062
grad norm 0.294 0.117 0.078 0.064 0.056 0.050 0.050

Flat over the last 2000 steps with the cosine fully annealed, so this is converged rather than truncated. The floor is higher than the 0.025 a plain fine-tune reached on the 50 fps version of this data, which is expected and not a regression: the staircase objective asks the model to denoise at every noise level simultaneously, including positions near pure noise, so its loss is not on the same scale as a shared-timestep run. Training loss, no held-out split.

Running it

--fps=30 is not optional. The policy has no notion of frame rate; it only learned the 33.3 ms action spacing present in the data. The task string must match verbatim — pi0.5 conditions on it as text and this checkpoint saw exactly one task.

πR² — one denoising step per call, continuous action stream

lerobot-rollout \
  --strategy.type=base \
  --policy.path=nepyope/pi05-can-30fps-staircase-12k \
  --inference.type=pir2 \
  --inference.pir2.max_delay=10 \
  --robot.type=unitree_g1 \
  --task="Bring the can to the white table" \
  --fps=30

Cap max_delay at the trained 10. The engine's own default is chunk_size // 2 = 25 and is not read from the checkpoint, so leaving it unset lets it choose a d the weights have never seen. It derives d from a rolling mean over the last 20 calls, so a generous cap does not inflate d when the link is healthy — it only stops the cap from binding when latency climbs.

RTC guided mode, as a fallback

lerobot-rollout \
  --strategy.type=base \
  --policy.path=nepyope/pi05-can-30fps-staircase-12k \
  --inference.type=rtc \
  --inference.rtc.mode=trained \
  --inference.rtc.execution_horizon=20 \
  --robot.type=unitree_g1 \
  --task="Bring the can to the white table" \
  --fps=30

mode=trained is available because rtc_training_max_delay=10 > 0. Keep the execution horizon inside [d, chunk_size - d] = [10, 40].

Known issue: the left gripper is dead

Inherited from the data. observation.state[29] and action[64] are identically 0.0 in every frame — the episodes were effectively recorded one-handed, and their q01/q99 are set by hand to 0/1 so quantile normalization does not divide by zero. This policy will never open or close the left hand. The right gripper is healthy, closed in 41% of frames.

Training command

export PYTHONPATH=/path/to/draccus-overlay:/path/to/lerobot-rtc-b2/src

accelerate launch --num_processes=8 --mixed_precision=bf16 \
  -m lerobot.scripts.lerobot_train \
  --policy.type=pi05 \
  --policy.pretrained_path=lerobot/pi05_base \
  --policy.max_state_dim=32 --policy.max_action_dim=66 \
  --policy.chunk_size=50 --policy.n_action_steps=50 \
  --policy.rtc_training_schedule=staircase \
  --policy.rtc_training_max_delay=10 \
  --policy.staircase_time_jitter=0.1 \
  --policy.train_expert_only=true \
  --policy.freeze_vision_encoder=false \
  --policy.gradient_checkpointing=true \
  --policy.compile_model=false \
  --policy.push_to_hub=false \
  --dataset.repo_id=nepyope/can_clean_final_30fps \
  --dataset.root=/path/to/can_clean_final_30fps \
  --dataset.use_imagenet_stats=false \
  --batch_size=16 --num_workers=10 --steps=12000 \
  --tolerance_s=0.001 \
  --use_policy_training_preset=false \
  --optimizer.type=adamw --optimizer.lr=1e-4 --optimizer.weight_decay=1e-4 \
  --optimizer.betas="[0.9,0.95]" --optimizer.grad_clip_norm=10.0 \
  --scheduler.type=cosine_decay_with_warmup \
  --scheduler.num_warmup_steps=500 --scheduler.num_decay_steps=12000 \
  --scheduler.peak_lr=1e-4 --scheduler.decay_lr=1e-5 \
  --wandb.enable=true --wandb.project=can-30fps-staircase --wandb.disable_artifact=true \
  --save_freq=2000 --log_freq=50 \
  --output_dir=/path/to/output --job_name=can30_staircase_12k

Four things this run needed

1. A patch to from_pretrained (required — without it you train from scratch). max_action_dim=66 disagrees with pi05_base's 32, so action_in_proj / action_out_proj cannot be loaded. In PI05Policy.from_pretrained the resulting load_state_dict error is swallowed by a broad except, which returns a randomly initialized model that then trains and logs perfectly normally. A local patch drops only the shape-mismatched tensors so the backbone loads and just those two projections start fresh:

Dropping 3 shape-mismatched keys (re-initialized):
  - model.action_in_proj.weight: ckpt (1024, 32) vs model (1024, 66)
  - model.action_out_proj.bias: ckpt (32,) vs model (66,)
  - model.action_out_proj.weight: ckpt (32, 1024) vs model (66, 1024)

num_learnable_params=693M of 4.14B confirms the frozen VLM did load. Verify those lines appear — without them the run silently trains from noise.

2. A newer draccus than the env had. This branch calls draccus.encode(self, PreTrainedConfig) with two arguments; draccus 0.8.0's encode takes one. The first attempt trained 2000 steps perfectly and then died with TypeError: encode() takes 1 positional argument but 2 were given at the moment it tried to write the checkpoint config, leaving a zero-byte config.json and no weights. Putting a newer draccus first on PYTHONPATH fixes it.

3. --dataset.use_imagenet_stats=false. make_dataset assumes every camera key has an entry in meta/stats.json and raises KeyError: 'observation.images.ego_view' otherwise; this dataset's stats cover only action, observation.state and the index columns. Harmless to disable, since pi0.5's VISUAL normalization is IDENTITY and image statistics are never read.

4. --tolerance_s=0.001. The default 1e-4 is tighter than float arithmetic on timestamps a thousand seconds into a video chunk, which trips FrameTimestampError on a 0.2 ms discrepancy — one hundredth of a frame interval.

--dataset.root bypasses Hub revision resolution and guarantees training reads 1ffc09d98.

Note on the logged throughput

Unlike runs on main, the updt_s / step_s / mem_gb figures in this branch's logs are correct. On main, MetricsTracker.reduce_across_ranks asks accelerate for a max reduction that it does not implement, so those metrics come out summed over ranks — inflated by the world size.

Downloads last month
25
Safetensors
Model size
4B params
Tensor type
F32
·
Video Preview
loading

Model tree for nepyope/pi05-can-30fps-staircase-12k

Finetuned
(725)
this model

Dataset used to train nepyope/pi05-can-30fps-staircase-12k

Paper for nepyope/pi05-can-30fps-staircase-12k