Instructions to use nepyope/pi05-can-30fps-staircase-12k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use nepyope/pi05-can-30fps-staircase-12k with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
pi0.5 G1 — bring the can to the white table, πR² staircase (step 12000)
π₀.₅ fine-tuned from lerobot/pi05_base on whole-body Unitree G1 teleoperation data, trained with
the πR² latency-adaptive staircase noise schedule (arXiv 2607.26055)
instead of the ordinary shared-timestep flow-matching objective.
This is the only checkpoint here that can run --inference.type=pir2, which spends one denoising
step per call and streams actions out continuously rather than planning a chunk at a time. The πR²
engine refuses prefix-trained checkpoints, so the plain fine-tunes cannot be used with it.
| base checkpoint | lerobot/pi05_base |
| dataset | nepyope/can_clean_final_30fps @ 1ffc09d981ff49a6793488d920ad5b20640c2618 — 105 episodes, 127,395 frames, 30 fps |
| robot | unitree_g1, Damiao CAN grippers on both hands |
| cameras | 3: ego_view, left_wrist, right_wrist at 480×640 |
| task | Bring the can to the white table |
| action dim | 66 = 64-D SONIC motion token + 2 grippers |
| state dim | 31 = 29 DOF + 2 grippers, padded to max_state_dim=32 |
chunk_size |
50 — at 30 fps this is 1.67 s of motion |
| noise schedule | staircase, rtc_training_max_delay=10, staircase_time_jitter=0.1, staircase_warmup_prob=0.2 |
| trainable params | 693M of 4.14B (train_expert_only=true, VLM frozen) |
| hardware | 8×H100 80GB, one node, 30.7 GB per GPU |
| batch | 16 per GPU × 8 = 128, no gradient accumulation |
| optimizer | AdamW, LR 1e-4, weight decay 1e-4, betas (0.9, 0.95), grad clip 10.0, cosine decay to 1e-5 with 500 warm-up steps over 12000 |
| throughput | 1.45 s/step, 88 samples/s — 12000 steps in 4h57m (12.06 epochs) |
| code | lerobot branch pir2-staircase @ 0a53c2f2e |
What the staircase schedule does
An ordinary pi0.5 fine-tune samples one noise level per example and shares it across all 50 chunk
positions. The staircase gives each position its own, ramping from clean at the front to pure noise
at the back — which reflects the fact that near-future actions are already committed while distant
ones are not. Concretely the per-position timestep is clamp((i - d) / (H - 2d), 0, 1): a clean
front of d positions, a linear ramp, then a pure-noise tail of d.
d is sampled uniformly from {0..10} during training, so rtc_training_max_delay=10 means the
weights have seen eleven ramp geometries. At inference d is the measured latency in frames — how
many actions the robot consumes while one call is in flight, and equally how many each call emits.
At 30 fps one unit of d is 33.3 ms, so this checkpoint covers per-call latencies up to 333 ms.
That is deliberately generous: the target deployment runs over wifi and showed lag spikes, and d
saturating at the cap means the robot drains the buffer faster than substeps refill it.
staircase_time_jitter=0.1 smears each position's timestep so the position index cannot become a
perfect proxy for its noise level, and the 20% staircase_warmup_prob reverts to the ordinary
shared-timestep objective so the weights can still denoise a chunk from pure noise — needed once per
episode to initialize the buffer.
Loss
| step | 50 | 2k | 4k | 6k | 8k | 10k | 12k |
|---|---|---|---|---|---|---|---|
| epoch | 0.05 | 2.01 | 4.02 | 6.03 | 8.04 | 10.05 | 12.06 |
| loss | 1.239 | 0.096 | 0.081 | 0.076 | 0.070 | 0.063 | 0.062 |
| grad norm | 0.294 | 0.117 | 0.078 | 0.064 | 0.056 | 0.050 | 0.050 |
Flat over the last 2000 steps with the cosine fully annealed, so this is converged rather than truncated. The floor is higher than the 0.025 a plain fine-tune reached on the 50 fps version of this data, which is expected and not a regression: the staircase objective asks the model to denoise at every noise level simultaneously, including positions near pure noise, so its loss is not on the same scale as a shared-timestep run. Training loss, no held-out split.
Running it
--fps=30 is not optional. The policy has no notion of frame rate; it only learned the 33.3 ms
action spacing present in the data. The task string must match verbatim — pi0.5 conditions on it as
text and this checkpoint saw exactly one task.
πR² — one denoising step per call, continuous action stream
lerobot-rollout \
--strategy.type=base \
--policy.path=nepyope/pi05-can-30fps-staircase-12k \
--inference.type=pir2 \
--inference.pir2.max_delay=10 \
--robot.type=unitree_g1 \
--task="Bring the can to the white table" \
--fps=30
Cap max_delay at the trained 10. The engine's own default is chunk_size // 2 = 25 and is not
read from the checkpoint, so leaving it unset lets it choose a d the weights have never seen. It
derives d from a rolling mean over the last 20 calls, so a generous cap does not inflate d when
the link is healthy — it only stops the cap from binding when latency climbs.
RTC guided mode, as a fallback
lerobot-rollout \
--strategy.type=base \
--policy.path=nepyope/pi05-can-30fps-staircase-12k \
--inference.type=rtc \
--inference.rtc.mode=trained \
--inference.rtc.execution_horizon=20 \
--robot.type=unitree_g1 \
--task="Bring the can to the white table" \
--fps=30
mode=trained is available because rtc_training_max_delay=10 > 0. Keep the execution horizon
inside [d, chunk_size - d] = [10, 40].
Known issue: the left gripper is dead
Inherited from the data. observation.state[29] and action[64] are identically 0.0 in every
frame — the episodes were effectively recorded one-handed, and their q01/q99 are set by hand to
0/1 so quantile normalization does not divide by zero. This policy will never open or close
the left hand. The right gripper is healthy, closed in 41% of frames.
Training command
export PYTHONPATH=/path/to/draccus-overlay:/path/to/lerobot-rtc-b2/src
accelerate launch --num_processes=8 --mixed_precision=bf16 \
-m lerobot.scripts.lerobot_train \
--policy.type=pi05 \
--policy.pretrained_path=lerobot/pi05_base \
--policy.max_state_dim=32 --policy.max_action_dim=66 \
--policy.chunk_size=50 --policy.n_action_steps=50 \
--policy.rtc_training_schedule=staircase \
--policy.rtc_training_max_delay=10 \
--policy.staircase_time_jitter=0.1 \
--policy.train_expert_only=true \
--policy.freeze_vision_encoder=false \
--policy.gradient_checkpointing=true \
--policy.compile_model=false \
--policy.push_to_hub=false \
--dataset.repo_id=nepyope/can_clean_final_30fps \
--dataset.root=/path/to/can_clean_final_30fps \
--dataset.use_imagenet_stats=false \
--batch_size=16 --num_workers=10 --steps=12000 \
--tolerance_s=0.001 \
--use_policy_training_preset=false \
--optimizer.type=adamw --optimizer.lr=1e-4 --optimizer.weight_decay=1e-4 \
--optimizer.betas="[0.9,0.95]" --optimizer.grad_clip_norm=10.0 \
--scheduler.type=cosine_decay_with_warmup \
--scheduler.num_warmup_steps=500 --scheduler.num_decay_steps=12000 \
--scheduler.peak_lr=1e-4 --scheduler.decay_lr=1e-5 \
--wandb.enable=true --wandb.project=can-30fps-staircase --wandb.disable_artifact=true \
--save_freq=2000 --log_freq=50 \
--output_dir=/path/to/output --job_name=can30_staircase_12k
Four things this run needed
1. A patch to from_pretrained (required — without it you train from scratch).
max_action_dim=66 disagrees with pi05_base's 32, so action_in_proj / action_out_proj cannot
be loaded. In PI05Policy.from_pretrained the resulting load_state_dict error is swallowed by a
broad except, which returns a randomly initialized model that then trains and logs perfectly
normally. A local patch drops only the shape-mismatched tensors so the backbone loads and just those
two projections start fresh:
Dropping 3 shape-mismatched keys (re-initialized):
- model.action_in_proj.weight: ckpt (1024, 32) vs model (1024, 66)
- model.action_out_proj.bias: ckpt (32,) vs model (66,)
- model.action_out_proj.weight: ckpt (32, 1024) vs model (66, 1024)
num_learnable_params=693M of 4.14B confirms the frozen VLM did load. Verify those lines appear —
without them the run silently trains from noise.
2. A newer draccus than the env had. This branch calls draccus.encode(self, PreTrainedConfig)
with two arguments; draccus 0.8.0's encode takes one. The first attempt trained 2000 steps
perfectly and then died with TypeError: encode() takes 1 positional argument but 2 were given at
the moment it tried to write the checkpoint config, leaving a zero-byte config.json and no
weights. Putting a newer draccus first on PYTHONPATH fixes it.
3. --dataset.use_imagenet_stats=false. make_dataset assumes every camera key has an entry in
meta/stats.json and raises KeyError: 'observation.images.ego_view' otherwise; this dataset's
stats cover only action, observation.state and the index columns. Harmless to disable, since
pi0.5's VISUAL normalization is IDENTITY and image statistics are never read.
4. --tolerance_s=0.001. The default 1e-4 is tighter than float arithmetic on timestamps a
thousand seconds into a video chunk, which trips FrameTimestampError on a 0.2 ms discrepancy —
one hundredth of a frame interval.
--dataset.root bypasses Hub revision resolution and guarantees training reads 1ffc09d98.
Note on the logged throughput
Unlike runs on main, the updt_s / step_s / mem_gb figures in this branch's logs are correct.
On main, MetricsTracker.reduce_across_ranks asks accelerate for a max reduction that it does
not implement, so those metrics come out summed over ranks — inflated by the world size.
- Downloads last month
- 25