Instructions to use nepyope/pi05-tshirt-staircase-8k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use nepyope/pi05-tshirt-staircase-8k with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
pi0.5 G1 t-shirt — πR² staircase schedule (step 8000)
π₀.₅ fine-tuned on the πR² latency-adaptive staircase noise schedule
(arXiv 2607.26055) rather than the usual shared-timestep flow
objective. Instead of drawing one noise level per sample and sharing it across all 50 chunk
positions, each position gets its own, ramping from clean at the front to pure noise at the back.
That is the schedule the πR² inference engine reproduces as a fixed point, so one denoising step
per call finishes d actions which are emitted immediately, the buffer slides, and d fresh-noise
slots are appended. The robot is fed continuously instead of waiting for a whole chunk.
20% of examples fall back to the ordinary shared-timestep objective (staircase_warmup_prob=0.2), so
the weights retain the ability to denoise a chunk from pure noise — needed once to initialize the
buffer at the start of an episode.
This checkpoint is usable with --inference.type=pir2, which refuses prefix-trained checkpoints.
| base checkpoint | lerobot/pi05_base |
| dataset | nepyope/t-shirt_pick_and_place_clean — 47 episodes, 101,254 frames, 50 fps |
| robot | unitree_g1, 3 cameras (ego_view, left_wrist, right_wrist) at 480×640 |
| task | put the t-shirt on the table |
| action / state dim | 66 (64 joints + 2 grippers) / 31 (29 DOF + 2 grippers), padded to max_state_dim=32 |
| code | huggingface/lerobot @ 0a53c2f2e (pir2-staircase, PR #4427), stacked on training-time RTC (PR #4056) |
rtc_training_schedule |
staircase |
rtc_training_max_delay |
1 (delays drawn uniformly from {0, 1} per example) |
staircase_time_jitter |
0.1 |
staircase_warmup_prob |
0.2 (default) |
chunk_size |
50 — at 50 fps this is 1.0 s of motion |
| trainable params | 693M of 4.14B (train_expert_only=true, VLM frozen) |
| hardware | 4×H100 80GB, one node, 39.9 GB per GPU |
| batch | 32 per GPU × 4 = 128, no gradient accumulation |
| optimizer | AdamW, LR 1e-4, weight decay 1e-4, betas (0.9, 0.95), cosine decay to 1e-5 with 500 warm-up steps over 8000 |
| throughput | 2.70 s/step, 47 samples/s — 8000 steps in 6.0 h (10.1 epochs) |
Training command
accelerate launch --num_processes=4 --mixed_precision=bf16 \
-m lerobot.scripts.lerobot_train \
--policy.type=pi05 \
--policy.pretrained_path=lerobot/pi05_base \
--policy.max_state_dim=32 --policy.max_action_dim=66 \
--policy.rtc_training_schedule=staircase \
--policy.rtc_training_max_delay=1 \
--policy.staircase_time_jitter=0.1 \
--policy.train_expert_only=true \
--policy.freeze_vision_encoder=false \
--policy.gradient_checkpointing=true \
--policy.push_to_hub=false \
--policy.chunk_size=50 --policy.n_action_steps=50 \
--dataset.repo_id=nepyope/t-shirt_pick_and_place_clean \
--dataset.root=/path/to/t-shirt_pick_and_place_clean \
--batch_size=32 --num_workers=10 --steps=8000 \
--use_policy_training_preset=false \
--optimizer.type=adamw --optimizer.lr=1e-4 --optimizer.weight_decay=1e-4 \
--optimizer.betas="[0.9,0.95]" \
--scheduler.type=cosine_decay_with_warmup \
--scheduler.num_warmup_steps=500 --scheduler.num_decay_steps=8000 \
--scheduler.peak_lr=1e-4 --scheduler.decay_lr=1e-5 \
--wandb.enable=true --wandb.project=tshirt-staircase --wandb.disable_artifact=true \
--save_freq=2000 --log_freq=50 \
--output_dir=/path/to/output --job_name=tshirt_staircase_8k
Two things this run needed that are not on the branch
Quantile stats. pi0.5 normalizes state and action with NormalizationMode.QUANTILES, but the
dataset shipped with only count/max/mean/min/std, so training aborts on the first batch with
QUANTILES normalization mode requires q01 and q99 stats. Fixed by copying the dataset to a
writable location (dereferencing symlinks, so the shared HF cache blobs are never written through)
and running:
python src/lerobot/scripts/augment_dataset_quantile_stats.py \
--repo-id=nepyope/t-shirt_pick_and_place_clean \
--root=/path/to/writable/copy --skip-images --overwrite
--skip-images is safe and fast here because pi0.5 maps VISUAL to IDENTITY, so no video needs
decoding. Note the script calls push_to_hub() unconditionally after writing meta/stats.json
locally, so it can appear to fail after having already done the useful work. The resulting stats are
baked into this checkpoint's preprocessor.
Shape-mismatch handling on load. max_action_dim=66 disagrees with pi05_base's 32, so
action_in_proj / action_out_proj cannot be loaded. In PI05Policy.from_pretrained the resulting
load_state_dict error is swallowed by a broad except, which returns a randomly initialized
model that then trains and logs normally. A local patch drops only the shape-mismatched tensors so
the backbone loads and just those two projections start fresh. Without it this run would have
silently trained from scratch.
Loss
| epoch | 0.06 | 0.19 | 1.64 | 3.22 | 4.80 | 6.38 | 7.96 | 10.11 |
|---|---|---|---|---|---|---|---|---|
| loss | 1.245 | 0.633 | 0.111 | 0.094 | 0.085 | 0.078 | 0.075 | 0.073 |
| grad norm | 0.288 | 1.346 | 0.190 | 0.092 | 0.072 | 0.061 | 0.056 | 0.054 |
Flat from about epoch 8 onward at 0.072–0.073 with the cosine schedule fully annealed to 1e-5, so this is a converged run rather than a truncated one.
Inference
lerobot-rollout \
--strategy.type=base \
--policy.path=<this repo or a local download> \
--inference.type=pir2 \
--inference.max_delay=1 \
--inference.latency_window=20 \
--robot.type=unitree_g1 \
--task="put the t-shirt on the table" \
--fps=50
--fps=50 is not optional: the dataset is 50 fps, so consecutive actions are 20 ms apart and the
policy has no notion of frame rate beyond that spacing. Running at 30 fps executes the motion 1.67×
slower than demonstrated.
--inference.max_delay=1 matters. The engine derives the per-call delay as
max(1, min(max_delay, round(mean_latency / (1/fps)))) from a rolling window of measured action-head
latencies, and max_delay defaults to chunk_size // 2 = 25 — it is not clamped to the value used
in training. Since this checkpoint only ever saw d ∈ {0, 1}, letting the engine pick d=2 or more
extrapolates past the trained schedule. A 20 ms tick is tight (one expert substep against a cached
prefix measures ~12 ms on an RTX 5090 laptop, before capture and transport overhead), so that is a
realistic risk rather than a theoretical one. Pinning max_delay=1 keeps inference in distribution;
if calls do overrun, the buffer simply runs a cycle behind rather than going off-schedule.
A future run with --policy.rtc_training_max_delay=3 would cover d up to 3 (60 ms per call at
50 fps) while keeping d=1 in distribution, and would be the more deployable choice.
- Downloads last month
- 12