Robotics
LeRobot
Safetensors
paligemma
pi05
piperx
intervention
fine-tuning

Pi0.5 PiperX: D3 + own human-intervention repair

Inference weights after 8,884 new fine-tuning steps, initialized from the original 3h demonstration model at step 12,637. Task: Insert the copper screw into the black sleeve.

Data and training

  • Replay: original seed-1000 D3 subset, 149 episodes / 323,507 frames (~2.9954h). Exact episode selection and full deterministic ranking are included.
  • Own-policy intervention source: linked D3 full-rollout dataset, pinned revision 91c68b1fbfc832ced5a64299f2986b344a4f1874 (82 full episodes / 463,814 frames).
  • All 113,713 human-controlled frames (~1h3m10s) were extracted into 748 continuous segments. Automatic frames excluded; all 11 segments shorter than 50 frames retained (minimum 16 frames).
  • Strict 1:1 sampling: each rank draws 4 demo + 4 intervention examples, 8 GPUs, global batch 64. Seed 1000.
  • 2.5 intervention exposure epochs: ceil(2.5 * 113713 / 32) = 8884 updates, not 2.5 combined-dataset epochs.
  • Full-parameter FP32, AMP/compile disabled, gradient checkpointing enabled. AdamW betas (0.9,0.95), epsilon 1e-8, weight decay 0.01, gradient clip 1.0.
  • Fresh optimizer/scheduler: LR 2.5e-6 decaying to 2.5e-7, warmup 200. Final checkpoint only.
  • Temporal action padding masked from loss. Original D3 normalization retained; intervention statistics recorded separately.
  • LeRobot 0.6.1, PyTorch 2.11.0+cu128, torchvision 0.26.0+cu128, Transformers 5.5.4, Accelerate 1.14.0; eight RTX PRO 6000 GPUs.
  • Inputs were read directly from shared persistent storage. Training completed 8884 updates at 2026-09-14 04:55 China time; final logged loss at step 8880 was 0.004.

Inference export and integrity

The final checkpoint save encountered a 10-minute distributed synchronization timeout while other ranks waited for the saving rank, and the training process exited abnormally. This history is preserved in the supplied log/status. The exported policy weights were independently reloaded strictly (813 tensors), checked for finite values, and successfully used for real-observation GPU inference producing finite [1,50,14] actions. These are software integrity checks, not a robot success-rate evaluation. The optimizer-resume state is not certified and is not uploaded.

Weights and normalization tensors are unchanged. Only the training-specific mask_action_padding_loss field is removed from the inference config for stock LeRobot compatibility; the original config is included under reproducibility/. This flag does not affect inference. Tokenizer files are bundled at repository root and the processor reference points to this repository. For fully offline use, override tokenizer_processor.tokenizer_name to the local downloaded repository directory.

Use the saved processors and suitable camera rename map. Canonical cameras: observation.images.base_0_rgb, observation.images.left_wrist_0_rgb, observation.images.right_wrist_0_rgb. Robot state/action dimension 14; action horizon 50.

Visible reproducibility files include the D3 seed manifest and ranking, intervention source manifest and segment mapping, normalization statistics, frozen-code hashes and patch, actual direct-read training command, scripts, logs and validation reports. Machine-specific paths in historical scripts require adaptation on another host. This is an inference export, not an exact optimizer-resume bundle.

Model SHA256: f22d871e947c497fbbad9a750912f68d8d873c11b9ce8ab9d380543519a3d95d.

Downloads last month
28
Safetensors
Model size
4B params
Tensor type
F32
·
Video Preview
loading

Model tree for Elvinky/pi05-piperx-3h-intervention-1to1-mask-8884s-seed1000-fp32

Finetuned
(1)
this model

Datasets used to train Elvinky/pi05-piperx-3h-intervention-1to1-mask-8884s-seed1000-fp32