DiscoDemo-Stage3_SFT-StackCube

A visuomotor π0.5 policy for the StackCube task (stack the red cube on the blue cube), fine-tuned on demonstrations generated by the DiscoDemo generator (diversity weight α = 0.3) of DiscoDemo. It acts from camera images and joint states only, and transfers to the real Franka FR3 robot without real-world fine-tuning.

Stage Repository
Generator (RL policy) DAVIAN-Robotics/DiscoDemo-Stage1_RL-StackCube
Generated dataset DAVIAN-Robotics/DiscoDemo-Stage2_GenData-StackCube
Imitation policy (π0.5) DAVIAN-Robotics/DiscoDemo-Stage3_SFT-StackCube (this repository)

Training

Base model DAVIAN-Robotics/pi05_droid_jointpos (π0.5 DROID, joint-position actions)
Data DAVIAN-Robotics/DiscoDemo-Stage2_GenData-StackCube
Steps 20,000
Batch size 16
Action chunk 15 (relative to the current joint state)
Cameras over-the-shoulder and wrist
Seed 1000

Usage

The policy is a LeRobot π0.5 checkpoint. Its pre- and post-processing pipeline uses DiscoDemo processing steps (relative actions, state tokenization), so loading it requires the DiscoDemo code: https://github.com/DAVIAN-Robotics/DiscoDemo.

License

Apache-2.0, following the base model (openpi).

Citation

@article{park2026discodemo,
  title   = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
  author  = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
  journal = {arXiv preprint},
  year    = {2026}
}
Downloads last month
9
Safetensors
Model size
4B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for DAVIAN-Robotics/DiscoDemo-Stage3_SFT-StackCube

Finetuned
(13)
this model

Collection including DAVIAN-Robotics/DiscoDemo-Stage3_SFT-StackCube