DiscoDemo-Stage1_RL-PnP_Banana

The data generator of DiscoDemo for the PnP-Banana task (pick up a banana and place it in a bowl): a state-based reinforcement-learning policy trained in simulation — the DiscoDemo generator (diversity weight α = 0.3). Rolling it out generates demonstrations such as those in the Stage2 dataset below.

Stage Repository
Generator (RL policy) DAVIAN-Robotics/DiscoDemo-Stage1_RL-PnP_Banana (this repository)
Generated dataset DAVIAN-Robotics/DiscoDemo-Stage2_GenData-PnP_Banana
Imitation policy (π0.5) DAVIAN-Robotics/DiscoDemo-Stage3_SFT-PnP_Banana

Files

File Content
actor.pt Actor network weights (network_state_dict). Optimizer state is not included.
normalizer.pt Observation bounds used to normalize the policy input.
config.yaml Full training configuration (environment, curriculum, agent).

Only what is needed to roll the policy out is released; the critic and other training-only state are omitted. The actor is conditioned on a skill vector z sampled once per episode from a standard normal distribution.

Environment steps at this checkpoint: 50.0 M.

Usage

Loading requires the DiscoDemo training code: https://github.com/DAVIAN-Robotics/DiscoDemo.

Citation

@article{park2026discodemo,
  title   = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
  author  = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
  journal = {arXiv preprint},
  year    = {2026}
}
Downloads last month
21
Video Preview
loading

Collection including DAVIAN-Robotics/DiscoDemo-Stage1_RL-PnP_Banana