DiscoDemo
Collection
Datasets, RL generators and pi0.5 SFT policies from DiscoDemo. Project page: https://davian-robotics.github.io/DiscoDemo/ • 28 items • Updated
The data generator of DiscoDemo for the FMB-SqCircle task (insert a two-pronged (cylinder and square) peg into its hole on the FMB board): a state-based reinforcement-learning policy trained in simulation — the DiscoDemo generator (diversity weight α = 0.3). Rolling it out generates demonstrations such as those in the Stage2 dataset below.
| Stage | Repository |
|---|---|
| Generator (RL policy) | DAVIAN-Robotics/DiscoDemo-Stage1_RL-FMB_SqCircle (this repository) |
| Generated dataset | DAVIAN-Robotics/DiscoDemo-Stage2_GenData-FMB_SqCircle |
| Imitation policy (π0.5) | DAVIAN-Robotics/DiscoDemo-Stage3_SFT-FMB_SqCircle |
| File | Content |
|---|---|
actor.pt |
Actor network weights (network_state_dict). Optimizer state is not included. |
normalizer.pt |
Observation bounds used to normalize the policy input. |
config.yaml |
Full training configuration (environment, curriculum, agent). |
Only what is needed to roll the policy out is released; the critic and other training-only state are omitted. The actor is conditioned on a skill vector z sampled once per episode from a standard normal distribution.
Environment steps at this checkpoint: 32.5 M.
Loading requires the DiscoDemo training code: https://github.com/DAVIAN-Robotics/DiscoDemo.
@article{park2026discodemo,
title = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
author = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
journal = {arXiv preprint},
year = {2026}
}