DiscoDemo
Collection
Datasets, RL generators and pi0.5 SFT policies from DiscoDemo. Project page: https://davian-robotics.github.io/DiscoDemo/ • 28 items • Updated
The data generator of DiscoDemo for the FMB-Round task (insert a cylindrical peg into its hole on the FMB board): a state-based reinforcement-learning policy trained in simulation — the P-RFCL baseline generator (α = 0: the same reverse-curriculum RL without the diversity reward). Rolling it out generates demonstrations such as those in the Stage2 dataset below.
| Stage | Repository |
|---|---|
| Generator (RL policy) | DAVIAN-Robotics/DiscoDemo-Stage1_RL-FMB_Round-alpha0 (this repository) |
| Generated dataset | DAVIAN-Robotics/DiscoDemo-Stage2_GenData-FMB_Round-alpha0 |
| Imitation policy (π0.5) | DAVIAN-Robotics/DiscoDemo-Stage3_SFT-FMB_Round-alpha0 |
| File | Content |
|---|---|
actor.pt |
Actor network weights (network_state_dict). Optimizer state is not included. |
normalizer.pt |
Observation bounds used to normalize the policy input. |
config.yaml |
Full training configuration (environment, curriculum, agent). |
Only what is needed to roll the policy out is released; the critic and other training-only state are omitted.
Environment steps at this checkpoint: 46.5 M.
Loading requires the DiscoDemo training code: https://github.com/DAVIAN-Robotics/DiscoDemo.
@article{park2026discodemo,
title = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
author = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
journal = {arXiv preprint},
year = {2026}
}