--- license: apache-2.0 tags: - robotics - reinforcement-learning - discodemo - franka --- # DiscoDemo-Stage1_RL-FMB_Round-alpha0 The **data generator** of [DiscoDemo](https://davian-robotics.github.io/DiscoDemo/) for the **FMB-Round** task (insert a cylindrical peg into its hole on the FMB board): a state-based reinforcement-learning policy trained in simulation — the P-RFCL baseline generator (α = 0: the same reverse-curriculum RL without the diversity reward). Rolling it out generates demonstrations such as those in the Stage2 dataset below. | Stage | Repository | |---|---| | Generator (RL policy) | [`DAVIAN-Robotics/DiscoDemo-Stage1_RL-FMB_Round-alpha0`](https://huggingface.co/DAVIAN-Robotics/DiscoDemo-Stage1_RL-FMB_Round-alpha0) (this repository) | | Generated dataset | [`DAVIAN-Robotics/DiscoDemo-Stage2_GenData-FMB_Round-alpha0`](https://huggingface.co/datasets/DAVIAN-Robotics/DiscoDemo-Stage2_GenData-FMB_Round-alpha0) | | Imitation policy (π0.5) | [`DAVIAN-Robotics/DiscoDemo-Stage3_SFT-FMB_Round-alpha0`](https://huggingface.co/DAVIAN-Robotics/DiscoDemo-Stage3_SFT-FMB_Round-alpha0) | ## Files | File | Content | |---|---| | `actor.pt` | Actor network weights (`network_state_dict`). Optimizer state is not included. | | `normalizer.pt` | Observation bounds used to normalize the policy input. | | `config.yaml` | Full training configuration (environment, curriculum, agent). | Only what is needed to roll the policy out is released; the critic and other training-only state are omitted. Environment steps at this checkpoint: 50.0 M. ## Usage Loading requires the DiscoDemo training code: https://github.com/DAVIAN-Robotics/DiscoDemo. ## Citation ```bibtex @article{park2026discodemo, title = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning}, author = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul}, journal = {arXiv preprint}, year = {2026} } ```