File size: 2,046 Bytes
a38b99a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a1179c1
a38b99a
 
 
 
 
5d442e9
a38b99a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
---
license: apache-2.0
tags:
- robotics
- reinforcement-learning
- discodemo
- franka
---

# DiscoDemo-Stage1_RL-FMB_Round-alpha0

The **data generator** of [DiscoDemo](https://davian-robotics.github.io/DiscoDemo/) for the **FMB-Round** task (insert a cylindrical peg into its hole on the FMB board): a state-based
reinforcement-learning policy trained in simulation — the P-RFCL baseline generator (α = 0: the same reverse-curriculum RL without the diversity reward).
Rolling it out generates demonstrations such as those in the Stage2 dataset below.

| Stage | Repository |
|---|---|
| Generator (RL policy) | [`DAVIAN-Robotics/DiscoDemo-Stage1_RL-FMB_Round-alpha0`](https://huggingface.co/DAVIAN-Robotics/DiscoDemo-Stage1_RL-FMB_Round-alpha0) (this repository) |
| Generated dataset | [`DAVIAN-Robotics/DiscoDemo-Stage2_GenData-FMB_Round-alpha0`](https://huggingface.co/datasets/DAVIAN-Robotics/DiscoDemo-Stage2_GenData-FMB_Round-alpha0) |
| Imitation policy (π0.5) | [`DAVIAN-Robotics/DiscoDemo-Stage3_SFT-FMB_Round-alpha0`](https://huggingface.co/DAVIAN-Robotics/DiscoDemo-Stage3_SFT-FMB_Round-alpha0) |

## Files

| File | Content |
|---|---|
| `actor.pt` | Actor network weights (`network_state_dict`). Optimizer state is not included. |
| `normalizer.pt` | Observation bounds used to normalize the policy input. |
| `config.yaml` | Full training configuration (environment, curriculum, agent). |

Only what is needed to roll the policy out is released; the critic and other training-only state are omitted.


Environment steps at this checkpoint: 50.0 M.

## Usage

Loading requires the DiscoDemo training code: https://github.com/DAVIAN-Robotics/DiscoDemo.

## Citation

```bibtex
@article{park2026discodemo,
  title   = {DiscoDemo: Discovering Efficient and Diverse Robot Demonstrations for Imitation Learning},
  author  = {Park, Minho and Kim, Kinam and Kim, Donghu and Lee, Byungkun and Hwang, Dongyoon and Shin, Yongjae and Hyung, Junha and Lee, Hojoon and Choo, Jaegul},
  journal = {arXiv preprint},
  year    = {2026}
}
```