Add DMM checkpoints and model card
Browse files- DMM-08M.pt +3 -0
- DMM-3M.pt +3 -0
- DMM-MICPO-08M.pt +3 -0
- DMM-MICPO-3M.pt +3 -0
- README.md +27 -0
DMM-08M.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6d6017bd8d274ff62a1987ab943763952d97f2ca64f3a4b829dcb3e2853910e5
|
| 3 |
+
size 3076725
|
DMM-3M.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:20ac0bb4b88fd4904e49d15d8e88ab1d7b768c37fba43806ab9bc643c5aa5916
|
| 3 |
+
size 13028093
|
DMM-MICPO-08M.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6fcd798143fef80f2efd936875e5648c5bd505cdaee498601adfa11a44f38d3a
|
| 3 |
+
size 3079541
|
DMM-MICPO-3M.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:69fc0dd5c05ce394148646d225e8740ec385958e443ce400a58499ee8c77d38d
|
| 3 |
+
size 13036441
|
README.md
CHANGED
|
@@ -1,3 +1,30 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- pytorch
|
| 5 |
+
- multi-agent-path-finding
|
| 6 |
+
- mapf
|
| 7 |
+
- decentralized
|
| 8 |
---
|
| 9 |
+
|
| 10 |
+
# DMM: Decentralized Master-Mind
|
| 11 |
+
|
| 12 |
+
DMM is a decentralized multi-agent pathfinding policy that refines agents' action
|
| 13 |
+
intents over several local communication rounds before committing to actions.
|
| 14 |
+
The models are pretrained on expert solutions with imitation learning and
|
| 15 |
+
optionally fine-tuned with MICPO, a critic-free reinforcement learning method.
|
| 16 |
+
|
| 17 |
+
[Paper](https://arxiv.org/abs/2609.32019) · [Code](https://github.com/CognitiveAISystems/DMM)
|
| 18 |
+
|
| 19 |
+
| Checkpoint | Imitation pretraining iterations | MICPO optimizer updates |
|
| 20 |
+
| --- | ---: | ---: |
|
| 21 |
+
| [DMM-08M.pt](DMM-08M.pt) | 1,000,000 | — |
|
| 22 |
+
| [DMM-3M.pt](DMM-3M.pt) | 1,000,000 | — |
|
| 23 |
+
| [DMM-MICPO-08M.pt](DMM-MICPO-08M.pt) | 1,000,000 | 96,000 |
|
| 24 |
+
| [DMM-MICPO-3M.pt](DMM-MICPO-3M.pt) | 1,000,000 | 96,000 |
|
| 25 |
+
|
| 26 |
+
MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates).
|
| 27 |
+
Both model sizes use four communication rounds. Training details are reported in
|
| 28 |
+
[the paper](https://arxiv.org/html/2609.32019v1#S5.SS1).
|
| 29 |
+
|
| 30 |
+
See the [GitHub repository](https://github.com/CognitiveAISystems/DMM) for code, usage instructions, training, and evaluation.
|