tviskaron commited on
Commit
b8cc825
·
1 Parent(s): 642781c

Add DMM checkpoints and model card

Browse files
Files changed (5) hide show
  1. DMM-08M.pt +3 -0
  2. DMM-3M.pt +3 -0
  3. DMM-MICPO-08M.pt +3 -0
  4. DMM-MICPO-3M.pt +3 -0
  5. README.md +27 -0
DMM-08M.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6d6017bd8d274ff62a1987ab943763952d97f2ca64f3a4b829dcb3e2853910e5
3
+ size 3076725
DMM-3M.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:20ac0bb4b88fd4904e49d15d8e88ab1d7b768c37fba43806ab9bc643c5aa5916
3
+ size 13028093
DMM-MICPO-08M.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6fcd798143fef80f2efd936875e5648c5bd505cdaee498601adfa11a44f38d3a
3
+ size 3079541
DMM-MICPO-3M.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:69fc0dd5c05ce394148646d225e8740ec385958e443ce400a58499ee8c77d38d
3
+ size 13036441
README.md CHANGED
@@ -1,3 +1,30 @@
1
  ---
2
  license: mit
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ tags:
4
+ - pytorch
5
+ - multi-agent-path-finding
6
+ - mapf
7
+ - decentralized
8
  ---
9
+
10
+ # DMM: Decentralized Master-Mind
11
+
12
+ DMM is a decentralized multi-agent pathfinding policy that refines agents' action
13
+ intents over several local communication rounds before committing to actions.
14
+ The models are pretrained on expert solutions with imitation learning and
15
+ optionally fine-tuned with MICPO, a critic-free reinforcement learning method.
16
+
17
+ [Paper](https://arxiv.org/abs/2609.32019) · [Code](https://github.com/CognitiveAISystems/DMM)
18
+
19
+ | Checkpoint | Imitation pretraining iterations | MICPO optimizer updates |
20
+ | --- | ---: | ---: |
21
+ | [DMM-08M.pt](DMM-08M.pt) | 1,000,000 | — |
22
+ | [DMM-3M.pt](DMM-3M.pt) | 1,000,000 | — |
23
+ | [DMM-MICPO-08M.pt](DMM-MICPO-08M.pt) | 1,000,000 | 96,000 |
24
+ | [DMM-MICPO-3M.pt](DMM-MICPO-3M.pt) | 1,000,000 | 96,000 |
25
+
26
+ MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates).
27
+ Both model sizes use four communication rounds. Training details are reported in
28
+ [the paper](https://arxiv.org/html/2609.32019v1#S5.SS1).
29
+
30
+ See the [GitHub repository](https://github.com/CognitiveAISystems/DMM) for code, usage instructions, training, and evaluation.