zhaoyian01 commited on
Commit
52abd38
·
verified ·
1 Parent(s): 51f6144

Add files using upload-large-folder tool

Browse files
MiniWorld_0_5b_droid.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e4b118befe88cee7338400c5510fdd497212b9b1988034290030b3ed351ced32
3
+ size 2226856306
MiniWorld_0_5b_re10k.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:184c135964a212907de77a4a00e6a9e7ca6bae0eb74377ee8cd68dfca0e3fc9d
3
+ size 2122945383
MiniWorld_1b_droid.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0737ad9c88934b7e558fcd4a047624aed4e3a45bb2a37f532f2c699470ea07c2
3
+ size 3854030562
MiniWorld_1b_re10k.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:10cd8b641e59584d59592e94ef7e71c53e83c4bb807d8e7df982624f3c6fe04a
3
+ size 3668296921
README.md CHANGED
@@ -1,3 +1,216 @@
1
  ---
 
 
 
 
 
 
 
 
 
 
 
 
2
  license: apache-2.0
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ library_name: pytorch
3
+ tags:
4
+ - world-model
5
+ - video-generation
6
+ - streaming-generation
7
+ - robotics
8
+ - camera-control
9
+ - diffusion
10
+ - rectified-flow
11
+ - video-dit
12
+ - droid
13
+ - realestate10k
14
  license: apache-2.0
15
  ---
16
+
17
+ # MiniWorld
18
+
19
+ **MiniWorld: Democratizing the Training of Video World Models from Scratch**
20
+
21
+ <a href="https://zhao-yian.github.io/MiniWorld/"><img src="https://img.shields.io/badge/Project-Page-1f6feb?style=for-the-badge&logo=googlechrome&logoColor=white" alt="Project Page"></a>
22
+ <img src="https://img.shields.io/badge/arXiv-Coming%20Soon-b31b1b?style=for-the-badge&logo=arxiv&logoColor=white" alt="arXiv">
23
+ <a href="https://github.com/zhao-yian/MiniWorld"><img src="https://img.shields.io/badge/GitHub-Code-181717?style=for-the-badge&logo=github&logoColor=white" alt="GitHub"></a>
24
+
25
+ MiniWorld is a minimal and reproducible framework for training streaming video
26
+ world models from scratch. Instead of adapting a pretrained bidirectional video
27
+ generator, it directly learns causal next-state prediction with a block-causal
28
+ Video Diffusion Transformer and Rectified Flow.
29
+
30
+ The same architecture supports two control modalities:
31
+
32
+ - **DROID:** low-level robot actions for embodied world modeling.
33
+ - **RealEstate10K:** camera poses for controllable scene prediction.
34
+
35
+ This Hugging Face repository hosts the MiniWorld model checkpoints. Code,
36
+ training scripts, and evaluation utilities live in the GitHub repository.
37
+
38
+ ## Model Summary
39
+
40
+ MiniWorld uses a block-causal Video Diffusion Transformer trained with Rectified
41
+ Flow in the latent space of the Wan2.2 VAE. During inference, MiniWorld performs
42
+ streaming generation with a rolling KV cache and pipelined asynchronous
43
+ denoising, enabling long-horizon generation under bounded online computation.
44
+
45
+ Key components:
46
+
47
+ - **Block-causal Video DiT** with bidirectional attention inside each chunk and
48
+ causal attention across chunks.
49
+ - **Unified conditioning** for robot actions and camera poses through AdaLN-LoRA
50
+ modulation.
51
+ - **Chunk-oriented Probability Propagation (CoPP)** for stable non-decreasing
52
+ diffusion schedules.
53
+ - **Continued long-context training** from short clips to 253-frame sequences.
54
+ - **Structured rolling KV cache** with a persistent sink and FIFO history.
55
+ - **Pipelined asynchronous denoising** for a quality-throughput trade-off at
56
+ inference time.
57
+
58
+ The complete model can be trained in several days on a single 8-GPU server.
59
+
60
+ ## Released Checkpoints
61
+
62
+ Sampling requires matching the checkpoint with the corresponding dataset and
63
+ model scale.
64
+
65
+ | Dataset | Model | Status | Checkpoint |
66
+ | --- | --- | --- | --- |
67
+ | DROID | MiniWorld-0.5B | Available | [MiniWorld_0_5b_droid.pt](resolve/main/MiniWorld_0_5b_droid.pt) |
68
+ | DROID | MiniWorld-1B | Available | [MiniWorld_1b_droid.pt](resolve/main/MiniWorld_1b_droid.pt) |
69
+ | DROID | MiniWorld-3B | Coming soon | -- |
70
+ | RealEstate10K | MiniWorld-0.5B | Available | [MiniWorld_0_5b_re10k.pt](resolve/main/MiniWorld_0_5b_re10k.pt) |
71
+ | RealEstate10K | MiniWorld-1B | Available | [MiniWorld_1b_re10k.pt](resolve/main/MiniWorld_1b_re10k.pt) |
72
+ | RealEstate10K | MiniWorld-3B | Coming soon | -- |
73
+
74
+ Download a single checkpoint with:
75
+
76
+ ```bash
77
+ hf download zhaoyian01/MiniWorld \
78
+ --include "MiniWorld_1b_droid.pt" \
79
+ --local-dir checkpoints/miniworld
80
+ ```
81
+
82
+ ## Model Configurations
83
+
84
+ `MODEL` is the identifier expected by the training and sampling scripts in the
85
+ GitHub repository.
86
+
87
+ | Model | `MODEL` | Depth | Width | Heads | Parameters |
88
+ | --- | --- | ---: | ---: | ---: | ---: |
89
+ | MiniWorld-B | `B` | 12 | 768 | 12 | 0.12B |
90
+ | MiniWorld-L | `L` | 24 | 1024 | 16 | 0.39B |
91
+ | MiniWorld-0.5B | `0.5B` | 28 | 1152 | 16 | 0.55B |
92
+ | MiniWorld-1B | `1B` | 28 | 1536 | 12 | 1B |
93
+ | MiniWorld-3B | `3B` | 32 | 2560 | 20 | 3B |
94
+
95
+ ## Intended Use
96
+
97
+ MiniWorld is intended for research on streaming video world models, including:
98
+
99
+ - action-conditioned robot world modeling,
100
+ - camera-pose-conditioned scene prediction,
101
+ - long-horizon autoregressive video generation,
102
+ - temporal memory and KV-cache mechanisms,
103
+ - train-test alignment for streaming diffusion models.
104
+
105
+ MiniWorld is a research baseline and is not intended as a general-purpose
106
+ text-to-video model.
107
+
108
+ ## Requirements
109
+
110
+ Inference requires the MiniWorld codebase and the pretrained Wan2.2 VAE:
111
+
112
+ - Linux with an NVIDIA CUDA GPU
113
+ - Python 3.11
114
+ - CUDA-compatible PyTorch 2.x
115
+ - FlashAttention
116
+ - Wan2.2 VAE checkpoint from `Wan-AI/Wan2.2-TI2V-5B`
117
+
118
+ Download the VAE:
119
+
120
+ ```bash
121
+ hf download Wan-AI/Wan2.2-TI2V-5B \
122
+ --include "Wan2.2_VAE.pth" \
123
+ --local-dir checkpoints/wan2.2
124
+ ```
125
+
126
+ ## Usage
127
+
128
+ Clone the [MiniWorld codebase](https://github.com/zhao-yian/MiniWorld), install
129
+ its requirements, then download the desired checkpoint. All commands are run
130
+ from the repository root.
131
+
132
+ The default sampler uses one observed frame as initial context, eight in-flight
133
+ chunks and a 24-chunk rolling KV cache (a 64-frame active attention window), one
134
+ persistent sink frame, 100 denoising steps with classifier-free guidance at
135
+ scale 2.0, and a 64-latent-frame rollout corresponding to 253 RGB frames.
136
+ Generated videos are saved to `${SAMPLE_DIR}/pred/`.
137
+
138
+ ### DROID action-conditioned generation
139
+
140
+ ```bash
141
+ DATA_ROOT=/path/to/droid_lerobot \
142
+ CKPT=/path/to/MiniWorld_1b_droid.pt \
143
+ VAE_CKPT=checkpoints/wan2.2/Wan2.2_VAE.pth \
144
+ MODEL=1B \
145
+ bash scripts/sample_droid.sh
146
+ ```
147
+
148
+ ### RealEstate10K camera-conditioned generation
149
+
150
+ ```bash
151
+ DATA_ROOT=/path/to/re10k/videos \
152
+ POSE_DIR=/path/to/re10k/poses \
153
+ CKPT=/path/to/MiniWorld_1b_re10k.pt \
154
+ VAE_CKPT=checkpoints/wan2.2/Wan2.2_VAE.pth \
155
+ MODEL=1B \
156
+ bash scripts/sample_re10k.sh
157
+ ```
158
+
159
+ ### Common inference controls
160
+
161
+ ```bash
162
+ GPU=0 \
163
+ TOTAL_LEN=96 \
164
+ CFG_SCALE=2.0 \
165
+ SAMPLE_NUM_VIDEOS=10 \
166
+ STREAM_INFLIGHT_CHUNKS=8 \
167
+ STREAM_MAX_CACHE_CHUNKS=24 \
168
+ STREAM_SINK_SIZE=1 \
169
+ bash scripts/sample_droid.sh
170
+ ```
171
+
172
+ `TOTAL_LEN` sets the rollout length in latent frames and can exceed the trained
173
+ window, since streaming keeps the attention span bounded; `TOTAL_LEN=96` yields
174
+ 381 RGB frames from a 64-frame checkpoint. MiniWorld is a streaming model and
175
+ does not assume a fixed generation horizon.
176
+
177
+ ### Custom camera trajectories
178
+
179
+ A RealEstate10K checkpoint can also animate a single image along a procedural
180
+ camera trajectory, without any dataset on disk:
181
+
182
+ ```bash
183
+ PYTHONPATH=. python -m miniworld.sample \
184
+ --dataset re10k \
185
+ --init_image /path/to/first_frame.png \
186
+ --custom_camera_trajectory orbit_right \
187
+ --checkpoint /path/to/MiniWorld_1b_re10k.pt \
188
+ --vae_checkpoint checkpoints/wan2.2/Wan2.2_VAE.pth \
189
+ --sample_dir samples/re10k_orbit_right \
190
+ --wm_model 1B \
191
+ --total_len 64 \
192
+ --sample_num_videos 1 \
193
+ --trajectory_magnitude 3.0
194
+ ```
195
+
196
+ These checkpoints are trained on raw (unnormalized) translations, so
197
+ `--trajectory_magnitude` is worth tuning: `1.0` is almost static, `3.0` is a
198
+ good default at `--total_len 64`, and values above `5.0` degrade the second half
199
+ of the rollout. Scale it with the rollout length to keep the same apparent
200
+ speed. See the GitHub README for the full list of trajectories.
201
+
202
+
203
+
204
+ ## Limitations
205
+
206
+ MiniWorld is a research model trained and evaluated at modest resolution and on
207
+ limited domains. It may exhibit long-horizon drift, geometric errors, temporal
208
+ inconsistencies, and failures under out-of-distribution actions, poses, scenes,
209
+ or camera motions. It should not be used for safety-critical simulation or as a
210
+ faithful physical simulator.
211
+
212
+ ## License
213
+
214
+ These checkpoints are released under the Apache 2.0 license. Please also follow
215
+ the licenses and usage terms of the underlying datasets (DROID, RealEstate10K)
216
+ and of the Wan2.2 VAE.