hzeng412 commited on
Commit
7827733
·
1 Parent(s): 9c19932

docs: fix Flash-WAM model card usage

Browse files
Files changed (1) hide show
  1. README.md +193 -26
README.md CHANGED
@@ -1,50 +1,217 @@
1
  ---
2
  license: apache-2.0
3
- library_name: diffusers
 
 
 
 
 
4
  tags:
5
  - robotics
 
 
6
  - world-model
7
  - diffusion
8
  - step-distillation
9
  - lingbot-va
10
- pipeline_tag: robotics
 
11
  ---
12
 
13
- # Flash-WAM — RoboTwin (distilled)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
- Single-step distilled checkpoint for **Flash-WAM: Modality-Aware Distillation for World Action Models**, applied to LingBot-VA and evaluated on RoboTwin 2.0. Flash-WAM distills each modality with a consistency function matched to its noise regime (linear-gradient-scaling for the action stream, variance-preserving for the video stream), compressing inference to a single step per modality for up to a **23× speedup** while preserving teacher-level task success.
 
 
16
 
17
- This repository contains the **complete model** (distilled transformer + encoders):
18
 
19
- | Component | Description |
20
- | :--- | :--- |
21
- | `transformer/` | Distilled Flash-WAM student |
22
- | `vae/` | VAE (from the LingBot-VA teacher) |
23
- | `text_encoder/` | UMT5-XXL text encoder (from the teacher) |
24
- | `tokenizer/` | T5 tokenizer |
25
 
26
- ## Links
 
 
27
 
28
- - 📄 Paper: https://arxiv.org/abs/2606.05254
29
- - 🌐 Project page: https://flashwam.github.io
30
- - 💻 Code: https://github.com/NU-World-Model-Embodied-AI/Flash-WAM
 
 
 
31
 
32
- ## Usage
 
 
33
 
34
- For environment setup and evaluation, follow the [Flash-WAM repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) and [LingBot-VA](https://github.com/Robbyant/lingbot-va). Point the inference server at this checkpoint directory.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
 
36
  ## Citation
37
 
38
  ```bibtex
39
  @misc{akbari2026flashwammodalityawaredistillationworld,
40
- title={Flash-WAM: Modality-Aware Distillation for World Action Models},
41
- author={Arman Akbari and Ci Zhang and Arash Akbari and Lin Zhao and Yixiao Chen and Weiwei Chen and Xuan Zhang and Geng Yuan and Yanzhi Wang},
42
- year={2026},
43
- eprint={2606.05254},
44
- archivePrefix={arXiv},
45
- primaryClass={cs.LG},
46
- url={https://arxiv.org/abs/2606.05254},
47
  }
48
  ```
49
-
50
- License: Apache-2.0.
 
1
  ---
2
  license: apache-2.0
3
+ base_model: robbyant/lingbot-va-posttrain-robotwin
4
+ datasets:
5
+ - robbyant/robotwin-clean-and-aug-lerobot
6
+ pipeline_tag: robotics
7
+ language:
8
+ - en
9
  tags:
10
  - robotics
11
+ - embodied-ai
12
+ - world-action-model
13
  - world-model
14
  - diffusion
15
  - step-distillation
16
  - lingbot-va
17
+ - robotwin
18
+ - arxiv:2606.05254
19
  ---
20
 
21
+ # Flash-WAM RoboTwin: Distilled World-Action Model
22
+
23
+ [Project page](https://flashwam.github.io/) ·
24
+ [Paper](https://arxiv.org/abs/2606.05254) ·
25
+ [Code](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) ·
26
+ [LingBot-VA](https://github.com/Robbyant/lingbot-va)
27
+
28
+ This repository contains the complete RoboTwin checkpoint for **Flash-WAM:
29
+ Modality-Aware Distillation for World Action Models**. Flash-WAM distills the
30
+ joint video and action streams of LingBot-VA with consistency functions matched
31
+ to their different noise regimes.
32
+
33
+ The released student supports one-step video and one-step action generation.
34
+ Under the paper's RoboTwin 2.0 setup on a single NVIDIA L40S, this reduces
35
+ per-chunk latency from **8.1 seconds to 348 milliseconds**, a **23.3× speedup**.
36
+
37
+ > **Important:** this is a custom joint video-action robotics model, not a
38
+ > generic text-to-image or video `DiffusionPipeline`. Do not use
39
+ > `DiffusionPipeline.from_pretrained(...)`. Install the Flash-WAM/LingBot-VA
40
+ > code and use their RoboTwin server/client evaluation path.
41
+
42
+ ## Model details
43
+
44
+ | Field | Value |
45
+ | --- | --- |
46
+ | Base model | [LingBot-VA RoboTwin post-training checkpoint](https://huggingface.co/robbyant/lingbot-va-posttrain-robotwin) |
47
+ | Task | Joint future-video and robot-action prediction |
48
+ | Benchmark | RoboTwin 2.0 |
49
+ | Released student | 1 video step / 1 action step |
50
+ | Action dimension | 30 in the released transformer config |
51
+ | Reported latency hardware | 1 × NVIDIA L40S |
52
+ | Checkpoint license | Apache-2.0 |
53
+
54
+ ## Repository contents
55
+
56
+ | Directory | Description |
57
+ | --- | --- |
58
+ | `transformer/` | Distilled Flash-WAM student, approximately 10 GB |
59
+ | `vae/` | VAE inherited from the LingBot-VA teacher, approximately 2.8 GB |
60
+ | `text_encoder/` | UMT5-XXL text encoder, approximately 11.3 GB |
61
+ | `tokenizer/` | T5 tokenizer files |
62
+
63
+ The full snapshot is approximately 24 GB.
64
+
65
+ ## Download
66
+
67
+ ```bash
68
+ pip install -U huggingface_hub
69
+
70
+ hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \
71
+ --local-dir ./FlashWAM-RoboTwin
72
+ ```
73
+
74
+ To inspect configs without downloading the weights:
75
+
76
+ ```bash
77
+ hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \
78
+ README.md \
79
+ transformer/config.json vae/config.json text_encoder/config.json \
80
+ tokenizer/tokenizer_config.json \
81
+ --local-dir ./FlashWAM-RoboTwin-config
82
+ ```
83
+
84
+ ## Environment and evaluation
85
+
86
+ Flash-WAM uses the LingBot-VA environment and the same RoboTwin server/client
87
+ evaluation pipeline:
88
+
89
+ 1. Follow the [LingBot-VA installation and RoboTwin evaluation
90
+ instructions](https://github.com/Robbyant/lingbot-va).
91
+ 2. Clone the [Flash-WAM repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM)
92
+ so the custom `wan_va` model implementation is available.
93
+ 3. Download this snapshot and set the model path in the LingBot-VA/Flash-WAM
94
+ evaluation configuration to the local snapshot directory.
95
+ 4. For training or distillation, use the released commands in the Flash-WAM
96
+ repository; this checkpoint is the already-distilled student.
97
 
98
+ The public Flash-WAM repository does not yet include the real-world Unitree G1
99
+ deployment setup. Do not infer a supported real-robot deployment command from
100
+ the checkpoint layout alone.
101
 
102
+ ## Optional component-loading check
103
 
104
+ After installing the LingBot-VA environment and making the Flash-WAM repository
105
+ available on `PYTHONPATH`, the released helper functions can load the individual
106
+ components:
 
 
 
107
 
108
+ ```python
109
+ from pathlib import Path
110
+ import torch
111
 
112
+ from wan_va.modules.utils import (
113
+ load_text_encoder,
114
+ load_tokenizer,
115
+ load_transformer,
116
+ load_vae,
117
+ )
118
 
119
+ root = Path("/path/to/FlashWAM-RoboTwin")
120
+ device = "cuda"
121
+ dtype = torch.bfloat16
122
 
123
+ tokenizer = load_tokenizer(root / "tokenizer")
124
+ text_encoder = load_text_encoder(root / "text_encoder", dtype, device)
125
+ vae = load_vae(root / "vae", dtype, device)
126
+ transformer = load_transformer(root / "transformer", dtype, device)
127
+ ```
128
+
129
+ This verifies component compatibility; it is not a complete policy rollout.
130
+ Use the upstream server/client evaluation path for observations, action
131
+ normalization, temporal caching, and environment interaction.
132
+
133
+ ## Reported results
134
+
135
+ ### RoboTwin 2.0
136
+
137
+ | Method | Video steps | Action steps | Average success | Speedup |
138
+ | --- | ---: | ---: | ---: | ---: |
139
+ | LingBot-VA teacher | 25 | 50 | 91.25% | 1.0× |
140
+ | Naive joint LCM | 1 | 2 | 23.97% | — |
141
+ | **Flash-WAM** | **1** | **2** | **85.54%** | **19.0×** |
142
+ | Naive joint LCM | 1 | 1 | 36.32% | — |
143
+ | **Flash-WAM** | **1** | **1** | **81.41%** | **23.3×** |
144
+
145
+ ### LIBERO
146
+
147
+ | Method | Video steps | Action steps | Average success | Speedup |
148
+ | --- | ---: | ---: | ---: | ---: |
149
+ | LingBot-VA teacher | 20 | 50 | 98.6% | 1.0× |
150
+ | **Flash-WAM** | **1** | **2** | **95.7%** | **13.7×** |
151
+ | **Flash-WAM** | **1** | **1** | **95.1%** | **16.3×** |
152
+
153
+ ### Real-world Unitree G1
154
+
155
+ Three manipulation tasks were evaluated with 10 rollouts per task:
156
+
157
+ | Method | Video/action steps | T1 | T2 | T3 | Average |
158
+ | --- | ---: | ---: | ---: | ---: | ---: |
159
+ | LingBot-VA | 3 / 10 | 50% | 70% | 80% | 66.7% |
160
+ | **Flash-WAM** | **1 / 2** | **50%** | **60%** | **70%** | **60.0%** |
161
+ | **Flash-WAM** | **1 / 1** | **40%** | **50%** | **60%** | **50.0%** |
162
+
163
+ The fastest 1-video/1-action-step configuration and the 60% real-world result
164
+ are **not the same configuration**. Report step budgets together with every
165
+ success-rate or latency claim.
166
+
167
+ ## Intended use
168
+
169
+ This checkpoint is intended for:
170
+
171
+ - research on step distillation for joint video-action models;
172
+ - reproducing the reported RoboTwin results;
173
+ - comparing modality-aware and naive joint consistency objectives;
174
+ - studying latency/task-success trade-offs in world-action models.
175
+
176
+ It is not a drop-in controller for an arbitrary robot or task. Deployment on
177
+ physical robots requires task-specific observation processing, action
178
+ normalization, safety constraints, control integration, and validation.
179
+
180
+ ## Limitations and safety
181
+
182
+ - Results are specific to LingBot-VA, the released RoboTwin checkpoint, and the
183
+ paper's evaluation settings.
184
+ - Latency depends on GPU, software stack, precision, resolution, horizon, and
185
+ server/client overhead; 348 ms is not a universal runtime guarantee.
186
+ - The real-world evaluation covers three tasks and 30 rollouts per method.
187
+ - One-step generation still reduces task success relative to the teacher.
188
+ - Generated actions may be unsafe or incorrect. Use independent safeguards,
189
+ workspace limits, emergency stops, and supervised testing before any
190
+ physical deployment.
191
+ - The real-world G1 deployment setup is not included in the public code release.
192
+
193
+ ## Licenses
194
+
195
+ The **checkpoint in this Hugging Face repository** is released under
196
+ Apache-2.0. It includes components derived from LingBot-VA, whose released
197
+ model and bundled upstream components are also Apache-2.0.
198
+
199
+ The separate [Flash-WAM GitHub repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM)
200
+ uses different terms for different parts: the authors' distillation code,
201
+ documentation, and demo videos are CC BY-NC 4.0, while the bundled `wan_va/`
202
+ components remain Apache-2.0. Downloading this checkpoint does not replace the
203
+ license notices of the code or other assets used with it.
204
 
205
  ## Citation
206
 
207
  ```bibtex
208
  @misc{akbari2026flashwammodalityawaredistillationworld,
209
+ title = {Flash-WAM: Modality-Aware Distillation for World Action Models},
210
+ author = {Arman Akbari and Ci Zhang and Arash Akbari and Lin Zhao and Yixiao Chen and Weiwei Chen and Xuan Zhang and Geng Yuan and Yanzhi Wang},
211
+ year = {2026},
212
+ eprint = {2606.05254},
213
+ archivePrefix = {arXiv},
214
+ primaryClass = {cs.LG},
215
+ url = {https://arxiv.org/abs/2606.05254}
216
  }
217
  ```