File size: 7,974 Bytes
d26c24e
 
0c05b2f
7827733
 
 
 
 
 
d26c24e
 
7827733
 
d26c24e
 
 
 
7827733
 
d26c24e
 
7827733
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d26c24e
7827733
 
 
d26c24e
7827733
d26c24e
7827733
 
 
d26c24e
7827733
 
 
d26c24e
7827733
 
 
 
 
 
d26c24e
7827733
 
 
d26c24e
7827733
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d26c24e
 
 
 
 
7827733
 
 
 
 
 
 
d26c24e
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
---
license: apache-2.0
library_name: custom
base_model: robbyant/lingbot-va-posttrain-robotwin
datasets:
- robbyant/robotwin-clean-and-aug-lerobot
pipeline_tag: robotics
language:
- en
tags:
- robotics
- embodied-ai
- world-action-model
- world-model
- diffusion
- step-distillation
- lingbot-va
- robotwin
- arxiv:2606.05254
---

# Flash-WAM RoboTwin: Distilled World-Action Model

[Project page](https://flashwam.github.io/) ·
[Paper](https://arxiv.org/abs/2606.05254) ·
[Code](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) ·
[LingBot-VA](https://github.com/Robbyant/lingbot-va)

This repository contains the complete RoboTwin checkpoint for **Flash-WAM:
Modality-Aware Distillation for World Action Models**. Flash-WAM distills the
joint video and action streams of LingBot-VA with consistency functions matched
to their different noise regimes.

The released student supports one-step video and one-step action generation.
Under the paper's RoboTwin 2.0 setup on a single NVIDIA L40S, this reduces
per-chunk latency from **8.1 seconds to 348 milliseconds**, a **23.3× speedup**.

> **Important:** this is a custom joint video-action robotics model, not a
> generic text-to-image or video `DiffusionPipeline`. Do not use
> `DiffusionPipeline.from_pretrained(...)`. Install the Flash-WAM/LingBot-VA
> code and use their RoboTwin server/client evaluation path.

## Model details

| Field | Value |
| --- | --- |
| Base model | [LingBot-VA RoboTwin post-training checkpoint](https://huggingface.co/robbyant/lingbot-va-posttrain-robotwin) |
| Task | Joint future-video and robot-action prediction |
| Benchmark | RoboTwin 2.0 |
| Released student | 1 video step / 1 action step |
| Action dimension | 30 in the released transformer config |
| Reported latency hardware | 1 × NVIDIA L40S |
| Checkpoint license | Apache-2.0 |

## Repository contents

| Directory | Description |
| --- | --- |
| `transformer/` | Distilled Flash-WAM student, approximately 10 GB |
| `vae/` | VAE inherited from the LingBot-VA teacher, approximately 2.8 GB |
| `text_encoder/` | UMT5-XXL text encoder, approximately 11.3 GB |
| `tokenizer/` | T5 tokenizer files |

The full snapshot is approximately 24 GB.

## Download

```bash
pip install -U huggingface_hub

hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \
  --local-dir ./FlashWAM-RoboTwin
```

To inspect configs without downloading the weights:

```bash
hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \
  README.md \
  transformer/config.json vae/config.json text_encoder/config.json \
  tokenizer/tokenizer_config.json \
  --local-dir ./FlashWAM-RoboTwin-config
```

## Environment and evaluation

Flash-WAM uses the LingBot-VA environment and the same RoboTwin server/client
evaluation pipeline:

1. Follow the [LingBot-VA installation and RoboTwin evaluation
   instructions](https://github.com/Robbyant/lingbot-va).
2. Clone the [Flash-WAM repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM)
   so the custom `wan_va` model implementation is available.
3. Download this snapshot and set the model path in the LingBot-VA/Flash-WAM
   evaluation configuration to the local snapshot directory.
4. For training or distillation, use the released commands in the Flash-WAM
   repository; this checkpoint is the already-distilled student.

The public Flash-WAM repository does not yet include the real-world Unitree G1
deployment setup. Do not infer a supported real-robot deployment command from
the checkpoint layout alone.

## Optional component-loading check

After installing the LingBot-VA environment and making the Flash-WAM repository
available on `PYTHONPATH`, the released helper functions can load the individual
components:

```python
from pathlib import Path
import torch

from wan_va.modules.utils import (
    load_text_encoder,
    load_tokenizer,
    load_transformer,
    load_vae,
)

root = Path("/path/to/FlashWAM-RoboTwin")
device = "cuda"
dtype = torch.bfloat16

tokenizer = load_tokenizer(root / "tokenizer")
text_encoder = load_text_encoder(root / "text_encoder", dtype, device)
vae = load_vae(root / "vae", dtype, device)
transformer = load_transformer(root / "transformer", dtype, device)
```

This verifies component compatibility; it is not a complete policy rollout.
Use the upstream server/client evaluation path for observations, action
normalization, temporal caching, and environment interaction.

## Reported results

### RoboTwin 2.0

| Method | Video steps | Action steps | Average success | Speedup |
| --- | ---: | ---: | ---: | ---: |
| LingBot-VA teacher | 25 | 50 | 91.25% | 1.0× |
| Naive joint LCM | 1 | 2 | 23.97% | — |
| **Flash-WAM** | **1** | **2** | **85.54%** | **19.0×** |
| Naive joint LCM | 1 | 1 | 36.32% | — |
| **Flash-WAM** | **1** | **1** | **81.41%** | **23.3×** |

### LIBERO

| Method | Video steps | Action steps | Average success | Speedup |
| --- | ---: | ---: | ---: | ---: |
| LingBot-VA teacher | 20 | 50 | 98.6% | 1.0× |
| **Flash-WAM** | **1** | **2** | **95.7%** | **13.7×** |
| **Flash-WAM** | **1** | **1** | **95.1%** | **16.3×** |

### Real-world Unitree G1

Three manipulation tasks were evaluated with 10 rollouts per task:

| Method | Video/action steps | T1 | T2 | T3 | Average |
| --- | ---: | ---: | ---: | ---: | ---: |
| LingBot-VA | 3 / 10 | 50% | 70% | 80% | 66.7% |
| **Flash-WAM** | **1 / 2** | **50%** | **60%** | **70%** | **60.0%** |
| **Flash-WAM** | **1 / 1** | **40%** | **50%** | **60%** | **50.0%** |

The fastest 1-video/1-action-step configuration and the 60% real-world result
are **not the same configuration**. Report step budgets together with every
success-rate or latency claim.

## Intended use

This checkpoint is intended for:

- research on step distillation for joint video-action models;
- reproducing the reported RoboTwin results;
- comparing modality-aware and naive joint consistency objectives;
- studying latency/task-success trade-offs in world-action models.

It is not a drop-in controller for an arbitrary robot or task. Deployment on
physical robots requires task-specific observation processing, action
normalization, safety constraints, control integration, and validation.

## Limitations and safety

- Results are specific to LingBot-VA, the released RoboTwin checkpoint, and the
  paper's evaluation settings.
- Latency depends on GPU, software stack, precision, resolution, horizon, and
  server/client overhead; 348 ms is not a universal runtime guarantee.
- The real-world evaluation covers three tasks and 30 rollouts per method.
- One-step generation still reduces task success relative to the teacher.
- Generated actions may be unsafe or incorrect. Use independent safeguards,
  workspace limits, emergency stops, and supervised testing before any
  physical deployment.
- The real-world G1 deployment setup is not included in the public code release.

## Licenses

The **checkpoint in this Hugging Face repository** is released under
Apache-2.0. It includes components derived from LingBot-VA, whose released
model and bundled upstream components are also Apache-2.0.

The separate [Flash-WAM GitHub repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM)
uses different terms for different parts: the authors' distillation code,
documentation, and demo videos are CC BY-NC 4.0, while the bundled `wan_va/`
components remain Apache-2.0. Downloading this checkpoint does not replace the
license notices of the code or other assets used with it.

## Citation

```bibtex
@misc{akbari2026flashwammodalityawaredistillationworld,
  title         = {Flash-WAM: Modality-Aware Distillation for World Action Models},
  author        = {Arman Akbari and Ci Zhang and Arash Akbari and Lin Zhao and Yixiao Chen and Weiwei Chen and Xuan Zhang and Geng Yuan and Yanzhi Wang},
  year          = {2026},
  eprint        = {2606.05254},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2606.05254}
}
```