File size: 7,974 Bytes
d26c24e 0c05b2f 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e 7827733 d26c24e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 | ---
license: apache-2.0
library_name: custom
base_model: robbyant/lingbot-va-posttrain-robotwin
datasets:
- robbyant/robotwin-clean-and-aug-lerobot
pipeline_tag: robotics
language:
- en
tags:
- robotics
- embodied-ai
- world-action-model
- world-model
- diffusion
- step-distillation
- lingbot-va
- robotwin
- arxiv:2606.05254
---
# Flash-WAM RoboTwin: Distilled World-Action Model
[Project page](https://flashwam.github.io/) ·
[Paper](https://arxiv.org/abs/2606.05254) ·
[Code](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) ·
[LingBot-VA](https://github.com/Robbyant/lingbot-va)
This repository contains the complete RoboTwin checkpoint for **Flash-WAM:
Modality-Aware Distillation for World Action Models**. Flash-WAM distills the
joint video and action streams of LingBot-VA with consistency functions matched
to their different noise regimes.
The released student supports one-step video and one-step action generation.
Under the paper's RoboTwin 2.0 setup on a single NVIDIA L40S, this reduces
per-chunk latency from **8.1 seconds to 348 milliseconds**, a **23.3× speedup**.
> **Important:** this is a custom joint video-action robotics model, not a
> generic text-to-image or video `DiffusionPipeline`. Do not use
> `DiffusionPipeline.from_pretrained(...)`. Install the Flash-WAM/LingBot-VA
> code and use their RoboTwin server/client evaluation path.
## Model details
| Field | Value |
| --- | --- |
| Base model | [LingBot-VA RoboTwin post-training checkpoint](https://huggingface.co/robbyant/lingbot-va-posttrain-robotwin) |
| Task | Joint future-video and robot-action prediction |
| Benchmark | RoboTwin 2.0 |
| Released student | 1 video step / 1 action step |
| Action dimension | 30 in the released transformer config |
| Reported latency hardware | 1 × NVIDIA L40S |
| Checkpoint license | Apache-2.0 |
## Repository contents
| Directory | Description |
| --- | --- |
| `transformer/` | Distilled Flash-WAM student, approximately 10 GB |
| `vae/` | VAE inherited from the LingBot-VA teacher, approximately 2.8 GB |
| `text_encoder/` | UMT5-XXL text encoder, approximately 11.3 GB |
| `tokenizer/` | T5 tokenizer files |
The full snapshot is approximately 24 GB.
## Download
```bash
pip install -U huggingface_hub
hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \
--local-dir ./FlashWAM-RoboTwin
```
To inspect configs without downloading the weights:
```bash
hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \
README.md \
transformer/config.json vae/config.json text_encoder/config.json \
tokenizer/tokenizer_config.json \
--local-dir ./FlashWAM-RoboTwin-config
```
## Environment and evaluation
Flash-WAM uses the LingBot-VA environment and the same RoboTwin server/client
evaluation pipeline:
1. Follow the [LingBot-VA installation and RoboTwin evaluation
instructions](https://github.com/Robbyant/lingbot-va).
2. Clone the [Flash-WAM repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM)
so the custom `wan_va` model implementation is available.
3. Download this snapshot and set the model path in the LingBot-VA/Flash-WAM
evaluation configuration to the local snapshot directory.
4. For training or distillation, use the released commands in the Flash-WAM
repository; this checkpoint is the already-distilled student.
The public Flash-WAM repository does not yet include the real-world Unitree G1
deployment setup. Do not infer a supported real-robot deployment command from
the checkpoint layout alone.
## Optional component-loading check
After installing the LingBot-VA environment and making the Flash-WAM repository
available on `PYTHONPATH`, the released helper functions can load the individual
components:
```python
from pathlib import Path
import torch
from wan_va.modules.utils import (
load_text_encoder,
load_tokenizer,
load_transformer,
load_vae,
)
root = Path("/path/to/FlashWAM-RoboTwin")
device = "cuda"
dtype = torch.bfloat16
tokenizer = load_tokenizer(root / "tokenizer")
text_encoder = load_text_encoder(root / "text_encoder", dtype, device)
vae = load_vae(root / "vae", dtype, device)
transformer = load_transformer(root / "transformer", dtype, device)
```
This verifies component compatibility; it is not a complete policy rollout.
Use the upstream server/client evaluation path for observations, action
normalization, temporal caching, and environment interaction.
## Reported results
### RoboTwin 2.0
| Method | Video steps | Action steps | Average success | Speedup |
| --- | ---: | ---: | ---: | ---: |
| LingBot-VA teacher | 25 | 50 | 91.25% | 1.0× |
| Naive joint LCM | 1 | 2 | 23.97% | — |
| **Flash-WAM** | **1** | **2** | **85.54%** | **19.0×** |
| Naive joint LCM | 1 | 1 | 36.32% | — |
| **Flash-WAM** | **1** | **1** | **81.41%** | **23.3×** |
### LIBERO
| Method | Video steps | Action steps | Average success | Speedup |
| --- | ---: | ---: | ---: | ---: |
| LingBot-VA teacher | 20 | 50 | 98.6% | 1.0× |
| **Flash-WAM** | **1** | **2** | **95.7%** | **13.7×** |
| **Flash-WAM** | **1** | **1** | **95.1%** | **16.3×** |
### Real-world Unitree G1
Three manipulation tasks were evaluated with 10 rollouts per task:
| Method | Video/action steps | T1 | T2 | T3 | Average |
| --- | ---: | ---: | ---: | ---: | ---: |
| LingBot-VA | 3 / 10 | 50% | 70% | 80% | 66.7% |
| **Flash-WAM** | **1 / 2** | **50%** | **60%** | **70%** | **60.0%** |
| **Flash-WAM** | **1 / 1** | **40%** | **50%** | **60%** | **50.0%** |
The fastest 1-video/1-action-step configuration and the 60% real-world result
are **not the same configuration**. Report step budgets together with every
success-rate or latency claim.
## Intended use
This checkpoint is intended for:
- research on step distillation for joint video-action models;
- reproducing the reported RoboTwin results;
- comparing modality-aware and naive joint consistency objectives;
- studying latency/task-success trade-offs in world-action models.
It is not a drop-in controller for an arbitrary robot or task. Deployment on
physical robots requires task-specific observation processing, action
normalization, safety constraints, control integration, and validation.
## Limitations and safety
- Results are specific to LingBot-VA, the released RoboTwin checkpoint, and the
paper's evaluation settings.
- Latency depends on GPU, software stack, precision, resolution, horizon, and
server/client overhead; 348 ms is not a universal runtime guarantee.
- The real-world evaluation covers three tasks and 30 rollouts per method.
- One-step generation still reduces task success relative to the teacher.
- Generated actions may be unsafe or incorrect. Use independent safeguards,
workspace limits, emergency stops, and supervised testing before any
physical deployment.
- The real-world G1 deployment setup is not included in the public code release.
## Licenses
The **checkpoint in this Hugging Face repository** is released under
Apache-2.0. It includes components derived from LingBot-VA, whose released
model and bundled upstream components are also Apache-2.0.
The separate [Flash-WAM GitHub repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM)
uses different terms for different parts: the authors' distillation code,
documentation, and demo videos are CC BY-NC 4.0, while the bundled `wan_va/`
components remain Apache-2.0. Downloading this checkpoint does not replace the
license notices of the code or other assets used with it.
## Citation
```bibtex
@misc{akbari2026flashwammodalityawaredistillationworld,
title = {Flash-WAM: Modality-Aware Distillation for World Action Models},
author = {Arman Akbari and Ci Zhang and Arash Akbari and Lin Zhao and Yixiao Chen and Weiwei Chen and Xuan Zhang and Geng Yuan and Yanzhi Wang},
year = {2026},
eprint = {2606.05254},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2606.05254}
}
```
|