|
Download README.md from NU-World-Model-Embodied-AI/FlashWAM-RoboTwin: direct link, hf CLI and curl.
- Browser
- Download file 7.97 kB
-
https://huggingface.co/NU-World-Model-Embodied-AI/FlashWAM-RoboTwin/resolve/main/README.md
- Command line
-
hf download hf://NU-World-Model-Embodied-AI/FlashWAM-RoboTwin/README.md
-
curl -L -o README.md https://huggingface.co/NU-World-Model-Embodied-AI/FlashWAM-RoboTwin/resolve/main/README.md
7.97 kB
| license: apache-2.0 | |
| library_name: custom | |
| base_model: robbyant/lingbot-va-posttrain-robotwin | |
| datasets: | |
| - robbyant/robotwin-clean-and-aug-lerobot | |
| pipeline_tag: robotics | |
| language: | |
| - en | |
| tags: | |
| - robotics | |
| - embodied-ai | |
| - world-action-model | |
| - world-model | |
| - diffusion | |
| - step-distillation | |
| - lingbot-va | |
| - robotwin | |
| - arxiv:2606.05254 | |
| # Flash-WAM RoboTwin: Distilled World-Action Model | |
| [Project page](https://flashwam.github.io/) · | |
| [Paper](https://arxiv.org/abs/2606.05254) · | |
| [Code](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) · | |
| [LingBot-VA](https://github.com/Robbyant/lingbot-va) | |
| This repository contains the complete RoboTwin checkpoint for **Flash-WAM: | |
| Modality-Aware Distillation for World Action Models**. Flash-WAM distills the | |
| joint video and action streams of LingBot-VA with consistency functions matched | |
| to their different noise regimes. | |
| The released student supports one-step video and one-step action generation. | |
| Under the paper's RoboTwin 2.0 setup on a single NVIDIA L40S, this reduces | |
| per-chunk latency from **8.1 seconds to 348 milliseconds**, a **23.3× speedup**. | |
| > **Important:** this is a custom joint video-action robotics model, not a | |
| > generic text-to-image or video `DiffusionPipeline`. Do not use | |
| > `DiffusionPipeline.from_pretrained(...)`. Install the Flash-WAM/LingBot-VA | |
| > code and use their RoboTwin server/client evaluation path. | |
| ## Model details | |
| | Field | Value | | |
| | --- | --- | | |
| | Base model | [LingBot-VA RoboTwin post-training checkpoint](https://huggingface.co/robbyant/lingbot-va-posttrain-robotwin) | | |
| | Task | Joint future-video and robot-action prediction | | |
| | Benchmark | RoboTwin 2.0 | | |
| | Released student | 1 video step / 1 action step | | |
| | Action dimension | 30 in the released transformer config | | |
| | Reported latency hardware | 1 × NVIDIA L40S | | |
| | Checkpoint license | Apache-2.0 | | |
| ## Repository contents | |
| | Directory | Description | | |
| | --- | --- | | |
| | `transformer/` | Distilled Flash-WAM student, approximately 10 GB | | |
| | `vae/` | VAE inherited from the LingBot-VA teacher, approximately 2.8 GB | | |
| | `text_encoder/` | UMT5-XXL text encoder, approximately 11.3 GB | | |
| | `tokenizer/` | T5 tokenizer files | | |
| The full snapshot is approximately 24 GB. | |
| ## Download | |
| ```bash | |
| pip install -U huggingface_hub | |
| hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \ | |
| --local-dir ./FlashWAM-RoboTwin | |
| ``` | |
| To inspect configs without downloading the weights: | |
| ```bash | |
| hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \ | |
| README.md \ | |
| transformer/config.json vae/config.json text_encoder/config.json \ | |
| tokenizer/tokenizer_config.json \ | |
| --local-dir ./FlashWAM-RoboTwin-config | |
| ``` | |
| ## Environment and evaluation | |
| Flash-WAM uses the LingBot-VA environment and the same RoboTwin server/client | |
| evaluation pipeline: | |
| 1. Follow the [LingBot-VA installation and RoboTwin evaluation | |
| instructions](https://github.com/Robbyant/lingbot-va). | |
| 2. Clone the [Flash-WAM repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) | |
| so the custom `wan_va` model implementation is available. | |
| 3. Download this snapshot and set the model path in the LingBot-VA/Flash-WAM | |
| evaluation configuration to the local snapshot directory. | |
| 4. For training or distillation, use the released commands in the Flash-WAM | |
| repository; this checkpoint is the already-distilled student. | |
| The public Flash-WAM repository does not yet include the real-world Unitree G1 | |
| deployment setup. Do not infer a supported real-robot deployment command from | |
| the checkpoint layout alone. | |
| ## Optional component-loading check | |
| After installing the LingBot-VA environment and making the Flash-WAM repository | |
| available on `PYTHONPATH`, the released helper functions can load the individual | |
| components: | |
| ```python | |
| from pathlib import Path | |
| import torch | |
| from wan_va.modules.utils import ( | |
| load_text_encoder, | |
| load_tokenizer, | |
| load_transformer, | |
| load_vae, | |
| ) | |
| root = Path("/path/to/FlashWAM-RoboTwin") | |
| device = "cuda" | |
| dtype = torch.bfloat16 | |
| tokenizer = load_tokenizer(root / "tokenizer") | |
| text_encoder = load_text_encoder(root / "text_encoder", dtype, device) | |
| vae = load_vae(root / "vae", dtype, device) | |
| transformer = load_transformer(root / "transformer", dtype, device) | |
| ``` | |
| This verifies component compatibility; it is not a complete policy rollout. | |
| Use the upstream server/client evaluation path for observations, action | |
| normalization, temporal caching, and environment interaction. | |
| ## Reported results | |
| ### RoboTwin 2.0 | |
| | Method | Video steps | Action steps | Average success | Speedup | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | LingBot-VA teacher | 25 | 50 | 91.25% | 1.0× | | |
| | Naive joint LCM | 1 | 2 | 23.97% | — | | |
| | **Flash-WAM** | **1** | **2** | **85.54%** | **19.0×** | | |
| | Naive joint LCM | 1 | 1 | 36.32% | — | | |
| | **Flash-WAM** | **1** | **1** | **81.41%** | **23.3×** | | |
| ### LIBERO | |
| | Method | Video steps | Action steps | Average success | Speedup | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | LingBot-VA teacher | 20 | 50 | 98.6% | 1.0× | | |
| | **Flash-WAM** | **1** | **2** | **95.7%** | **13.7×** | | |
| | **Flash-WAM** | **1** | **1** | **95.1%** | **16.3×** | | |
| ### Real-world Unitree G1 | |
| Three manipulation tasks were evaluated with 10 rollouts per task: | |
| | Method | Video/action steps | T1 | T2 | T3 | Average | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | LingBot-VA | 3 / 10 | 50% | 70% | 80% | 66.7% | | |
| | **Flash-WAM** | **1 / 2** | **50%** | **60%** | **70%** | **60.0%** | | |
| | **Flash-WAM** | **1 / 1** | **40%** | **50%** | **60%** | **50.0%** | | |
| The fastest 1-video/1-action-step configuration and the 60% real-world result | |
| are **not the same configuration**. Report step budgets together with every | |
| success-rate or latency claim. | |
| ## Intended use | |
| This checkpoint is intended for: | |
| - research on step distillation for joint video-action models; | |
| - reproducing the reported RoboTwin results; | |
| - comparing modality-aware and naive joint consistency objectives; | |
| - studying latency/task-success trade-offs in world-action models. | |
| It is not a drop-in controller for an arbitrary robot or task. Deployment on | |
| physical robots requires task-specific observation processing, action | |
| normalization, safety constraints, control integration, and validation. | |
| ## Limitations and safety | |
| - Results are specific to LingBot-VA, the released RoboTwin checkpoint, and the | |
| paper's evaluation settings. | |
| - Latency depends on GPU, software stack, precision, resolution, horizon, and | |
| server/client overhead; 348 ms is not a universal runtime guarantee. | |
| - The real-world evaluation covers three tasks and 30 rollouts per method. | |
| - One-step generation still reduces task success relative to the teacher. | |
| - Generated actions may be unsafe or incorrect. Use independent safeguards, | |
| workspace limits, emergency stops, and supervised testing before any | |
| physical deployment. | |
| - The real-world G1 deployment setup is not included in the public code release. | |
| ## Licenses | |
| The **checkpoint in this Hugging Face repository** is released under | |
| Apache-2.0. It includes components derived from LingBot-VA, whose released | |
| model and bundled upstream components are also Apache-2.0. | |
| The separate [Flash-WAM GitHub repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) | |
| uses different terms for different parts: the authors' distillation code, | |
| documentation, and demo videos are CC BY-NC 4.0, while the bundled `wan_va/` | |
| components remain Apache-2.0. Downloading this checkpoint does not replace the | |
| license notices of the code or other assets used with it. | |
| ## Citation | |
| ```bibtex | |
| @misc{akbari2026flashwammodalityawaredistillationworld, | |
| title = {Flash-WAM: Modality-Aware Distillation for World Action Models}, | |
| author = {Arman Akbari and Ci Zhang and Arash Akbari and Lin Zhao and Yixiao Chen and Weiwei Chen and Xuan Zhang and Geng Yuan and Yanzhi Wang}, | |
| year = {2026}, | |
| eprint = {2606.05254}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.LG}, | |
| url = {https://arxiv.org/abs/2606.05254} | |
| } | |
| ``` | |