--- license: apache-2.0 library_name: custom base_model: robbyant/lingbot-va-posttrain-robotwin datasets: - robbyant/robotwin-clean-and-aug-lerobot pipeline_tag: robotics language: - en tags: - robotics - embodied-ai - world-action-model - world-model - diffusion - step-distillation - lingbot-va - robotwin - arxiv:2606.05254 --- # Flash-WAM RoboTwin: Distilled World-Action Model [Project page](https://flashwam.github.io/) · [Paper](https://arxiv.org/abs/2606.05254) · [Code](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) · [LingBot-VA](https://github.com/Robbyant/lingbot-va) This repository contains the complete RoboTwin checkpoint for **Flash-WAM: Modality-Aware Distillation for World Action Models**. Flash-WAM distills the joint video and action streams of LingBot-VA with consistency functions matched to their different noise regimes. The released student supports one-step video and one-step action generation. Under the paper's RoboTwin 2.0 setup on a single NVIDIA L40S, this reduces per-chunk latency from **8.1 seconds to 348 milliseconds**, a **23.3× speedup**. > **Important:** this is a custom joint video-action robotics model, not a > generic text-to-image or video `DiffusionPipeline`. Do not use > `DiffusionPipeline.from_pretrained(...)`. Install the Flash-WAM/LingBot-VA > code and use their RoboTwin server/client evaluation path. ## Model details | Field | Value | | --- | --- | | Base model | [LingBot-VA RoboTwin post-training checkpoint](https://huggingface.co/robbyant/lingbot-va-posttrain-robotwin) | | Task | Joint future-video and robot-action prediction | | Benchmark | RoboTwin 2.0 | | Released student | 1 video step / 1 action step | | Action dimension | 30 in the released transformer config | | Reported latency hardware | 1 × NVIDIA L40S | | Checkpoint license | Apache-2.0 | ## Repository contents | Directory | Description | | --- | --- | | `transformer/` | Distilled Flash-WAM student, approximately 10 GB | | `vae/` | VAE inherited from the LingBot-VA teacher, approximately 2.8 GB | | `text_encoder/` | UMT5-XXL text encoder, approximately 11.3 GB | | `tokenizer/` | T5 tokenizer files | The full snapshot is approximately 24 GB. ## Download ```bash pip install -U huggingface_hub hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \ --local-dir ./FlashWAM-RoboTwin ``` To inspect configs without downloading the weights: ```bash hf download NU-World-Model-Embodied-AI/FlashWAM-RoboTwin \ README.md \ transformer/config.json vae/config.json text_encoder/config.json \ tokenizer/tokenizer_config.json \ --local-dir ./FlashWAM-RoboTwin-config ``` ## Environment and evaluation Flash-WAM uses the LingBot-VA environment and the same RoboTwin server/client evaluation pipeline: 1. Follow the [LingBot-VA installation and RoboTwin evaluation instructions](https://github.com/Robbyant/lingbot-va). 2. Clone the [Flash-WAM repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) so the custom `wan_va` model implementation is available. 3. Download this snapshot and set the model path in the LingBot-VA/Flash-WAM evaluation configuration to the local snapshot directory. 4. For training or distillation, use the released commands in the Flash-WAM repository; this checkpoint is the already-distilled student. The public Flash-WAM repository does not yet include the real-world Unitree G1 deployment setup. Do not infer a supported real-robot deployment command from the checkpoint layout alone. ## Optional component-loading check After installing the LingBot-VA environment and making the Flash-WAM repository available on `PYTHONPATH`, the released helper functions can load the individual components: ```python from pathlib import Path import torch from wan_va.modules.utils import ( load_text_encoder, load_tokenizer, load_transformer, load_vae, ) root = Path("/path/to/FlashWAM-RoboTwin") device = "cuda" dtype = torch.bfloat16 tokenizer = load_tokenizer(root / "tokenizer") text_encoder = load_text_encoder(root / "text_encoder", dtype, device) vae = load_vae(root / "vae", dtype, device) transformer = load_transformer(root / "transformer", dtype, device) ``` This verifies component compatibility; it is not a complete policy rollout. Use the upstream server/client evaluation path for observations, action normalization, temporal caching, and environment interaction. ## Reported results ### RoboTwin 2.0 | Method | Video steps | Action steps | Average success | Speedup | | --- | ---: | ---: | ---: | ---: | | LingBot-VA teacher | 25 | 50 | 91.25% | 1.0× | | Naive joint LCM | 1 | 2 | 23.97% | — | | **Flash-WAM** | **1** | **2** | **85.54%** | **19.0×** | | Naive joint LCM | 1 | 1 | 36.32% | — | | **Flash-WAM** | **1** | **1** | **81.41%** | **23.3×** | ### LIBERO | Method | Video steps | Action steps | Average success | Speedup | | --- | ---: | ---: | ---: | ---: | | LingBot-VA teacher | 20 | 50 | 98.6% | 1.0× | | **Flash-WAM** | **1** | **2** | **95.7%** | **13.7×** | | **Flash-WAM** | **1** | **1** | **95.1%** | **16.3×** | ### Real-world Unitree G1 Three manipulation tasks were evaluated with 10 rollouts per task: | Method | Video/action steps | T1 | T2 | T3 | Average | | --- | ---: | ---: | ---: | ---: | ---: | | LingBot-VA | 3 / 10 | 50% | 70% | 80% | 66.7% | | **Flash-WAM** | **1 / 2** | **50%** | **60%** | **70%** | **60.0%** | | **Flash-WAM** | **1 / 1** | **40%** | **50%** | **60%** | **50.0%** | The fastest 1-video/1-action-step configuration and the 60% real-world result are **not the same configuration**. Report step budgets together with every success-rate or latency claim. ## Intended use This checkpoint is intended for: - research on step distillation for joint video-action models; - reproducing the reported RoboTwin results; - comparing modality-aware and naive joint consistency objectives; - studying latency/task-success trade-offs in world-action models. It is not a drop-in controller for an arbitrary robot or task. Deployment on physical robots requires task-specific observation processing, action normalization, safety constraints, control integration, and validation. ## Limitations and safety - Results are specific to LingBot-VA, the released RoboTwin checkpoint, and the paper's evaluation settings. - Latency depends on GPU, software stack, precision, resolution, horizon, and server/client overhead; 348 ms is not a universal runtime guarantee. - The real-world evaluation covers three tasks and 30 rollouts per method. - One-step generation still reduces task success relative to the teacher. - Generated actions may be unsafe or incorrect. Use independent safeguards, workspace limits, emergency stops, and supervised testing before any physical deployment. - The real-world G1 deployment setup is not included in the public code release. ## Licenses The **checkpoint in this Hugging Face repository** is released under Apache-2.0. It includes components derived from LingBot-VA, whose released model and bundled upstream components are also Apache-2.0. The separate [Flash-WAM GitHub repository](https://github.com/NU-World-Model-Embodied-AI/Flash-WAM) uses different terms for different parts: the authors' distillation code, documentation, and demo videos are CC BY-NC 4.0, while the bundled `wan_va/` components remain Apache-2.0. Downloading this checkpoint does not replace the license notices of the code or other assets used with it. ## Citation ```bibtex @misc{akbari2026flashwammodalityawaredistillationworld, title = {Flash-WAM: Modality-Aware Distillation for World Action Models}, author = {Arman Akbari and Ci Zhang and Arash Akbari and Lin Zhao and Yixiao Chen and Weiwei Chen and Xuan Zhang and Geng Yuan and Yanzhi Wang}, year = {2026}, eprint = {2606.05254}, archivePrefix = {arXiv}, primaryClass = {cs.LG}, url = {https://arxiv.org/abs/2606.05254} } ```