--- license: gemma pipeline_tag: robotics language: - en inference: false base_model: - lerobot/pi05_base datasets: - TianxingChen/RoboTwin2.0 tags: - flashvla - vision-language-action - robotics - flow-matching - streaming-action-decoding - pi05 --- # FlashVLA · π0.5 · RoboTwin 2.0 **Streaming Action Decoding for Fast and Asynchronous VLA Inference** [![Paper](https://img.shields.io/badge/arXiv-2608.27384-b31b1b.svg)](https://arxiv.org/abs/2608.27384) [![GitHub](https://img.shields.io/badge/GitHub-FlashVLA-181717?logo=github)](https://github.com/z-lab/flashvla) [![Blog](https://img.shields.io/badge/Blog-FlashVLA-blue)](https://z-lab.ai/projects/flashvla/) [![Models](https://img.shields.io/badge/%F0%9F%A4%97-Models-yellow)](https://huggingface.co/collections/z-lab/flashvla) A **π0.5** flow-matching vision-language-action policy finetuned on **RoboTwin 2.0** (50-task multitask) and served with [FlashVLA](https://github.com/z-lab/flashvla) streaming action decoding for fast, asynchronous inference. - **Base model:** [`lerobot/pi05_base`](https://huggingface.co/lerobot/pi05_base) - **Method:** [FlashVLA](https://github.com/z-lab/flashvla), streaming action decoding for flow-matching VLAs (async chunk-overlap execution) - **Benchmark:** RoboTwin 2.0, 50-task multitask (clean / randomized) ## Results RoboTwin 2.0 50-task multitask success rate (%). `d` is the async step delay: `d=0` is synchronous, `d=1` and `d=2` overlap the next chunk's inference with execution. | Model | Clean | Random | Avg | |:--|:--:|:--:|:--:| | π0.5 (base) | 82.74 | 76.76 | 79.75 | | **+FlashVLA** (`d=0`) | 90.64 | 90.06 | 90.35 | | **+FlashVLA** (`d=1`) | **91.14** | **90.60** | **90.87** | | **+FlashVLA** (`d=2`) | 90.20 | 89.66 | 89.93 | ## Usage Install FlashVLA: ```bash git clone https://github.com/z-lab/flashvla.git cd flashvla conda env create -f environment.yml conda activate flashvla ``` RoboTwin 2.0 evaluation uses a server/client split across two environments. After the one-time setup in [`sim_eval/robotwin/`](https://github.com/z-lab/flashvla/tree/main/sim_eval/robotwin), from `$ROBOTWIN/policy/pi05_flashvla/`: ```bash bash eval_server.sh # terminal 1, flashvla env, starts the policy server ROBOTWIN_VENV=... bash eval_client.sh # terminal 2, RoboTwin env, runs the SAPIEN sim ``` Training configs for this checkpoint are in [`train/configs/pi05/robotwin/`](https://github.com/z-lab/flashvla/tree/main/train/configs/pi05/robotwin). ## License These weights are finetuned from [`lerobot/pi05_base`](https://huggingface.co/lerobot/pi05_base), which is released under the [Gemma Terms of Use](https://ai.google.dev/gemma/terms). Those terms govern model derivatives, so they apply to this checkpoint and to anything derived from it, including the [Gemma Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy). If you redistribute this checkpoint or a derivative of it, you must pass the same terms along. The [FlashVLA](https://github.com/z-lab/flashvla) inference and training code is separately released under the [Apache 2.0 License](https://github.com/z-lab/flashvla/blob/main/LICENSE). ## Citation ```bibtex @inproceedings{li2026flashvla, title = {{FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference}}, author = {Li, Zekai and Tang, Jiaming and Liu, Zhijian}, booktitle = {Conference on Robot Learning (CoRL)}, year = {2026} } ```