Convert to the alias-free PI0.5 layout (813 tensors)
Browse files- README.md +19 -11
- model.safetensors +2 -2
README.md
CHANGED
|
@@ -9,36 +9,44 @@ base_model:
|
|
| 9 |
---
|
| 10 |
# FlashVLA · π0.5 · RoboTwin 2.0
|
| 11 |
|
| 12 |
-
|
| 13 |
A **π0.5** flow-matching vision-language-action policy finetuned on **RoboTwin 2.0**
|
| 14 |
(50-task multitask) and served with [**FlashVLA**](https://github.com/z-lab/flashvla)
|
| 15 |
streaming action decoding for fast, asynchronous inference.
|
| 16 |
|
| 17 |
-
|
| 18 |
- **Base model:** [`lerobot/pi05_base`](https://huggingface.co/lerobot/pi05_base)
|
| 19 |
- **Method:** [FlashVLA](https://github.com/z-lab/flashvla) — streaming action decoding for flow-matching VLAs (async chunk-overlap execution)
|
| 20 |
- **Benchmark:** RoboTwin 2.0, 50-task multitask (clean / randomized)
|
| 21 |
- **License:** Apache-2.0
|
| 22 |
|
| 23 |
-
|
| 24 |
## Results
|
| 25 |
|
| 26 |
-
|
| 27 |
RoboTwin 2.0 50-task multitask success rate (%). `d` is the async step delay:
|
| 28 |
`d=0` is synchronous, `d=1`/`d=2` overlap the next chunk's inference with execution.
|
| 29 |
|
| 30 |
-
|
| 31 |
| Model | Clean | Random | Avg |
|
| 32 |
|:--|:--:|:--:|:--:|
|
| 33 |
-
| π0.5 (base) |
|
| 34 |
-
| **+FlashVLA** (`d=0`) | 90.
|
| 35 |
-
| **+FlashVLA** (`d=1`) | **91.
|
| 36 |
-
| **+FlashVLA** (`d=2`) |
|
| 37 |
|
| 38 |
|
|
|
|
| 39 |
|
|
|
|
| 40 |
|
| 41 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
|
|
|
|
| 43 |
|
| 44 |
-
|
|
|
|
| 9 |
---
|
| 10 |
# FlashVLA · π0.5 · RoboTwin 2.0
|
| 11 |
|
|
|
|
| 12 |
A **π0.5** flow-matching vision-language-action policy finetuned on **RoboTwin 2.0**
|
| 13 |
(50-task multitask) and served with [**FlashVLA**](https://github.com/z-lab/flashvla)
|
| 14 |
streaming action decoding for fast, asynchronous inference.
|
| 15 |
|
|
|
|
| 16 |
- **Base model:** [`lerobot/pi05_base`](https://huggingface.co/lerobot/pi05_base)
|
| 17 |
- **Method:** [FlashVLA](https://github.com/z-lab/flashvla) — streaming action decoding for flow-matching VLAs (async chunk-overlap execution)
|
| 18 |
- **Benchmark:** RoboTwin 2.0, 50-task multitask (clean / randomized)
|
| 19 |
- **License:** Apache-2.0
|
| 20 |
|
|
|
|
| 21 |
## Results
|
| 22 |
|
|
|
|
| 23 |
RoboTwin 2.0 50-task multitask success rate (%). `d` is the async step delay:
|
| 24 |
`d=0` is synchronous, `d=1`/`d=2` overlap the next chunk's inference with execution.
|
| 25 |
|
|
|
|
| 26 |
| Model | Clean | Random | Avg |
|
| 27 |
|:--|:--:|:--:|:--:|
|
| 28 |
+
| π0.5 (base) | 82.74 | 76.76 | 79.75 |
|
| 29 |
+
| **+FlashVLA** (`d=0`) | 90.64 | 90.06 | 90.35 |
|
| 30 |
+
| **+FlashVLA** (`d=1`) | **91.14** | **90.60** | **90.87** |
|
| 31 |
+
| **+FlashVLA** (`d=2`) | 90.20 | 89.66 | 89.93 |
|
| 32 |
|
| 33 |
|
| 34 |
+
## Usage
|
| 35 |
|
| 36 |
+
Install the [FlashVLA](https://github.com/z-lab/flashvla) library.
|
| 37 |
|
| 38 |
+
See [`sim_eval/robotwin/`](https://github.com/z-lab/flashvla/tree/main/sim_eval/robotwin) for evaluation setup.
|
| 39 |
+
|
| 40 |
+
## Citation
|
| 41 |
+
|
| 42 |
+
```bibtex
|
| 43 |
+
@article{flashvla2026,
|
| 44 |
+
title = {{FlashVLA: Streaming Action Decoding for Fast and Asynchronous VLA Inference}},
|
| 45 |
+
author = {Li, Zekai and Tang, Jiaming and Liu, Zhijian},
|
| 46 |
+
year = {2026}
|
| 47 |
+
}
|
| 48 |
+
```
|
| 49 |
|
| 50 |
+
## License
|
| 51 |
|
| 52 |
+
Apache-2.0
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:143af9b4890dedad78bb1246dd144069f7fdb50872e964de2421bb88a455730e
|
| 3 |
+
size 16573735272
|