maanavdalal commited on
Commit
3d0887b
·
1 Parent(s): 4c8386b

Update FLUX 3 Action model card

Browse files
Files changed (1) hide show
  1. README.md +7 -7
README.md CHANGED
@@ -16,17 +16,17 @@ FLUX 3 Action is an open weights 7B world action model. It takes camera frames,
16
 
17
  For more information, read the [documentation](https://docs.bfl.ai/flux_3/flux3_action_overview).
18
 
19
- Fine-tuned on [DROID](https://droid-dataset.github.io/), FLUX 3 Action places first on the RoboLab-120 benchmark at 42.6% task success.
20
 
21
  ## Evaluation
22
 
23
- RoboLab-120 is 120 tabletop tasks in Isaac Sim, 10 trials each, on a DROID-style Franka setup; a trial succeeds only if the task is completed as instructed. The full board is on the [RoboLab leaderboard](https://research.nvidia.com/labs/srl/projects/robolab/leaderboard.html). Latencies were measured on one H200.
24
 
25
- | Model | Type | Success | Parameters | Plan latency |
26
- | --- | --- | --- | --- | --- |
27
- | FLUX 3 Action | WAM | 42.6% | 7B | 481 ms bf16, 290 ms FP8 |
28
- | Cosmos3-Nano-Policy | WAM | 36.8% | 16B | 724 ms bf16 |
29
- | π0.5 | VLA | 28.0% | 3.3B | |
30
 
31
  WAM: predicts future frames and actions together. VLA: a vision-language model that outputs actions directly.
32
 
 
16
 
17
  For more information, read the [documentation](https://docs.bfl.ai/flux_3/flux3_action_overview).
18
 
19
+ Fine-tuned on [DROID](https://droid-dataset.github.io/), FLUX 3 Action places first on the RoboLab-120 benchmark at 42.92% task success.
20
 
21
  ## Evaluation
22
 
23
+ RoboLab-120 is 120 tabletop tasks in Isaac Sim, 10 trials each, on a DROID-style Franka setup; a trial succeeds only if the task is completed as instructed. The full board is on the [RoboLab leaderboard](https://research.nvidia.com/labs/srl/projects/robolab/leaderboard.html).
24
 
25
+ | Model | Type | Success | Parameters |
26
+ | --- | --- | --- | --- |
27
+ | FLUX 3 Action | WAM | 42.92% | 7B |
28
+ | Cosmos3-Nano-Policy | WAM | 36.8% | 16B |
29
+ | π0.5 | VLA | 28.0% | 3.3B |
30
 
31
  WAM: predicts future frames and actions together. VLA: a vision-language model that outputs actions directly.
32