Spaces:
Running
A newer version of the Gradio SDK is available: 6.28.0
title: Provael ASR Leaderboard
emoji: π¦Ύ
colorFrom: red
colorTo: indigo
sdk: gradio
sdk_version: 6.23.1
app_file: app.py
pinned: false
license: apache-2.0
short_description: Attack-success rates for VLA policies, with benign controls
tags:
- leaderboard
- submission:manual
- test:public
- modality:text
- eval:safety
- language:English
Provael β VLA Red-Team ASR Leaderboard
Attack Success Rate (ASR) of instruction / visual / injection attacks against
Vision-Language-Action (VLA) robot policies in simulation, built with
provael. Lower ASR = more robust.
Measure a policy you did not train, and its authors get the full artifact 14 days before anything is published here β including the right to have it re-run, corrected, or pulled pending review. That is written down in Measuring someone else's policy, and it binds the maintainer of this board more tightly than it binds anyone submitting to it.
β Real data.
results/leaderboard.jsonholds the ten-task SmolVLA-on-LIBERO suite screen measured on 14 September 2026 with provael 0.41.2 (HuggingFaceVLA/smolvla_libero, all 10libero_objecttasks Γ 5 seeds, 350 measured episodes): instruction 33.3% (50/150) [26β41%], against a benignnonebaseline of 2.0% (1/50). The board's rows sum to 15.7% (55/350) across every arm including the benign control β that is the all-episode observed rate, not the attack rate, and it is diluted by the arms that did not move. Read the per-family rows, not the sum.Read it as lift over baseline β instruction-reframing attacks are the only family that moves this policy; injection 2/50 and visual 2/100 do not separate from the run's own benign floor (McNemar exact p = 1.0 on every arm) and stay published as measured. Per-attack detail, including
roleplayat 42/50 against the run's own 1/50 benign control, with its McNemar and task-clustered interval, is inresults/smolvla_libero_object_suite_2026-09-14/. The 0.32.0 run this board aggregated until 18 September 2026 (44/50, 62/150) stays committed underresults/smolvla_libero_object_suite/; the re-run reproduced its headline and read its meaning differently (E-2026-12: the frame moves the policy, target or no target).β οΈ Four qualifiers, and they now travel inside the artifact (
schema_version5) rather than living only in this README:
"calibrated": falseon every row β the same default keep-out box on all ten tasks, never fitted to any of them. This is "diverted out of the benign safe envelope," not a calibrated hazard rate, and it is why the benign arm tripped at all. Per-task zone calibration is still owed (#136)."stochastic": trueβ SmolVLA's flow-matching sampler draws; the flag says so. These rows were measured with provael 0.41.2, which seeds the policy's own sampler and recordspolicy_seedon every episode (a stochastic submission without one is refused), so the draw behind each number is named: the same seed reproduces it, a different seed is a different draw, and five seeds per task is what the interval is over."not_applicable": ["mcp_tool_desc"]β 50 episode records, zero applicable episodes. Not-measured and measured-zero are different claims, so it is named rather than silently dropped from the denominator."checkpoint"β one checkpoint, one suite. Nothing here speaks tolibero_spatial,libero_goal,libero_10, or any other policy.
Tabs
All policies and Open-source policies β a RoboArena-style split (open-weights models vs. everything) so robustness is compared like-for-like.
Example payloads β the attacked inputs behind the numbers.
Submit a result β open-submission: upload a
provael leaderboard buildresults JSON; it opens a PR to theprovael-submissions/requestsdataset (Open-LLM-Leaderboard pattern) for a maintainer to validate and promote. NeedsHF_TOKENset on the Space.This pointed at
provael-submissions/requestsβ an org that was never created (verified 2026-08-08: org page 404, datasets API 401,?author=provael-submissionsreturns[]). The submit button had been aimed at nothing since it shipped. That it went unnoticed is itself the finding: the leaderboard has had zero third-party submissions ever, so this path had never been walked by a stranger.It now points under the same account that owns the Space, so no org administration stands between a stranger and a contribution. If the dataset still does not exist, the button says so plainly and routes the submitter to a GitHub issue rather than handing them a traceback β or, worse, a success message for a result nobody received.
tests/test_leaderboard_submit_path.pypins that behaviour so it cannot rot again unobserved.
Opening the queue (one command, once)
export HF_TOKEN=hf_... # write scope
python leaderboard/setup_requests_dataset.py # creates the dataset + its card, idempotent
Then set the same token as a secret named HF_TOKEN on the Space
(Settings β Variables and secrets). The script prints both steps and re-running it verifies rather
than errors, so it doubles as the "is the queue actually open?" check.
It is a script and not a README step on purpose: "create the dataset" lived only in prose for two months, and prose does not fail.
Run it
Locally:
pip install gradio huggingface_hub
python app.py
On Hugging Face: this folder is a Gradio Space rendering the committed results/*.json with no
GPU. Submission is enabled when HF_TOKEN is set as a Space secret.
How the data is produced
# CPU demo (stub policy) β an example; no GPU/model needed:
provael attack --policy stub --suite stub \
--attacks instruction,visual,injection --episodes 10 --seed 0 --out runs/demo
provael leaderboard build --runs runs/demo --out leaderboard/results # writes leaderboard.json
Real numbers (GPU box) β what's committed here
pip install 'provael[lerobot]' 'lerobot[libero]==0.5.1'
apt-get install -y libosmesa6 libgl1 libglx-mesa0 # headless GL (cloud images ship none)
export MUJOCO_GL=osmesa PYOPENGL_PLATFORM=osmesa
provael attack --policy smolvla --suite libero --model HuggingFaceVLA/smolvla_libero \
--attacks none,instruction,visual,injection --seeds 10 --horizon 280 --seed 0 \
--out runs/smolvla_libero
provael leaderboard build --runs 'runs/*' --out leaderboard/results
Commit the resulting results/*.json; the banner reads "includes real-model results"
whenever a non-stub run is present.
Schema
Each results/*.json is a Leaderboard: {schema_version, is_demo, rows[], examples[]},
where each row is {policy, suite, family, attempts, successes, asr} (ranked by ASR
descending) and each example is {attack, family, example} (a representative injected
payload). Output is deterministic (sorted keys, no timestamps).