CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
Abstract
Interactive CAPTCHAs remain challenging for computer-use agents, while existing datasets face trade-offs among type coverage, interaction fidelity, and trajectory supervision. To address these gaps, we present CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving. It contains 50K puzzles across 20 CAPTCHA types and 5 interaction modes, with every solution verified through execution. CaptchaArena provides 50K screenshot-action trajectories, including 46K with step-by-step reasoning annotations. It also includes fine-grained pixel-mask annotations for irregular targets. Using CaptchaArena, we train CaptchaAgent, a single 9B policy for all 20 CAPTCHA types, with supervised fine-tuning followed by reinforcement learning. The environment verifier directly provides the RL reward. Supervised fine-tuning reaches 70.5 Pass@1, and reinforcement learning further improves it to 71.7, while also improving performance on two external benchmarks. These results demonstrate the value of large-scale, fine-grained computer-use supervision for training interactive CAPTCHA agents. We release CaptchaArena and CaptchaAgent at https://github.com/X0X0X00/CaptchaArena.
Community
We introduce CaptchaArena, a large-scale, fine-grained dataset for interactive CAPTCHA solving, with 50K puzzles across 20 types and 5 interaction modes, including 46K screenshot-action trajectories with step-by-step reasoning. Using CaptchaArena, we train CaptchaAgent, a unified 9B computer-use policy with supervised fine-tuning and reinforcement learning.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent (2026)
- Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs (2026)
- Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom (2026)
- StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning (2026)
- Momentum-Coupled Rubric Adaptation for Detailed Image Captioning (2026)
- RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing (2026)
- RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 1
Datasets citing this paper 2
ZHEN-04/CaptchaArena-Trajectories
ZHEN-04/CaptchaArena
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper