SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 15 days ago • 51
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published 17 days ago • 18
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 21 days ago • 71
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Paper • 2608.06352 • Published 28 days ago • 23
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published Jul 23 • 154
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 62
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published Jul 9 • 77
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding Paper • 2606.31315 • Published Jun 30 • 77
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published Jun 26 • 48
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published Jun 28 • 24
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 53
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Paper • 2606.15932 • Published Jun 16 • 38
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published Jun 24 • 18
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 160
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions Paper • 2606.23654 • Published Jun 22 • 80
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 37