SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 13 days ago • 51
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents Paper • 2602.13379 • Published Feb 13 • 3
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 13 days ago • 51
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 106
Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace Paper • 2605.10913 • Published May 11 • 4
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games Paper • 2603.09022 • Published Mar 9 • 25
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Paper • 2604.14142 • Published Apr 15 • 30
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation Paper • 2603.18886 • Published Mar 19 • 6
SPICE: Self-Play In Corpus Environments Improves Reasoning Paper • 2510.24684 • Published Oct 28, 2025 • 18
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity Paper • 2510.01171 • Published Oct 1, 2025 • 19
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution Paper • 2510.08697 • Published Oct 9, 2025 • 40