MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models Paper • 2605.28009 • Published May 27 • 3
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 21 days ago • 114
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 27 days ago • 40
Trimming the Long-Tail of Visual World Modeling Evaluation Paper • 2606.24256 • Published Jun 23 • 43
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96
Brick-Composer: Using MLLMs for Assembly with Diverse Bricks Paper • 2606.05445 • Published Jun 3 • 8
AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints Paper • 2606.05622 • Published Jun 4 • 44
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing Paper • 2605.02910 • Published May 6 • 24
Toward Cognitive Supersensing in Multimodal Large Language Model Paper • 2602.01541 • Published Feb 2 • 16
MedSAM3: Delving into Segment Anything with Medical Concepts Paper • 2511.19046 • Published Nov 24, 2025 • 56
Analyzing and Internalizing Complex Policy Documents for LLM Agents Paper • 2510.11588 • Published Oct 13, 2025 • 1
Multimodal Policy Internalization for Conversational Agents Paper • 2510.09474 • Published Oct 10, 2025 • 5