When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents Paper • 2609.32520 • Published 11 days ago • 20
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents Paper • 2609.32536 • Published 11 days ago • 12
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads Paper • 2608.04570 • Published Aug 5 • 41
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents Paper • 2608.04574 • Published Aug 5 • 16
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants Paper • 2607.26611 • Published Jul 29 • 33
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering Paper • 2603.28583 • Published Jul 14 • 10
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Paper • 2605.14906 • Published May 14 • 78
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? Paper • 2605.06527 • Published May 7 • 48
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Paper • 2605.13831 • Published May 13 • 90