ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs Paper • 2609.10895 • Published 6 days ago • 37 • 3
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 1 day ago • 134 • 2
TempCloze: Can Video-LLMs Identify the Missing Middle? Paper • 2609.01515 • Published 14 days ago • 31 • 3
Steering Geometry: Validating Human Value Geometry in LLM Steering Space Paper • 2609.06289 • Published 10 days ago • 33 • 4
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Paper • 2609.10016 • Published 6 days ago • 33 • 5
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation Paper • 2609.11486 • Published 5 days ago • 34 • 3
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 6 days ago • 38 • 2
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 5 days ago • 44 • 3
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Paper • 2609.06245 • Published 10 days ago • 34 • 3
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published 8 days ago • 135 • 3
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 11 days ago • 41 • 3
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Paper • 2609.05903 • Published 10 days ago • 59 • 3
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 5 days ago • 250 • 3
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 6 days ago • 308 • 3
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 7 days ago • 41 • 2
Scaling Automatic Research Agents via World Models Paper • 2608.12564 • Published 17 days ago • 452 • 3
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published 7 days ago • 94 • 3