Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning Paper • 2608.09926 • Published 22 days ago • 14
WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes Paper • 2605.15843 • Published May 15 • 6
DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo Paper • 2605.16257 • Published May 15 • 55
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models Paper • 2605.15961 • Published May 15 • 10
Map2World: Segment Map Conditioned Text to 3D World Generation Paper • 2605.00781 • Published May 1 • 26
When Do Diffusion Models learn to Generate Multiple Objects? Paper • 2605.00273 • Published Apr 30 • 9
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models Paper • 2602.24264 • Published Feb 27 • 14
Enhancing Multi-Image Understanding through Delimiter Token Scaling Paper • 2602.01984 • Published Feb 2 • 5
DISCO: Diversifying Sample Condensation for Efficient Model Evaluation Paper • 2510.07959 • Published Oct 9, 2025 • 15
Does Data Scaling Lead to Visual Compositional Generalization? Paper • 2507.07102 • Published Jul 9, 2025 • 2