R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables Paper • 2607.02921 • Published Jul 3 • 8
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding Paper • 2608.17402 • Published 14 days ago • 17
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model Paper • 2309.16058 • Published Sep 27, 2023 • 56