X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 3 days ago • 28
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 3 days ago • 28
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 3 days ago • 28
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models Paper • 2608.22403 • Published 21 days ago • 1
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 6 days ago • 67
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning Paper • 2604.14922 • Published Apr 16 • 7
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework Paper • 2605.20373 • Published May 19
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10
A3D: Adaptive Affordance Assembly with Dual-Arm Manipulation Paper • 2601.11076 • Published Jan 16
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments Paper • 2605.18758 • Published Apr 3 • 17
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling Paper • 2604.19734 • Published Apr 21 • 34
DreamPolish: Domain Score Distillation With Progressive Geometry Generation Paper • 2411.01602 • Published Nov 3, 2024 • 11
GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Paper • 2507.01006 • Published Jul 1, 2025 • 257
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Paper • 2501.02955 • Published Jan 6, 2025 • 44