TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published 17 days ago • 58
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published 21 days ago • 150
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published Aug 1 • 17
FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Paper • 2608.04530 • Published Aug 5 • 14
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published Aug 5 • 40
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models Paper • 2608.04964 • Published Aug 5 • 13
OPD-V: Visual On-Policy Self-Distillation with Modality Balance Paper • 2608.05131 • Published Aug 6 • 14
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Paper • 2608.05042 • Published Aug 5 • 8
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 105
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 140
DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model Paper • 2606.30292 • Published Jun 29 • 16
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 217
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models Paper • 2512.24618 • Published Dec 31, 2025 • 156
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times Paper • 2512.16093 • Published Dec 18, 2025 • 96