What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
Abstract
World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the future must still be generated during inference is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: environmental perturbation, data efficiency, and task generalization. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from preparing the future, not generating it. We therefore propose Simple-WAM, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: https://zrporz.github.io/Simple-WAM-Web/
Community
Links
π paper: https://arxiv.org/abs/2609.34981
π project page: https://zrporz.github.io/Simple-WAM-Web/
π» code: https://github.com/LeapLabTHU/Simple-WAM
π€ model: https://huggingface.co/rpzhou/Simple-WAM
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models (2026)
- Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models (2026)
- Foresight Without Seeing: Latent Futures for World Action Models (2026)
- V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents (2026)
- World Tokens: Enhancing Embodied Policies with Training-Time World Modeling (2026)
- The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models (2026)
- An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34981 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper