World Action Models: The Next Frontier in Embodied AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Siyin, Shi, Junhao, Fu, Zhaoyang, He, Xinzhe, Liu, Feihong, Yang, Chenchen, Zhou, Yikang, Fei, Zhaoye, Gong, Jingjing, Fu, Jinlan, Shou, Mike Zheng, Huang, Xuanjing, Qiu, Xipeng, Jiang, Yu-Gang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911674287521792
author Wang, Siyin
Shi, Junhao
Fu, Zhaoyang
He, Xinzhe
Liu, Feihong
Yang, Chenchen
Zhou, Yikang
Fei, Zhaoye
Gong, Jingjing
Fu, Jinlan
Shou, Mike Zheng
Huang, Xuanjing
Qiu, Xipeng
Jiang, Yu-Gang
author_facet Wang, Siyin
Shi, Junhao
Fu, Zhaoyang
He, Xinzhe
Liu, Feihong
Yang, Chenchen
Zhou, Yikang
Fei, Zhaoye
Gong, Jingjing
Fu, Jinlan
Shou, Mike Zheng
Huang, Xuanjing
Qiu, Xipeng
Jiang, Yu-Gang
contents Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline. We term this emerging paradigm World Action Models (WAMs): embodied foundation models that unify predictive state modeling with action generation, targeting a joint distribution over future states and actions rather than actions alone. However, the literature remains fragmented across architectures, learning objectives, and application scenarios, lacking a unified conceptual framework. We formally define WAMs and disambiguate them from related concepts, and trace the foundations and early integration of VLA and world model research that gave rise to this paradigm. We organize existing methods into a structured taxonomy of Cascaded and Joint WAMs, with further subdivision by generation modality, conditioning mechanism, and action decoding strategy. We systematically analyze the data ecosystem fueling WAMs development, spanning robot teleoperation, portable human demonstrations, simulation, and internet-scale egocentric video, and synthesize emerging evaluation protocols organized around visual fidelity, physical commonsense, and action plausibility. Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12090
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle World Action Models: The Next Frontier in Embodied AI
Wang, Siyin
Shi, Junhao
Fu, Zhaoyang
He, Xinzhe
Liu, Feihong
Yang, Chenchen
Zhou, Yikang
Fei, Zhaoye
Gong, Jingjing
Fu, Jinlan
Shou, Mike Zheng
Huang, Xuanjing
Qiu, Xipeng
Jiang, Yu-Gang
Robotics
Computation and Language
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-action mappings without explicitly modeling how the physical world evolves under intervention. A growing body of work addresses this limitation by integrating world models, predictive models of environment dynamics, into the action generation pipeline. We term this emerging paradigm World Action Models (WAMs): embodied foundation models that unify predictive state modeling with action generation, targeting a joint distribution over future states and actions rather than actions alone. However, the literature remains fragmented across architectures, learning objectives, and application scenarios, lacking a unified conceptual framework. We formally define WAMs and disambiguate them from related concepts, and trace the foundations and early integration of VLA and world model research that gave rise to this paradigm. We organize existing methods into a structured taxonomy of Cascaded and Joint WAMs, with further subdivision by generation modality, conditioning mechanism, and action decoding strategy. We systematically analyze the data ecosystem fueling WAMs development, spanning robot teleoperation, portable human demonstrations, simulation, and internet-scale egocentric video, and synthesize emerging evaluation protocols organized around visual fidelity, physical commonsense, and action plausibility. Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.
title World Action Models: The Next Frontier in Embodied AI
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12090