World Guidance: World Modeling in Condition Space for Action Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Su, Yue, Chen, Sijin, Shi, Haixin, Liu, Mingyu, Zhang, Zhengshen, Huang, Ningyuan, Zhong, Weiheng, Zhu, Zhengbang, Liu, Yuxiao, Liu, Xihui
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914350652981248
author Su, Yue
Chen, Sijin
Shi, Haixin
Liu, Mingyu
Zhang, Zhengshen
Huang, Ningyuan
Zhong, Weiheng
Zhu, Zhengbang
Liu, Yuxiao
Liu, Xihui
author_facet Su, Yue
Chen, Sijin
Shi, Haixin
Liu, Mingyu
Zhang, Zhengshen
Huang, Ningyuan
Zhong, Weiheng
Zhu, Zhengbang
Liu, Yuxiao
Liu, Xihui
contents Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. The VLA is then trained to simultaneously predict these compressed conditions alongside future actions, thereby achieving effective world modeling within the condition space for action inference. We demonstrate that modeling and predicting this condition space not only facilitates fine-grained action generation but also exhibits superior generalization capabilities. Moreover, it learns effectively from substantial human manipulation videos. Extensive experiments across both simulation and real-world environments validate that our method significantly outperforms existing methods based on future prediction. Project page is available at: https://selen-suyue.github.io/WoGNet/
format Preprint
id arxiv_https___arxiv_org_abs_2602_22010
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle World Guidance: World Modeling in Condition Space for Action Generation
Su, Yue
Chen, Sijin
Shi, Haixin
Liu, Mingyu
Zhang, Zhengshen
Huang, Ningyuan
Zhong, Weiheng
Zhu, Zhengbang
Liu, Yuxiao
Liu, Xihui
Robotics
Computer Vision and Pattern Recognition
Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. The VLA is then trained to simultaneously predict these compressed conditions alongside future actions, thereby achieving effective world modeling within the condition space for action inference. We demonstrate that modeling and predicting this condition space not only facilitates fine-grained action generation but also exhibits superior generalization capabilities. Moreover, it learns effectively from substantial human manipulation videos. Extensive experiments across both simulation and real-world environments validate that our method significantly outperforms existing methods based on future prediction. Project page is available at: https://selen-suyue.github.io/WoGNet/
title World Guidance: World Modeling in Condition Space for Action Generation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.22010