X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Baolu, Qian, Jingyu, Guo, Rui, Chen, Yilun, Liu, Hanpeng, Lin, Yuan, Zhou, Junhong, Liu, Ruixin, Yang, Willow, Zheng, Yutong, Zhang, Zhenli, Tenglong, Gu, Ding, Zhuangzhuang, Zheng, Pengkun, Zhang, Yu, Liu, Xianming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910272809074688
author Li, Baolu
Qian, Jingyu
Guo, Rui
Chen, Yilun
Liu, Hanpeng
Lin, Yuan
Zhou, Junhong
Liu, Ruixin
Yang, Willow
Zheng, Yutong
Zhang, Zhenli
Tenglong
Gu
Ding, Zhuangzhuang
Zheng, Pengkun
Zhang, Yu
Liu, Xianming
author_facet Li, Baolu
Qian, Jingyu
Guo, Rui
Chen, Yilun
Liu, Hanpeng
Lin, Yuan
Zhou, Junhong
Liu, Ruixin
Yang, Willow
Zheng, Yutong
Zhang, Zhenli
Tenglong
Gu
Ding, Zhuangzhuang
Zheng, Pengkun
Zhang, Yu
Liu, Xianming
contents Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internalize physical dynamics and long-term causality by predicting future video from past observations. However, naive next-frame prediction faces two challenges: 1) unlike semantically distinct text tokens, video tokens are low-entropy and redundant, causing prediction to degenerate into trivial extrapolation. 2) world modeling poses a temporal dilemma: dense prediction captures instantaneous dynamics, but cannot efficiently model long-horizon causality. To learn world knowledge effectively, we introduce X-Foresight, a predictive world model integrated directly into the VLA architecture to jointly learn world modeling and real-time action control. At its core lies a long-horizon chunk-wise auto-regressive strategy that addresses both challenges: by predicting semantically distant chunks rather than adjacent frames, it escapes trivial extrapolation, while preserving dense intra-chunk frames for instantaneous dynamics and sparse inter-chunk transitions for long-term causality. A curriculum learning schedule progressively extends prediction horizons and stabilizes long-horizon training. To capture long-term causality effectively, we present temporal importance sampling, which concentrates supervision on safety-critical chunks identified by ego-motion and behavioral signals. We further delegate photorealistic synthesis to a diffusion-based multi-view renderer, improving photorealistic appearance. Comprehensive experiments demonstrate that X-Foresight significantly outperforms VLA baselines in planning performance while maintaining strong generative fidelity, establishing a robust paradigm for world-knowledge-driven autonomous systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24892
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
Li, Baolu
Qian, Jingyu
Guo, Rui
Chen, Yilun
Liu, Hanpeng
Lin, Yuan
Zhou, Junhong
Liu, Ruixin
Yang, Willow
Zheng, Yutong
Zhang, Zhenli
Tenglong
Gu
Ding, Zhuangzhuang
Zheng, Pengkun
Zhang, Yu
Liu, Xianming
Computer Vision and Pattern Recognition
Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internalize physical dynamics and long-term causality by predicting future video from past observations. However, naive next-frame prediction faces two challenges: 1) unlike semantically distinct text tokens, video tokens are low-entropy and redundant, causing prediction to degenerate into trivial extrapolation. 2) world modeling poses a temporal dilemma: dense prediction captures instantaneous dynamics, but cannot efficiently model long-horizon causality. To learn world knowledge effectively, we introduce X-Foresight, a predictive world model integrated directly into the VLA architecture to jointly learn world modeling and real-time action control. At its core lies a long-horizon chunk-wise auto-regressive strategy that addresses both challenges: by predicting semantically distant chunks rather than adjacent frames, it escapes trivial extrapolation, while preserving dense intra-chunk frames for instantaneous dynamics and sparse inter-chunk transitions for long-term causality. A curriculum learning schedule progressively extends prediction horizons and stabilizes long-horizon training. To capture long-term causality effectively, we present temporal importance sampling, which concentrates supervision on safety-critical chunks identified by ego-motion and behavioral signals. We further delegate photorealistic synthesis to a diffusion-based multi-view renderer, improving photorealistic appearance. Comprehensive experiments demonstrate that X-Foresight significantly outperforms VLA baselines in planning performance while maintaining strong generative fidelity, establishing a robust paradigm for world-knowledge-driven autonomous systems.
title X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.24892