X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Baolu, Qian, Jingyu, Guo, Rui, Chen, Yilun, Liu, Hanpeng, Lin, Yuan, Zhou, Junhong, Liu, Ruixin, Yang, Willow, Zheng, Yutong, Zhang, Zhenli, Tenglong, Gu, Ding, Zhuangzhuang, Zheng, Pengkun, Zhang, Yu, Liu, Xianming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
You Only Need Less Attention at Each Stage in Vision Transformers
by: Zhang, Shuoxi, et al.
Published: (2024)
by: Zhang, Shuoxi, et al.
Published: (2024)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026)
by: Zheng, Yupeng, et al.
Published: (2026)
Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
by: Zhang, Yin, et al.
Published: (2025)
by: Zhang, Yin, et al.
Published: (2025)
AstraNav-World: World Model for Foresight Control and Consistency
by: Chen, Jintao, et al.
Published: (2025)
by: Chen, Jintao, et al.
Published: (2025)
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026)
by: Park, Sangwu, et al.
Published: (2026)
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
by: Lin, Minghui, et al.
Published: (2025)
by: Lin, Minghui, et al.
Published: (2025)
Privileged Foresight Distillation: Zero-Cost Future Correction for World Action Models
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
Current Agents Fail to Leverage World Model as Tool for Foresight
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
Vehicles Swarm Intelligence: Cooperation in both Longitudinal and Lateral Dimensions
by: Hu, Jia, et al.
Published: (2024)
by: Hu, Jia, et al.
Published: (2024)
Joker: Joint Optimization Framework for Lightweight Kernel Machines
by: Zhang, Junhong, et al.
Published: (2025)
by: Zhang, Junhong, et al.
Published: (2025)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
Spatiotemporal Causal Decoupling Model for Air Quality Forecasting
by: Ma, Jiaming, et al.
Published: (2025)
by: Ma, Jiaming, et al.
Published: (2025)
On Active Privacy Auditing in Supervised Fine-tuning for White-Box Language Models
by: Sun, Qian, et al.
Published: (2024)
by: Sun, Qian, et al.
Published: (2024)
Some functors preserving exceptionality
by: Liu, Dajun, et al.
Published: (2025)
by: Liu, Dajun, et al.
Published: (2025)
A characterization of IE-closed subcategories via canonical twin support $τ$-tilting modules
by: Gao, Hanpeng, et al.
Published: (2026)
by: Gao, Hanpeng, et al.
Published: (2026)
Forecasting Supply Chain Disruptions with Foresight Learning
by: Turtel, Benjamin, et al.
Published: (2026)
by: Turtel, Benjamin, et al.
Published: (2026)
Learning Causal Structure Distributions for Robust Planning
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2025)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2025)
Action Recognition in Real-World Ambient Assisted Living Environment
by: Zakka, Vincent Gbouna, et al.
Published: (2025)
by: Zakka, Vincent Gbouna, et al.
Published: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry
by: Liu, Yuejiang, et al.
Published: (2026)
by: Liu, Yuejiang, et al.
Published: (2026)
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
by: Zhang, Chuye, et al.
Published: (2025)
by: Zhang, Chuye, et al.
Published: (2025)
Projectively Wakamatsu Tilting Modules over One-Point Extensions
by: Liu, Dajun, et al.
Published: (2026)
by: Liu, Dajun, et al.
Published: (2026)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
by: Xu, Xiaoxu, et al.
Published: (2026)
by: Xu, Xiaoxu, et al.
Published: (2026)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
by: Zeng, Yixiao, et al.
Published: (2026)
by: Zeng, Yixiao, et al.
Published: (2026)
Vision Language Models Are Not (Yet) Spelling Correctors
by: Liang, Junhong, et al.
Published: (2025)
by: Liang, Junhong, et al.
Published: (2025)
Large Language Models for Medical Forecasting -- Foresight 2
by: Kraljevic, Zeljko, et al.
Published: (2024)
by: Kraljevic, Zeljko, et al.
Published: (2024)
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
by: Yan, Haodong, et al.
Published: (2026)
by: Yan, Haodong, et al.
Published: (2026)
Expressive Speech-driven Facial Animation with controllable emotions
by: Chen, Yutong, et al.
Published: (2023)
by: Chen, Yutong, et al.
Published: (2023)
Weberite Na$_2$MM'F$_7$ (M,M'=Redox-Active Metal) as Promising Fluoride-Based Sodium-Ion Battery Cathodes
by: Lu, Tenglong, et al.
Published: (2023)
by: Lu, Tenglong, et al.
Published: (2023)
Causal Path Alignment: Anchoring the Optimization Trajectory for Controllable In-Parameter Knowledge Editing
by: Liu, Xiyu, et al.
Published: (2025)
by: Liu, Xiyu, et al.
Published: (2025)
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
by: Ye, Angen, et al.
Published: (2026)
by: Ye, Angen, et al.
Published: (2026)
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
by: Zhai, Shaopeng, et al.
Published: (2025)
by: Zhai, Shaopeng, et al.
Published: (2025)
REPS: Reconstruction-based Point Cloud Sampling
by: Zhang, Guoqing, et al.
Published: (2024)
by: Zhang, Guoqing, et al.
Published: (2024)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
Similar Items
-
You Only Need Less Attention at Each Stage in Vision Transformers
by: Zhang, Shuoxi, et al.
Published: (2024) -
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
by: Yang, Yi, et al.
Published: (2025) -
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
by: Liu, Hanpeng, et al.
Published: (2026) -
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026) -
Causal-Tune: Mining Causal Factors from Vision Foundation Models for Domain Generalized Semantic Segmentation
by: Zhang, Yin, et al.
Published: (2025)