Saved in:
Bibliographic Details
Main Authors: Li, Lin, Zhang, Qihang, Luo, Yiming, Yang, Shuai, Wang, Ruilin, Han, Fei, Yu, Mingrui, Gao, Zelin, Xue, Nan, Zhu, Xing, Shen, Yujun, Xu, Yinghao
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.21998
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911533661945856
author Li, Lin
Zhang, Qihang
Luo, Yiming
Yang, Shuai
Wang, Ruilin
Han, Fei
Yu, Mingrui
Gao, Zelin
Xue, Nan
Zhu, Xing
Shen, Yujun
Xu, Yinghao
author_facet Li, Lin
Zhang, Qihang
Luo, Yiming
Yang, Shuai
Wang, Ruilin
Han, Fei
Yu, Mingrui
Gao, Zelin
Xue, Nan
Zhu, Xing
Shen, Yujun
Xu, Yinghao
contents This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by understanding the causality between actions and visual dynamics. Inspired by this, we introduce LingBot-VA, an autoregressive diffusion framework that learns frame prediction and policy execution simultaneously. Our model features three carefully crafted designs: (1) a shared latent space, integrating vision and action tokens, driven by a Mixture-of-Transformers (MoT) architecture, (2) a closed-loop rollout mechanism, allowing for ongoing acquisition of environmental feedback with ground-truth observations, (3) an asynchronous inference pipeline, parallelizing action prediction and motor execution to support efficient control. We evaluate our model on both simulation benchmarks and real-world scenarios, where it shows significant promise in long-horizon manipulation, data efficiency in post-training, and strong generalizability to novel configurations. The code and model are made publicly available to facilitate the community.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21998
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Causal World Modeling for Robot Control
Li, Lin
Zhang, Qihang
Luo, Yiming
Yang, Shuai
Wang, Ruilin
Han, Fei
Yu, Mingrui
Gao, Zelin
Xue, Nan
Zhu, Xing
Shen, Yujun
Xu, Yinghao
Computer Vision and Pattern Recognition
Robotics
This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by understanding the causality between actions and visual dynamics. Inspired by this, we introduce LingBot-VA, an autoregressive diffusion framework that learns frame prediction and policy execution simultaneously. Our model features three carefully crafted designs: (1) a shared latent space, integrating vision and action tokens, driven by a Mixture-of-Transformers (MoT) architecture, (2) a closed-loop rollout mechanism, allowing for ongoing acquisition of environmental feedback with ground-truth observations, (3) an asynchronous inference pipeline, parallelizing action prediction and motor execution to support efficient control. We evaluate our model on both simulation benchmarks and real-world scenarios, where it shows significant promise in long-horizon manipulation, data efficiency in post-training, and strong generalizability to novel configurations. The code and model are made publicly available to facilitate the community.
title Causal World Modeling for Robot Control
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2601.21998