Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Qiqi, Xu, Huan, Li, Jingyu, Sun, Bin, Hao, Zhihui, She, Dangen, Zhu, Xiatian, Zhang, Li
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911550333255680
author Liu, Qiqi
Xu, Huan
Li, Jingyu
Sun, Bin
Hao, Zhihui
She, Dangen
Zhu, Xiatian
Zhang, Li
author_facet Liu, Qiqi
Xu, Huan
Li, Jingyu
Sun, Bin
Hao, Zhihui
She, Dangen
Zhu, Xiatian
Zhang, Li
contents Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes first and plan afterwards, resulting in open-loop imagination that may drift from the actual decision process. In this paper, we present Uni-World VLA, a unified vision-language-action (VLA) model that tightly interleaves future frame prediction and trajectory planning. Instead of generating a full world rollout before planning, our model alternates between predicting future frames and ego actions step by step, allowing planning decisions to be continuously conditioned on the imagined future observations. This interleaved generation forms a closed-loop interaction between world modeling and control, enabling more adaptive decision-making in dynamic traffic scenarios. In addition, we incorporate monocular depth information into frames to provide stronger geometric cues for world modeling, improving long-horizon scene prediction. Experiments on the NAVSIM benchmark show that our approach achieves competitive closed-loop planning performance while producing high-fidelity future frame predictions. These results demonstrate that tightly coupling world prediction and planning is a promising direction for scalable VLA driving systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27287
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving
Liu, Qiqi
Xu, Huan
Li, Jingyu
Sun, Bin
Hao, Zhihui
She, Dangen
Zhu, Xiatian
Zhang, Li
Robotics
Computer Vision and Pattern Recognition
Autonomous driving requires reasoning about how the environment evolves and planning actions accordingly. Existing world-model-based approaches typically predict future scenes first and plan afterwards, resulting in open-loop imagination that may drift from the actual decision process. In this paper, we present Uni-World VLA, a unified vision-language-action (VLA) model that tightly interleaves future frame prediction and trajectory planning. Instead of generating a full world rollout before planning, our model alternates between predicting future frames and ego actions step by step, allowing planning decisions to be continuously conditioned on the imagined future observations. This interleaved generation forms a closed-loop interaction between world modeling and control, enabling more adaptive decision-making in dynamic traffic scenarios. In addition, we incorporate monocular depth information into frames to provide stronger geometric cues for world modeling, improving long-horizon scene prediction. Experiments on the NAVSIM benchmark show that our approach achieves competitive closed-loop planning performance while producing high-fidelity future frame predictions. These results demonstrate that tightly coupling world prediction and planning is a promising direction for scalable VLA driving systems.
title Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27287