X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Chaoda, Li, Sean, Deng, Jinhao, Wang, Zhennan, Chen, Shijia, Xiao, Liqiang, Chi, Ziheng, Lin, Hongbin, Chen, Kangjie, Wang, Boyang, Zhang, Yu, Liu, Xianming
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917371254407168
author Zheng, Chaoda
Li, Sean
Deng, Jinhao
Wang, Zhennan
Chen, Shijia
Xiao, Liqiang
Chi, Ziheng
Lin, Hongbin
Chen, Kangjie
Wang, Boyang
Zhang, Yu
Liu, Xianming
author_facet Zheng, Chaoda
Li, Sean
Deng, Jinhao
Wang, Zhennan
Chen, Shijia
Xiao, Liqiang
Chi, Ziheng
Lin, Hongbin
Chen, Kangjie
Wang, Boyang
Zhang, Yu
Liu, Xianming
contents Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realistic future observations under proposed actions, while remaining controllable and stable over long horizons. We present X-World, an action-conditioned multi-camera generative world model that simulates future observations directly in video space. Given synchronized multi-view camera history and a future action sequence, X-World generates future multi-camera video streams that follow the commanded actions. To ensure reproducible and editable scene rollouts, X-World further supports optional controls over dynamic traffic agents and static road elements, and retains a text-prompt interface for appearance-level control (e.g., weather and time of day). Beyond world simulation, X-World also enables video style transfer by conditioning on appearance prompts while preserving the underlying action and scene dynamics. At the core of X-World is a multi-view latent video generator designed to explicitly encourage cross-view geometric consistency and temporal coherence under diverse control signals. Experiments show that X-World achieves high-quality multi-view video generation with (i) strong view consistency across cameras, (ii) stable temporal dynamics over long rollouts, and (iii) high controllability with strict action following and faithful adherence to optional scene controls. These properties make X-World a practical foundation for scalable and reproducible evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19979
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
Zheng, Chaoda
Li, Sean
Deng, Jinhao
Wang, Zhennan
Chen, Shijia
Xiao, Liqiang
Chi, Ziheng
Lin, Hongbin
Chen, Kangjie
Wang, Boyang
Zhang, Yu
Liu, Xianming
Computer Vision and Pattern Recognition
Artificial Intelligence
Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realistic future observations under proposed actions, while remaining controllable and stable over long horizons. We present X-World, an action-conditioned multi-camera generative world model that simulates future observations directly in video space. Given synchronized multi-view camera history and a future action sequence, X-World generates future multi-camera video streams that follow the commanded actions. To ensure reproducible and editable scene rollouts, X-World further supports optional controls over dynamic traffic agents and static road elements, and retains a text-prompt interface for appearance-level control (e.g., weather and time of day). Beyond world simulation, X-World also enables video style transfer by conditioning on appearance prompts while preserving the underlying action and scene dynamics. At the core of X-World is a multi-view latent video generator designed to explicitly encourage cross-view geometric consistency and temporal coherence under diverse control signals. Experiments show that X-World achieves high-quality multi-view video generation with (i) strong view consistency across cameras, (ii) stable temporal dynamics over long rollouts, and (iii) high controllability with strict action following and faithful adherence to optional scene controls. These properties make X-World a practical foundation for scalable and reproducible evaluation.
title X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.19979