UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Guosheng, Wang, Yaozeng, Wang, Xiaofeng, Zhu, Zheng, Yu, Tingdong, Huang, Guan, Zai, Yongchen, Jiao, Ji, Xue, Changliang, Wang, Xiaole, Yang, Zhen, Zhu, Futang, Wang, Xingang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912868371267584
author Zhao, Guosheng
Wang, Yaozeng
Wang, Xiaofeng
Zhu, Zheng
Yu, Tingdong
Huang, Guan
Zai, Yongchen
Jiao, Ji
Xue, Changliang
Wang, Xiaole
Yang, Zhen
Zhu, Futang
Wang, Xingang
author_facet Zhao, Guosheng
Wang, Yaozeng
Wang, Xiaofeng
Zhu, Zheng
Yu, Tingdong
Huang, Guan
Zai, Yongchen
Jiao, Ji
Xue, Changliang
Wang, Xiaole
Yang, Zhen
Zhu, Futang
Wang, Xingang
contents World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera video or LiDAR sequence synthesis. In this paper, we propose UniDriveDreamer, a single-stage unified multimodal world model for autonomous driving, which directly generates multimodal future observations without relying on intermediate representations or cascaded modules. Our framework introduces a LiDAR-specific variational autoencoder (VAE) designed to encode input LiDAR sequences, alongside a video VAE for multi-camera images. To ensure cross-modal compatibility and training stability, we propose Unified Latent Anchoring (ULA), which explicitly aligns the latent distributions of the two modalities. The aligned features are fused and processed by a diffusion transformer that jointly models their geometric correspondence and temporal evolution. Additionally, structured scene layout information is projected per modality as a conditioning signal to guide the synthesis. Extensive experiments demonstrate that UniDriveDreamer outperforms previous state-of-the-art methods in both video and LiDAR generation, while also yielding measurable improvements in downstream
format Preprint
id arxiv_https___arxiv_org_abs_2602_02002
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
Zhao, Guosheng
Wang, Yaozeng
Wang, Xiaofeng
Zhu, Zheng
Yu, Tingdong
Huang, Guan
Zai, Yongchen
Jiao, Ji
Xue, Changliang
Wang, Xiaole
Yang, Zhen
Zhu, Futang
Wang, Xingang
Computer Vision and Pattern Recognition
World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera video or LiDAR sequence synthesis. In this paper, we propose UniDriveDreamer, a single-stage unified multimodal world model for autonomous driving, which directly generates multimodal future observations without relying on intermediate representations or cascaded modules. Our framework introduces a LiDAR-specific variational autoencoder (VAE) designed to encode input LiDAR sequences, alongside a video VAE for multi-camera images. To ensure cross-modal compatibility and training stability, we propose Unified Latent Anchoring (ULA), which explicitly aligns the latent distributions of the two modalities. The aligned features are fused and processed by a diffusion transformer that jointly models their geometric correspondence and temporal evolution. Additionally, structured scene layout information is projected per modality as a conditioning signal to guide the synthesis. Extensive experiments demonstrate that UniDriveDreamer outperforms previous state-of-the-art methods in both video and LiDAR generation, while also yielding measurable improvements in downstream
title UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02002