Generating Multimodal Driving Scenes via Next-Scene Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yanhao, Zhang, Haoyang, Lin, Tianwei, Huang, Lichao, Luo, Shujie, Wu, Rui, Qiu, Congpei, Ke, Wei, Zhang, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915213594329088
author Wu, Yanhao
Zhang, Haoyang
Lin, Tianwei
Huang, Lichao
Luo, Shujie
Wu, Rui
Qiu, Congpei
Ke, Wei
Zhang, Tong
author_facet Wu, Yanhao
Zhang, Haoyang
Lin, Tianwei
Huang, Lichao
Luo, Shujie
Wu, Rui
Qiu, Congpei
Ke, Wei
Zhang, Tong
contents Generative models in Autonomous Driving (AD) enable diverse scene creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multimodal generation framework that incorporates four major data modalities, including a novel addition of map modality. With tokenized modalities, our scene sequence generation framework autoregressively predicts each scene while managing computational demands through a two-stage approach. The Temporal AutoRegressive (TAR) component captures inter-frame dynamics for each modality while the Ordered AutoRegressive (OAR) component aligns modalities within each scene by sequentially predicting tokens in a fixed order. To maintain coherence between map and ego-action modalities, we introduce the Action-aware Map Alignment (AMA) module, which applies a transformation based on the ego-action to maintain coherence between these modalities. Our framework effectively generates complex, realistic driving scenes over extended sequences, ensuring multimodal consistency and offering fine-grained control over scene elements. Project page: https://yanhaowu.github.io/UMGen/
format Preprint
id arxiv_https___arxiv_org_abs_2503_14945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Multimodal Driving Scenes via Next-Scene Prediction
Wu, Yanhao
Zhang, Haoyang
Lin, Tianwei
Huang, Lichao
Luo, Shujie
Wu, Rui
Qiu, Congpei
Ke, Wei
Zhang, Tong
Computer Vision and Pattern Recognition
Generative models in Autonomous Driving (AD) enable diverse scene creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multimodal generation framework that incorporates four major data modalities, including a novel addition of map modality. With tokenized modalities, our scene sequence generation framework autoregressively predicts each scene while managing computational demands through a two-stage approach. The Temporal AutoRegressive (TAR) component captures inter-frame dynamics for each modality while the Ordered AutoRegressive (OAR) component aligns modalities within each scene by sequentially predicting tokens in a fixed order. To maintain coherence between map and ego-action modalities, we introduce the Action-aware Map Alignment (AMA) module, which applies a transformation based on the ego-action to maintain coherence between these modalities. Our framework effectively generates complex, realistic driving scenes over extended sequences, ensuring multimodal consistency and offering fine-grained control over scene elements. Project page: https://yanhaowu.github.io/UMGen/
title Generating Multimodal Driving Scenes via Next-Scene Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.14945