Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Xiangyu, Wu, Zhanqian, Xiong, Kaixin, Xu, Ziyang, Zhou, Lijun, Xu, Gangwei, Xu, Shaoqing, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun, Liu, Wenyu, Wang, Xinggang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913903288516608
author Guo, Xiangyu
Wu, Zhanqian
Xiong, Kaixin
Xu, Ziyang
Zhou, Lijun
Xu, Gangwei
Xu, Shaoqing
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Liu, Wenyu
Wang, Xinggang
author_facet Guo, Xiangyu
Wu, Zhanqian
Xiong, Kaixin
Xu, Ziyang
Zhou, Lijun
Xu, Gangwei
Xu, Shaoqing
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Liu, Wenyu
Wang, Xinggang
contents We present Genesis, a unified framework for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-aware LiDAR generator with NeRF-based rendering and adaptive sampling. Both modalities are directly coupled through a shared latent space, enabling coherent evolution across visual and geometric domains. To guide the generation with structured semantics, we introduce DataCrafter, a captioning module built on vision-language models that provides scene-level and instance-level supervision. Extensive experiments on the nuScenes benchmark demonstrate that Genesis achieves state-of-the-art performance across video and LiDAR metrics (FVD 16.95, FID 4.24, Chamfer 0.611), and benefits downstream tasks including segmentation and 3D detection, validating the semantic fidelity and practical utility of the generated data.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
Guo, Xiangyu
Wu, Zhanqian
Xiong, Kaixin
Xu, Ziyang
Zhou, Lijun
Xu, Gangwei
Xu, Shaoqing
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Liu, Wenyu
Wang, Xinggang
Computer Vision and Pattern Recognition
We present Genesis, a unified framework for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-aware LiDAR generator with NeRF-based rendering and adaptive sampling. Both modalities are directly coupled through a shared latent space, enabling coherent evolution across visual and geometric domains. To guide the generation with structured semantics, we introduce DataCrafter, a captioning module built on vision-language models that provides scene-level and instance-level supervision. Extensive experiments on the nuScenes benchmark demonstrate that Genesis achieves state-of-the-art performance across video and LiDAR metrics (FVD 16.95, FID 4.24, Chamfer 0.611), and benefits downstream tasks including segmentation and 3D detection, validating the semantic fidelity and practical utility of the generated data.
title Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07497