Saved in:
Bibliographic Details
Main Authors: Ji, Yishen, Zhu, Ziyue, Zhu, Zhenxin, Xiong, Kaixin, Lu, Ming, Li, Zhiqi, Zhou, Lijun, Sun, Haiyang, Wang, Bing, Lu, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.22231
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Recent progress in driving video generation has shown significant potential for enhancing self-driving systems by providing scalable and controllable training data. Although pretrained state-of-the-art generation models, guided by 2D layout conditions (e.g., HD maps and bounding boxes), can produce photorealistic driving videos, achieving controllable multi-view videos with high 3D consistency remains a major challenge. To tackle this, we introduce a novel spatial adaptive generation framework, CoGen, which leverages advances in 3D generation to improve performance in two key aspects: (i) To ensure 3D consistency, we first generate high-quality, controllable 3D conditions that capture the geometry of driving scenes. By replacing coarse 2D conditions with these fine-grained 3D representations, our approach significantly enhances the spatial consistency of the generated videos. (ii) Additionally, we introduce a consistency adapter module to strengthen the robustness of the model to multi-condition control. The results demonstrate that this method excels in preserving geometric fidelity and visual realism, offering a reliable video generation solution for autonomous driving.