MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Ruiyuan, Chen, Kai, Li, Zhihao, Hong, Lanqing, Li, Zhenguo, Xu, Qiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913959346438144
author Gao, Ruiyuan
Chen, Kai
Li, Zhihao
Hong, Lanqing
Li, Zhenguo
Xu, Qiang
author_facet Gao, Ruiyuan
Chen, Kai
Li, Zhihao
Hong, Lanqing
Li, Zhenguo
Xu, Qiang
contents Controllable generative models for images and videos have seen significant success, yet 3D scene generation, especially in unbounded scenarios like autonomous driving, remains underdeveloped. Existing methods lack flexible controllability and often rely on dense view data collection in controlled environments, limiting their generalizability across common datasets (e.g., nuScenes). In this paper, we introduce MagicDrive3D, a novel framework for controllable 3D street scene generation that combines video-based view synthesis with 3D representation (3DGS) generation. It supports multi-condition control, including road maps, 3D objects, and text descriptions. Unlike previous approaches that require 3D representation before training, MagicDrive3D first trains a multi-view video generation model to synthesize diverse street views. This method utilizes routinely collected autonomous driving data, reducing data acquisition challenges and enriching 3D scene generation. In the 3DGS generation step, we introduce Fault-Tolerant Gaussian Splatting to address minor errors and use monocular depth for better initialization, alongside appearance modeling to manage exposure discrepancies across viewpoints. Experiments show that MagicDrive3D generates diverse, high-quality 3D driving scenes, supports any-view rendering, and enhances downstream tasks like BEV segmentation, demonstrating its potential for autonomous driving simulation and beyond.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14475
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
Gao, Ruiyuan
Chen, Kai
Li, Zhihao
Hong, Lanqing
Li, Zhenguo
Xu, Qiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Controllable generative models for images and videos have seen significant success, yet 3D scene generation, especially in unbounded scenarios like autonomous driving, remains underdeveloped. Existing methods lack flexible controllability and often rely on dense view data collection in controlled environments, limiting their generalizability across common datasets (e.g., nuScenes). In this paper, we introduce MagicDrive3D, a novel framework for controllable 3D street scene generation that combines video-based view synthesis with 3D representation (3DGS) generation. It supports multi-condition control, including road maps, 3D objects, and text descriptions. Unlike previous approaches that require 3D representation before training, MagicDrive3D first trains a multi-view video generation model to synthesize diverse street views. This method utilizes routinely collected autonomous driving data, reducing data acquisition challenges and enriching 3D scene generation. In the 3DGS generation step, we introduce Fault-Tolerant Gaussian Splatting to address minor errors and use monocular depth for better initialization, alongside appearance modeling to manage exposure discrepancies across viewpoints. Experiments show that MagicDrive3D generates diverse, high-quality 3D driving scenes, supports any-view rendering, and enhances downstream tasks like BEV segmentation, demonstrating its potential for autonomous driving simulation and beyond.
title MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.14475