MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Ruiyuan, Chen, Kai, Xiao, Bo, Hong, Lanqing, Li, Zhenguo, Xu, Qiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912500482572288
author Gao, Ruiyuan
Chen, Kai
Xiao, Bo
Hong, Lanqing
Li, Zhenguo
Xu, Qiang
author_facet Gao, Ruiyuan
Chen, Kai
Xiao, Bo
Hong, Lanqing
Li, Zhenguo
Xu, Qiang
contents The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework for video generation, it introduces challenges in controllable driving video generation, especially for geometry control, rendering existing control methods ineffective. To address these issues, we propose MagicDrive-V2, a novel approach that integrates the MVDiT block and spatial-temporal conditional encoding to enable multi-view video generation and precise geometric control. Additionally, we introduce an efficient method for obtaining contextual descriptions for videos to support diverse textual control, along with a progressive training strategy using mixed video data to enhance training efficiency and generalizability. Consequently, MagicDrive-V2 enables multi-view driving video synthesis with $3.3\times$ resolution and $4\times$ frame count (compared to current SOTA), rich contextual control, and geometric controls. Extensive experiments demonstrate MagicDrive-V2's ability, unlocking broader applications in autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13807
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
Gao, Ruiyuan
Chen, Kai
Xiao, Bo
Hong, Lanqing
Li, Zhenguo
Xu, Qiang
Computer Vision and Pattern Recognition
The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework for video generation, it introduces challenges in controllable driving video generation, especially for geometry control, rendering existing control methods ineffective. To address these issues, we propose MagicDrive-V2, a novel approach that integrates the MVDiT block and spatial-temporal conditional encoding to enable multi-view video generation and precise geometric control. Additionally, we introduce an efficient method for obtaining contextual descriptions for videos to support diverse textual control, along with a progressive training strategy using mixed video data to enhance training efficiency and generalizability. Consequently, MagicDrive-V2 enables multi-view driving video synthesis with $3.3\times$ resolution and $4\times$ frame count (compared to current SOTA), rich contextual control, and geometric controls. Extensive experiments demonstrate MagicDrive-V2's ability, unlocking broader applications in autonomous driving.
title MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13807