World-consistent Video Diffusion with Explicit 3D Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Qihang, Zhai, Shuangfei, Bautista, Miguel Angel, Miao, Kevin, Toshev, Alexander, Susskind, Joshua, Gu, Jiatao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909412378017792
author Zhang, Qihang
Zhai, Shuangfei
Bautista, Miguel Angel
Miao, Kevin
Toshev, Alexander
Susskind, Joshua
Gu, Jiatao
author_facet Zhang, Qihang
Zhai, Shuangfei
Bautista, Miguel Angel
Miao, Kevin
Toshev, Alexander
Susskind, Joshua
Gu, Jiatao
contents Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propose World-consistent Video Diffusion (WVD), a novel framework that incorporates explicit 3D supervision using XYZ images, which encode global 3D coordinates for each image pixel. More specifically, we train a diffusion transformer to learn the joint distribution of RGB and XYZ frames. This approach supports multi-task adaptability via a flexible inpainting strategy. For example, WVD can estimate XYZ frames from ground-truth RGB or generate novel RGB frames using XYZ projections along a specified camera trajectory. In doing so, WVD unifies tasks like single-image-to-3D generation, multi-view stereo, and camera-controlled video generation. Our approach demonstrates competitive performance across multiple benchmarks, providing a scalable solution for 3D-consistent video and image generation with a single pretrained model.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle World-consistent Video Diffusion with Explicit 3D Modeling
Zhang, Qihang
Zhai, Shuangfei
Bautista, Miguel Angel
Miao, Kevin
Toshev, Alexander
Susskind, Joshua
Gu, Jiatao
Computer Vision and Pattern Recognition
Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propose World-consistent Video Diffusion (WVD), a novel framework that incorporates explicit 3D supervision using XYZ images, which encode global 3D coordinates for each image pixel. More specifically, we train a diffusion transformer to learn the joint distribution of RGB and XYZ frames. This approach supports multi-task adaptability via a flexible inpainting strategy. For example, WVD can estimate XYZ frames from ground-truth RGB or generate novel RGB frames using XYZ projections along a specified camera trajectory. In doing so, WVD unifies tasks like single-image-to-3D generation, multi-view stereo, and camera-controlled video generation. Our approach demonstrates competitive performance across multiple benchmarks, providing a scalable solution for 3D-consistent video and image generation with a single pretrained model.
title World-consistent Video Diffusion with Explicit 3D Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.01821