SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yiming, Yao, Chun-Han, Voleti, Vikram, Jiang, Huaizu, Jampani, Varun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915176116125696
author Xie, Yiming
Yao, Chun-Han
Voleti, Vikram
Jiang, Huaizu
Jampani, Varun
author_facet Xie, Yiming
Yao, Chun-Han
Voleti, Vikram
Jiang, Huaizu
Jampani, Varun
contents We present Stable Video 4D (SV4D), a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17470
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
Xie, Yiming
Yao, Chun-Han
Voleti, Vikram
Jiang, Huaizu
Jampani, Varun
Computer Vision and Pattern Recognition
We present Stable Video 4D (SV4D), a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works.
title SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.17470