SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dai, Peng, Tan, Feitong, Xu, Qiangeng, Futschik, David, Du, Ruofei, Fanello, Sean, Qi, Xiaojuan, Zhang, Yinda
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929404185149440
author Dai, Peng
Tan, Feitong
Xu, Qiangeng
Futschik, David
Du, Ruofei
Fanello, Sean
Qi, Xiaojuan
Zhang, Yinda
author_facet Dai, Peng
Tan, Feitong
Xu, Qiangeng
Futschik, David
Du, Ruofei
Fanello, Sean
Qi, Xiaojuan
Zhang, Yinda
contents Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating 3D stereoscopic videos using an off-the-shelf monocular video generation model. Our method warps a generated monocular video into camera views on stereoscopic baseline using estimated video depth, and employs a novel frame matrix video inpainting framework. The framework leverages the video generation model to inpaint frames observed from different timestamps and views. This effective approach generates consistent and semantically coherent stereoscopic videos without scene optimization or model fine-tuning. Moreover, we develop a disocclusion boundary re-injection scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, including Sora [4 ], Lumiere [2], WALT [8 ], and Zeroscope [ 42]. The experiments demonstrate that our method has a significant improvement over previous methods. The code will be released at \url{https://daipengwa.github.io/SVG_ProjectPage}.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00367
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix
Dai, Peng
Tan, Feitong
Xu, Qiangeng
Futschik, David
Du, Ruofei
Fanello, Sean
Qi, Xiaojuan
Zhang, Yinda
Computer Vision and Pattern Recognition
Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating 3D stereoscopic videos using an off-the-shelf monocular video generation model. Our method warps a generated monocular video into camera views on stereoscopic baseline using estimated video depth, and employs a novel frame matrix video inpainting framework. The framework leverages the video generation model to inpaint frames observed from different timestamps and views. This effective approach generates consistent and semantically coherent stereoscopic videos without scene optimization or model fine-tuning. Moreover, we develop a disocclusion boundary re-injection scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, including Sora [4 ], Lumiere [2], WALT [8 ], and Zeroscope [ 42]. The experiments demonstrate that our method has a significant improvement over previous methods. The code will be released at \url{https://daipengwa.github.io/SVG_ProjectPage}.
title SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.00367