S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Peng, Tan, Feitong, Xu, Qiangeng, Huang, Yihua, Futschik, David, Du, Ruofei, Fanello, Sean, Zhang, Yinda, Qi, Xiaojuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913984629702656
author Dai, Peng
Tan, Feitong
Xu, Qiangeng
Huang, Yihua
Futschik, David
Du, Ruofei
Fanello, Sean
Zhang, Yinda
Qi, Xiaojuan
author_facet Dai, Peng
Tan, Feitong
Xu, Qiangeng
Huang, Yihua
Futschik, David
Du, Ruofei
Fanello, Sean
Zhang, Yinda
Qi, Xiaojuan
contents While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored challenge. We present a pose-free and training-free method that leverages an off-the-shelf monocular video generation model to produce immersive 3D videos. Our approach first warps the generated monocular video into pre-defined camera viewpoints using estimated depth information, then applies a novel \textit{frame matrix} inpainting framework. This framework utilizes the original video generation model to synthesize missing content across different viewpoints and timestamps, ensuring spatial and temporal consistency without requiring additional model fine-tuning. Moreover, we develop a \dualupdate~scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. The resulting multi-view videos are then adapted into stereoscopic pairs or optimized into 4D Gaussians for spatial video synthesis. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, such as Sora, Lumiere, WALT, and Zeroscope. The experiments demonstrate that our method has a significant improvement over previous methods. Project page at: https://daipengwa.github.io/S-2VG_ProjectPage/
format Preprint
id arxiv_https___arxiv_org_abs_2508_08048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
Dai, Peng
Tan, Feitong
Xu, Qiangeng
Huang, Yihua
Futschik, David
Du, Ruofei
Fanello, Sean
Zhang, Yinda
Qi, Xiaojuan
Computer Vision and Pattern Recognition
While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored challenge. We present a pose-free and training-free method that leverages an off-the-shelf monocular video generation model to produce immersive 3D videos. Our approach first warps the generated monocular video into pre-defined camera viewpoints using estimated depth information, then applies a novel \textit{frame matrix} inpainting framework. This framework utilizes the original video generation model to synthesize missing content across different viewpoints and timestamps, ensuring spatial and temporal consistency without requiring additional model fine-tuning. Moreover, we develop a \dualupdate~scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. The resulting multi-view videos are then adapted into stereoscopic pairs or optimized into 4D Gaussians for spatial video synthesis. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, such as Sora, Lumiere, WALT, and Zeroscope. The experiments demonstrate that our method has a significant improvement over previous methods. Project page at: https://daipengwa.github.io/S-2VG_ProjectPage/
title S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08048