Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Geyer, Michal, Tov, Omer, Jin, Linyi, Tucker, Richard, Mosseri, Inbar, Dekel, Tali, Snavely, Noah
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916715703566336
author Geyer, Michal
Tov, Omer
Jin, Linyi
Tucker, Richard
Mosseri, Inbar
Dekel, Tali
Snavely, Noah
author_facet Geyer, Michal
Tov, Omer
Jin, Linyi
Tucker, Richard
Mosseri, Inbar
Dekel, Tali
Snavely, Noah
contents The rising popularity of immersive visual experiences has increased interest in stereoscopic 3D video generation. Despite significant advances in video synthesis, creating 3D videos remains challenging due to the relative scarcity of 3D video data. We propose a simple approach for transforming a text-to-video generator into a video-to-stereo generator. Given an input video, our framework automatically produces the video frames from a shifted viewpoint, enabling a compelling 3D effect. Prior and concurrent approaches for this task typically operate in multiple phases, first estimating video disparity or depth, then warping the video accordingly to produce a second view, and finally inpainting the disoccluded regions. This approach inherently fails when the scene involves specular surfaces or transparent objects. In such cases, single-layer disparity estimation is insufficient, resulting in artifacts and incorrect pixel shifts during warping. Our work bypasses these restrictions by directly synthesizing the new viewpoint, avoiding any intermediate steps. This is achieved by leveraging a pre-trained video model's priors on geometry, object materials, optics, and semantics, without relying on external geometry models or manually disentangling geometry from the synthesis process. We demonstrate the advantages of our approach in complex, real-world scenarios featuring diverse object materials and compositions. See videos on https://video-eye2eye.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2505_00135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis
Geyer, Michal
Tov, Omer
Jin, Linyi
Tucker, Richard
Mosseri, Inbar
Dekel, Tali
Snavely, Noah
Computer Vision and Pattern Recognition
The rising popularity of immersive visual experiences has increased interest in stereoscopic 3D video generation. Despite significant advances in video synthesis, creating 3D videos remains challenging due to the relative scarcity of 3D video data. We propose a simple approach for transforming a text-to-video generator into a video-to-stereo generator. Given an input video, our framework automatically produces the video frames from a shifted viewpoint, enabling a compelling 3D effect. Prior and concurrent approaches for this task typically operate in multiple phases, first estimating video disparity or depth, then warping the video accordingly to produce a second view, and finally inpainting the disoccluded regions. This approach inherently fails when the scene involves specular surfaces or transparent objects. In such cases, single-layer disparity estimation is insufficient, resulting in artifacts and incorrect pixel shifts during warping. Our work bypasses these restrictions by directly synthesizing the new viewpoint, avoiding any intermediate steps. This is achieved by leveraging a pre-trained video model's priors on geometry, object materials, optics, and semantics, without relying on external geometry models or manually disentangling geometry from the synthesis process. We demonstrate the advantages of our approach in complex, real-world scenarios featuring diverse object materials and compositions. See videos on https://video-eye2eye.github.io
title Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.00135