M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shvetsova, Nina, Bhat, Goutam, Truong, Prune, Kuehne, Hilde, Tombari, Federico
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918216797782016
author Shvetsova, Nina
Bhat, Goutam
Truong, Prune
Kuehne, Hilde
Tombari, Federico
author_facet Shvetsova, Nina
Bhat, Goutam
Truong, Prune
Kuehne, Hilde
Tombari, Federico
contents We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right view obtained by depth-based reprojection of the input left view. We extend the Stable Video Diffusion (SVD) model to utilize the input left video, the warped right video, and the disocclusion masks as conditioning input to generate a high-quality right camera view. In order to effectively exploit information from neighboring frames for inpainting, we modify the attention layers in SVD to compute full attention for discoccluded pixels. Our model is trained to generate the right view video in an end-to-end manner without iterative diffusion steps by minimizing image space losses to ensure high-quality generation. Our approach outperforms previous state-of-the-art methods, being ranked best 2.6x more often than the second-place method in a user study, while being 6x faster.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16565
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
Shvetsova, Nina
Bhat, Goutam
Truong, Prune
Kuehne, Hilde
Tombari, Federico
Computer Vision and Pattern Recognition
We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right view obtained by depth-based reprojection of the input left view. We extend the Stable Video Diffusion (SVD) model to utilize the input left video, the warped right video, and the disocclusion masks as conditioning input to generate a high-quality right camera view. In order to effectively exploit information from neighboring frames for inpainting, we modify the attention layers in SVD to compute full attention for discoccluded pixels. Our model is trained to generate the right view video in an end-to-end manner without iterative diffusion steps by minimizing image space losses to ensure high-quality generation. Our approach outperforms previous state-of-the-art methods, being ranked best 2.6x more often than the second-place method in a user study, while being 6x faster.
title M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16565