StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Behrens, Tjark, Obukhov, Anton, Ke, Bingxin, Tosi, Fabio, Poggi, Matteo, Schindler, Konrad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917446405849088
author Behrens, Tjark
Obukhov, Anton
Ke, Bingxin
Tosi, Fabio
Poggi, Matteo
Schindler, Konrad
author_facet Behrens, Tjark
Obukhov, Anton
Ke, Bingxin
Tosi, Fabio
Poggi, Matteo
Schindler, Konrad
contents We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning, without explicit depth or warping. A canonical rectified space and the conditioning guide the generator to infer correspondences and fill disocclusions end-to-end. To ensure fair and leakage-free evaluation, we introduce an end-to-end protocol that excludes any ground truth or proxy geometry estimates at test time. The protocol emphasizes metrics reflecting downstream relevance: iSQoE for perceptual comfort and MEt3R for geometric consistency. StereoSpace surpasses other methods from the warp & inpaint, latent-warping, and warped-conditioning categories, achieving sharp parallax and strong robustness on layered and non-Lambertian scenes. This establishes viewpoint-conditioned diffusion as a scalable, depth-free solution for stereo generation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10959
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
Behrens, Tjark
Obukhov, Anton
Ke, Bingxin
Tosi, Fabio
Poggi, Matteo
Schindler, Konrad
Computer Vision and Pattern Recognition
We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning, without explicit depth or warping. A canonical rectified space and the conditioning guide the generator to infer correspondences and fill disocclusions end-to-end. To ensure fair and leakage-free evaluation, we introduce an end-to-end protocol that excludes any ground truth or proxy geometry estimates at test time. The protocol emphasizes metrics reflecting downstream relevance: iSQoE for perceptual comfort and MEt3R for geometric consistency. StereoSpace surpasses other methods from the warp & inpaint, latent-warping, and warped-conditioning categories, achieving sharp parallax and strong robustness on layered and non-Lambertian scenes. This establishes viewpoint-conditioned diffusion as a scalable, depth-free solution for stereo generation.
title StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.10959