Stereo Any Video: Temporally Consistent Stereo Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jing, Junpeng, Luo, Weixun, Mao, Ye, Mikolajczyk, Krystian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908458516742144
author Jing, Junpeng
Luo, Weixun
Mao, Ye
Mikolajczyk, Krystian
author_facet Jing, Junpeng
Luo, Weixun
Mao, Ye
Mikolajczyk, Krystian
contents This paper introduces Stereo Any Video, a powerful framework for video stereo matching. It can estimate spatially accurate and temporally consistent disparities without relying on auxiliary information such as camera poses or optical flow. The strong capability is driven by rich priors from monocular video depth models, which are integrated with convolutional features to produce stable representations. To further enhance performance, key architectural innovations are introduced: all-to-all-pairs correlation, which constructs smooth and robust matching cost volumes, and temporal convex upsampling, which improves temporal coherence. These components collectively ensure robustness, accuracy, and temporal consistency, setting a new standard in video stereo matching. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple datasets both qualitatively and quantitatively in zero-shot settings, as well as strong generalization to real-world indoor and outdoor scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stereo Any Video: Temporally Consistent Stereo Matching
Jing, Junpeng
Luo, Weixun
Mao, Ye
Mikolajczyk, Krystian
Computer Vision and Pattern Recognition
This paper introduces Stereo Any Video, a powerful framework for video stereo matching. It can estimate spatially accurate and temporally consistent disparities without relying on auxiliary information such as camera poses or optical flow. The strong capability is driven by rich priors from monocular video depth models, which are integrated with convolutional features to produce stable representations. To further enhance performance, key architectural innovations are introduced: all-to-all-pairs correlation, which constructs smooth and robust matching cost volumes, and temporal convex upsampling, which improves temporal coherence. These components collectively ensure robustness, accuracy, and temporal consistency, setting a new standard in video stereo matching. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple datasets both qualitatively and quantitatively in zero-shot settings, as well as strong generalization to real-world indoor and outdoor scenarios.
title Stereo Any Video: Temporally Consistent Stereo Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.05549