Generalized Geometry Encoding Volume for Real-time Stereo Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiaxin, Xu, Gangwei, Wang, Xianqi, Zhang, Chengliang, Yang, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918236115697664
author Liu, Jiaxin
Xu, Gangwei
Wang, Xianqi
Zhang, Chengliang
Yang, Xin
author_facet Liu, Jiaxin
Xu, Gangwei
Wang, Xianqi
Zhang, Chengliang
Yang, Xin
contents Real-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In contrast, recent stereo foundation models leverage monocular foundation models (MFMs) to improve generalization, but typically suffer from substantial inference latency. To address this trade-off, we propose Generalized Geometry Encoding Volume (GGEV), a novel real-time stereo matching network that achieves strong generalization. We first extract depth-aware features that encode domain-invariant structural priors as guidance for cost aggregation. Subsequently, we introduce a Depth-aware Dynamic Cost Aggregation (DDCA) module that adaptively incorporates these priors into each disparity hypothesis, effectively enhancing fragile matching relationships in unseen scenes. Both steps are lightweight and complementary, leading to the construction of a generalized geometry encoding volume with strong generalization capability. Experimental results demonstrate that our GGEV surpasses all existing real-time methods in zero-shot generalization capability, and achieves state-of-the-art performance on the KITTI 2012, KITTI 2015, and ETH3D benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalized Geometry Encoding Volume for Real-time Stereo Matching
Liu, Jiaxin
Xu, Gangwei
Wang, Xianqi
Zhang, Chengliang
Yang, Xin
Computer Vision and Pattern Recognition
Real-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In contrast, recent stereo foundation models leverage monocular foundation models (MFMs) to improve generalization, but typically suffer from substantial inference latency. To address this trade-off, we propose Generalized Geometry Encoding Volume (GGEV), a novel real-time stereo matching network that achieves strong generalization. We first extract depth-aware features that encode domain-invariant structural priors as guidance for cost aggregation. Subsequently, we introduce a Depth-aware Dynamic Cost Aggregation (DDCA) module that adaptively incorporates these priors into each disparity hypothesis, effectively enhancing fragile matching relationships in unseen scenes. Both steps are lightweight and complementary, leading to the construction of a generalized geometry encoding volume with strong generalization capability. Experimental results demonstrate that our GGEV surpasses all existing real-time methods in zero-shot generalization capability, and achieves state-of-the-art performance on the KITTI 2012, KITTI 2015, and ETH3D benchmarks.
title Generalized Geometry Encoding Volume for Real-time Stereo Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.06793