Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Zhaoxing, Cheng, Junda, Xu, Gangwei, Wang, Xiaoxiang, Zhang, Can, Yang, Xin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910864781606912
author Zhang, Zhaoxing
Cheng, Junda
Xu, Gangwei
Wang, Xiaoxiang
Zhang, Can
Yang, Xin
author_facet Zhang, Zhaoxing
Cheng, Junda
Xu, Gangwei
Wang, Xiaoxiang
Zhang, Can
Yang, Xin
contents Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to handle challenging scenarios and long-sequence estimation. To overcome these challenges, we introduce Spatio-Temporal Visual Odometry (STVO), a novel deep network architecture that effectively leverages inherent spatio-temporal cues to enhance the accuracy and consistency of multi-frame flow matching. With more accurate and consistent flow matching, STVO can achieve better pose estimation through the bundle adjustment (BA). Specifically, STVO introduces two innovative components: 1) the Temporal Propagation Module that utilizes multi-frame information to extract and propagate temporal cues across adjacent frames, maintaining temporal consistency; 2) the Spatial Activation Module that utilizes geometric priors from the depth maps to enhance spatial consistency while filtering out excessive noise and incorrect matches. Our STVO achieves state-of-the-art performance on TUM-RGBD, EuRoc MAV, ETH3D and KITTI Odometry benchmarks. Notably, it improves accuracy by 77.8% on ETH3D benchmark and 38.9% on KITTI Odometry benchmark over the previous best methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16923
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
Zhang, Zhaoxing
Cheng, Junda
Xu, Gangwei
Wang, Xiaoxiang
Zhang, Can
Yang, Xin
Computer Vision and Pattern Recognition
Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to handle challenging scenarios and long-sequence estimation. To overcome these challenges, we introduce Spatio-Temporal Visual Odometry (STVO), a novel deep network architecture that effectively leverages inherent spatio-temporal cues to enhance the accuracy and consistency of multi-frame flow matching. With more accurate and consistent flow matching, STVO can achieve better pose estimation through the bundle adjustment (BA). Specifically, STVO introduces two innovative components: 1) the Temporal Propagation Module that utilizes multi-frame information to extract and propagate temporal cues across adjacent frames, maintaining temporal consistency; 2) the Spatial Activation Module that utilizes geometric priors from the depth maps to enhance spatial consistency while filtering out excessive noise and incorrect matches. Our STVO achieves state-of-the-art performance on TUM-RGBD, EuRoc MAV, ETH3D and KITTI Odometry benchmarks. Notably, it improves accuracy by 77.8% on ETH3D benchmark and 38.9% on KITTI Odometry benchmark over the previous best methods.
title Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.16923