BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yufei, Lu, Sha, Han, Fuzhang, Xiong, Rong, Wang, Yue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913578804576256
author Wei, Yufei
Lu, Sha
Han, Fuzhang
Xiong, Rong
Wang, Yue
author_facet Wei, Yufei
Lu, Sha
Han, Fuzhang
Xiong, Rong
Wang, Yue
contents Monocular visual odometry (MVO) is vital in autonomous navigation and robotics, providing a cost-effective and flexible motion tracking solution, but the inherent scale ambiguity in monocular setups often leads to cumulative errors over time. In this paper, we present BEV-ODOM, a novel MVO framework leveraging the Bird's Eye View (BEV) Representation to address scale drift. Unlike existing approaches, BEV-ODOM integrates a depth-based perspective-view (PV) to BEV encoder, a correlation feature extraction neck, and a CNN-MLP-based decoder, enabling it to estimate motion across three degrees of freedom without the need for depth supervision or complex optimization techniques. Our framework reduces scale drift in long-term sequences and achieves accurate motion estimation across various datasets, including NCLT, Oxford, and KITTI. The results indicate that BEV-ODOM outperforms current MVO methods, demonstrating reduced scale drift and higher accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation
Wei, Yufei
Lu, Sha
Han, Fuzhang
Xiong, Rong
Wang, Yue
Robotics
Monocular visual odometry (MVO) is vital in autonomous navigation and robotics, providing a cost-effective and flexible motion tracking solution, but the inherent scale ambiguity in monocular setups often leads to cumulative errors over time. In this paper, we present BEV-ODOM, a novel MVO framework leveraging the Bird's Eye View (BEV) Representation to address scale drift. Unlike existing approaches, BEV-ODOM integrates a depth-based perspective-view (PV) to BEV encoder, a correlation feature extraction neck, and a CNN-MLP-based decoder, enabling it to estimate motion across three degrees of freedom without the need for depth supervision or complex optimization techniques. Our framework reduces scale drift in long-term sequences and achieves accurate motion estimation across various datasets, including NCLT, Oxford, and KITTI. The results indicate that BEV-ODOM outperforms current MVO methods, demonstrating reduced scale drift and higher accuracy.
title BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation
topic Robotics
url https://arxiv.org/abs/2411.10195