FVO: Fast Visual Odometry with Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yugay, Vlardimir, Nguyen, Duy-Kien, Gevers, Theo, Snoek, Cees G. M., Oswald, Martin R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915844440719360
author Yugay, Vlardimir
Nguyen, Duy-Kien
Gevers, Theo
Snoek, Cees G. M.
Oswald, Martin R.
author_facet Yugay, Vlardimir
Nguyen, Duy-Kien
Gevers, Theo
Snoek, Cees G. M.
Oswald, Martin R.
contents Hybrid pipelines that combine deep learning with classical optimization have established themselves as the dominant approach to visual odometry (VO). By integrating neural network predictions with bundle adjustment, these models estimate camera trajectories with high accuracy. Still, hybrid VO methods fall short of the speed and capabilities of pure end-to-end approaches. Current hybrid frameworks rely on massive, pre-trained 3D networks to predict geometry. Because these backends are trained to be scale-ambiguous and frozen rather than retrained, the pipelines essentially inherit this limitation and, by design, fails to estimate absolute scale. Furthermore, their slow optimization and post-processing steps bottleneck the pipeline's inference speed. We propose to replace post-processing entirely by formulating monocular visual odometry as a direct relative pose regression problem. This formulation enables us to train a fast, high-capacity transformer to predict relative camera poses and corresponding confidences using only camera poses as supervision. More importantly, it allows us to employ a confidence-aware inference scheme that aggregates overlapping pose predictions for robust trajectory estimation. We demonstrate on multiple visual odometry benchmarks that our method, Fast Visual Odometry (FVO), successfully leverages diverse data to achieve competitive or superior performance while being nearly 2 times faster than the fastest baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03348
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FVO: Fast Visual Odometry with Transformers
Yugay, Vlardimir
Nguyen, Duy-Kien
Gevers, Theo
Snoek, Cees G. M.
Oswald, Martin R.
Computer Vision and Pattern Recognition
Hybrid pipelines that combine deep learning with classical optimization have established themselves as the dominant approach to visual odometry (VO). By integrating neural network predictions with bundle adjustment, these models estimate camera trajectories with high accuracy. Still, hybrid VO methods fall short of the speed and capabilities of pure end-to-end approaches. Current hybrid frameworks rely on massive, pre-trained 3D networks to predict geometry. Because these backends are trained to be scale-ambiguous and frozen rather than retrained, the pipelines essentially inherit this limitation and, by design, fails to estimate absolute scale. Furthermore, their slow optimization and post-processing steps bottleneck the pipeline's inference speed. We propose to replace post-processing entirely by formulating monocular visual odometry as a direct relative pose regression problem. This formulation enables us to train a fast, high-capacity transformer to predict relative camera poses and corresponding confidences using only camera poses as supervision. More importantly, it allows us to employ a confidence-aware inference scheme that aggregates overlapping pose predictions for robust trajectory estimation. We demonstrate on multiple visual odometry benchmarks that our method, Fast Visual Odometry (FVO), successfully leverages diverse data to achieve competitive or superior performance while being nearly 2 times faster than the fastest baselines.
title FVO: Fast Visual Odometry with Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03348