Keyframe-Based Feed-Forward Visual Odometry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Weichen, Su, Wenhan, Kong, Da, Ming, Yuhang, Kong, Wanzeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911392256229376
author Dai, Weichen
Su, Wenhan
Kong, Da
Ming, Yuhang
Kong, Wanzeng
author_facet Dai, Weichen
Su, Wenhan
Kong, Da
Ming, Yuhang
Kong, Wanzeng
contents The emergence of visual foundation models has revolutionized visual odometry~(VO) and SLAM, enabling pose estimation and dense reconstruction within a single feed-forward network. However, unlike traditional pipelines that leverage keyframe methods to enhance efficiency and accuracy, current foundation model based methods, such as VGGT-Long, typically process raw image sequences indiscriminately. This leads to computational redundancy and degraded performance caused by low inter-frame parallax, which provides limited contextual stereo information. Integrating traditional geometric heuristics into these methods is non-trivial, as their performance depends on high-dimensional latent representations rather than explicit geometric metrics. To bridge this gap, we propose a novel keyframe-based feed-forward VO. Instead of relying on hand-crafted rules, our approach employs reinforcement learning to derive an adaptive keyframe policy in a data-driven manner, aligning selection with the intrinsic characteristics of the underlying foundation model. We train our agent on TartanAir dataset and conduct extensive evaluations across several real-world datasets. Experimental results demonstrate that the proposed method achieves consistent and substantial improvements over state-of-the-art feed-forward VO methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16020
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Keyframe-Based Feed-Forward Visual Odometry
Dai, Weichen
Su, Wenhan
Kong, Da
Ming, Yuhang
Kong, Wanzeng
Computer Vision and Pattern Recognition
Robotics
The emergence of visual foundation models has revolutionized visual odometry~(VO) and SLAM, enabling pose estimation and dense reconstruction within a single feed-forward network. However, unlike traditional pipelines that leverage keyframe methods to enhance efficiency and accuracy, current foundation model based methods, such as VGGT-Long, typically process raw image sequences indiscriminately. This leads to computational redundancy and degraded performance caused by low inter-frame parallax, which provides limited contextual stereo information. Integrating traditional geometric heuristics into these methods is non-trivial, as their performance depends on high-dimensional latent representations rather than explicit geometric metrics. To bridge this gap, we propose a novel keyframe-based feed-forward VO. Instead of relying on hand-crafted rules, our approach employs reinforcement learning to derive an adaptive keyframe policy in a data-driven manner, aligning selection with the intrinsic characteristics of the underlying foundation model. We train our agent on TartanAir dataset and conduct extensive evaluations across several real-world datasets. Experimental results demonstrate that the proposed method achieves consistent and substantial improvements over state-of-the-art feed-forward VO methods.
title Keyframe-Based Feed-Forward Visual Odometry
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2601.16020