From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yepeng, Li, Hao, Yang, Liwen, Li, Fangzhen, Ge, Xudi, Gu, Yuliang, Gao, kuang, Wang, Bing, Chen, Guang, Ye, Hangjun, Xu, Yongchao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917498126860288
author Liu, Yepeng
Li, Hao
Yang, Liwen
Li, Fangzhen
Ge, Xudi
Gu, Yuliang
Gao, kuang
Wang, Bing
Chen, Guang
Ye, Hangjun
Xu, Yongchao
author_facet Liu, Yepeng
Li, Hao
Yang, Liwen
Li, Fangzhen
Ge, Xudi
Gu, Yuliang
Gao, kuang
Wang, Bing
Chen, Guang
Ye, Hangjun
Xu, Yongchao
contents Keypoint-based matching is a fundamental component of modern 3D vision systems, such as Structure-from-Motion (SfM) and SLAM. Most existing learning-based methods are trained on image pairs, a paradigm that fails to explicitly optimize for the long-term trackability of keypoints across sequences under challenging viewpoint and illumination changes. In this paper, we reframe keypoint detection as a sequential decision-making problem. We introduce TraqPoint, a novel, end-to-end Reinforcement Learning (RL) framework designed to optimize the \textbf{Tra}ck-\textbf{q}uality (Traq) of keypoints directly on image sequences. Our core innovation is a track-aware reward mechanism that jointly encourages the consistency and distinctiveness of keypoints across multiple views, guided by a policy gradient method. Extensive evaluations on sparse matching benchmarks, including relative pose estimation and 3D reconstruction, demonstrate that TraqPoint significantly outperforms some state-of-the-art (SOTA) keypoint detection and description methods.The code will be available at https://github.com/xiaomi-research/traqpoint.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20630
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
Liu, Yepeng
Li, Hao
Yang, Liwen
Li, Fangzhen
Ge, Xudi
Gu, Yuliang
Gao, kuang
Wang, Bing
Chen, Guang
Ye, Hangjun
Xu, Yongchao
Computer Vision and Pattern Recognition
Keypoint-based matching is a fundamental component of modern 3D vision systems, such as Structure-from-Motion (SfM) and SLAM. Most existing learning-based methods are trained on image pairs, a paradigm that fails to explicitly optimize for the long-term trackability of keypoints across sequences under challenging viewpoint and illumination changes. In this paper, we reframe keypoint detection as a sequential decision-making problem. We introduce TraqPoint, a novel, end-to-end Reinforcement Learning (RL) framework designed to optimize the \textbf{Tra}ck-\textbf{q}uality (Traq) of keypoints directly on image sequences. Our core innovation is a track-aware reward mechanism that jointly encourages the consistency and distinctiveness of keypoints across multiple views, guided by a policy gradient method. Extensive evaluations on sparse matching benchmarks, including relative pose estimation and 3D reconstruction, demonstrate that TraqPoint significantly outperforms some state-of-the-art (SOTA) keypoint detection and description methods.The code will be available at https://github.com/xiaomi-research/traqpoint.
title From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.20630