RELO: Reinforcement Learning to Localize for Visual Object Tracking

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Xin, Sun, Chuanyu, Xu, Jiao, Peng, Houwen, Wang, Dong, Lu, Huchuan, Ma, Kede
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916024069128192
author Chen, Xin
Sun, Chuanyu
Xu, Jiao
Peng, Houwen
Wang, Dong
Lu, Huchuan
Ma, Kede
author_facet Chen, Xin
Sun, Chuanyu
Xu, Jiao
Peng, Houwen
Wang, Dong
Lu, Huchuan
Ma, Kede
contents Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining 57.5% AUC on LaSOText without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07379
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RELO: Reinforcement Learning to Localize for Visual Object Tracking
Chen, Xin
Sun, Chuanyu
Xu, Jiao
Peng, Houwen
Wang, Dong
Lu, Huchuan
Ma, Kede
Computer Vision and Pattern Recognition
Artificial Intelligence
Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly aligned with tracking optimization and evaluation metrics, such as intersection over union (IoU) and area under the success curve (AUC). Here, we introduce RELO, a REinforcement-learning-to-LOcalize method for visual object tracking that formulates target localization as a Markov decision process. Specifically, RELO replaces handcrafted spatial priors with a localization policy learned over spatial positions via reinforcement learning, with rewards combining frame-level IoU and sequence-level AUC. We additionally introduce layer-aligned temporal token propagation to improve semantic consistency across frames, with negligible computational overhead. Across multiple benchmarks, RELO achieves superior results, attaining 57.5% AUC on LaSOText without template updates. This confirms that reward-driven localization provides an effective alternative to prior-driven localization for visual object tracking.
title RELO: Reinforcement Learning to Localize for Visual Object Tracking
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.07379