Efficient Motion Prompt Learning for Robust Visual Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jie, Chen, Xin, Yuan, Yongsheng, Felsberg, Michael, Wang, Dong, Lu, Huchuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912952759615488
author Zhao, Jie
Chen, Xin
Yuan, Yongsheng
Felsberg, Michael
Wang, Dong
Lu, Huchuan
author_facet Zhao, Jie
Chen, Xin
Yuan, Yongsheng
Felsberg, Michael
Wang, Dong
Lu, Huchuan
contents Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existing vision-based trackers to build a joint tracking framework leveraging both motion and vision cues, thereby achieving robust tracking through efficient prompt learning. A motion encoder with three different positional encodings is proposed to encode the long-term motion trajectory into the visual embedding space, while a fusion decoder and an adaptive weight mechanism are designed to dynamically fuse visual and motion features. We integrate our motion module into three different trackers with five models in total. Experiments on seven challenging tracking benchmarks demonstrate that the proposed motion module significantly improves the robustness of vision-based trackers, with minimal training costs and negligible speed sacrifice. Code is available at https://github.com/zj5559/Motion-Prompt-Tracking.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Motion Prompt Learning for Robust Visual Tracking
Zhao, Jie
Chen, Xin
Yuan, Yongsheng
Felsberg, Michael
Wang, Dong
Lu, Huchuan
Computer Vision and Pattern Recognition
Due to the challenges of processing temporal information, most trackers depend solely on visual discriminability and overlook the unique temporal coherence of video data. In this paper, we propose a lightweight and plug-and-play motion prompt tracking method. It can be easily integrated into existing vision-based trackers to build a joint tracking framework leveraging both motion and vision cues, thereby achieving robust tracking through efficient prompt learning. A motion encoder with three different positional encodings is proposed to encode the long-term motion trajectory into the visual embedding space, while a fusion decoder and an adaptive weight mechanism are designed to dynamically fuse visual and motion features. We integrate our motion module into three different trackers with five models in total. Experiments on seven challenging tracking benchmarks demonstrate that the proposed motion module significantly improves the robustness of vision-based trackers, with minimal training costs and negligible speed sacrifice. Code is available at https://github.com/zj5559/Motion-Prompt-Tracking.
title Efficient Motion Prompt Learning for Robust Visual Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16321