SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Qianxiong, Zhu, Lanyun, Liu, Chenxi, Lin, Guosheng, Long, Cheng, Li, Ziyue, Zhao, Rui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913964889210880
author Xu, Qianxiong
Zhu, Lanyun
Liu, Chenxi
Lin, Guosheng
Long, Cheng
Li, Ziyue
Zhao, Rui
author_facet Xu, Qianxiong
Zhu, Lanyun
Liu, Chenxi
Lin, Guosheng
Long, Cheng
Li, Ziyue
Zhao, Rui
contents Visual Object Tracking (VOT) is widely used in applications like autonomous driving to continuously track targets in videos. Existing methods can be roughly categorized into template matching and autoregressive methods, where the former usually neglects the temporal dependencies across frames and the latter tends to get biased towards the object categories during training, showing weak generalizability to unseen classes. To address these issues, some methods propose to adapt the video foundation model SAM2 for VOT, where the tracking results of each frame would be encoded as memory for conditioning the rest of frames in an autoregressive manner. Nevertheless, existing methods fail to overcome the challenges of object occlusions and distractions, and do not have any measures to intercept the propagation of tracking errors. To tackle them, we present a SAMITE model, built upon SAM2 with additional modules, including: (1) Prototypical Memory Bank: We propose to quantify the feature-wise and position-wise correctness of each frame's tracking results, and select the best frames to condition subsequent frames. As the features of occluded and distracting objects are feature-wise and position-wise inaccurate, their scores would naturally be lower and thus can be filtered to intercept error propagation; (2) Positional Prompt Generator: To further reduce the impacts of distractors, we propose to generate positional mask prompts to provide explicit positional clues for the target, leading to more accurate tracking. Extensive experiments have been conducted on six benchmarks, showing the superiority of SAMITE. The code is available at https://github.com/Sam1224/SAMITE.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21732
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking
Xu, Qianxiong
Zhu, Lanyun
Liu, Chenxi
Lin, Guosheng
Long, Cheng
Li, Ziyue
Zhao, Rui
Computer Vision and Pattern Recognition
Visual Object Tracking (VOT) is widely used in applications like autonomous driving to continuously track targets in videos. Existing methods can be roughly categorized into template matching and autoregressive methods, where the former usually neglects the temporal dependencies across frames and the latter tends to get biased towards the object categories during training, showing weak generalizability to unseen classes. To address these issues, some methods propose to adapt the video foundation model SAM2 for VOT, where the tracking results of each frame would be encoded as memory for conditioning the rest of frames in an autoregressive manner. Nevertheless, existing methods fail to overcome the challenges of object occlusions and distractions, and do not have any measures to intercept the propagation of tracking errors. To tackle them, we present a SAMITE model, built upon SAM2 with additional modules, including: (1) Prototypical Memory Bank: We propose to quantify the feature-wise and position-wise correctness of each frame's tracking results, and select the best frames to condition subsequent frames. As the features of occluded and distracting objects are feature-wise and position-wise inaccurate, their scores would naturally be lower and thus can be filtered to intercept error propagation; (2) Positional Prompt Generator: To further reduce the impacts of distractors, we propose to generate positional mask prompts to provide explicit positional clues for the target, leading to more accurate tracking. Extensive experiments have been conducted on six benchmarks, showing the superiority of SAMITE. The code is available at https://github.com/Sam1224/SAMITE.
title SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.21732