SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ma, Yinchao, Yang, Dengqing, He, Zhangyu, Yang, Wenfei, Zhang, Tianzhu
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912868200349696
author Ma, Yinchao
Yang, Dengqing
He, Zhangyu
Yang, Wenfei
Zhang, Tianzhu
author_facet Ma, Yinchao
Yang, Dengqing
He, Zhangyu
Yang, Wenfei
Zhang, Tianzhu
contents Visual tracking aims to automatically estimate the state of a target object in a video sequence, which is challenging especially in dynamic scenarios. Thus, numerous methods are proposed to introduce temporal cues to enhance tracking robustness. However, conventional CNN and Transformer architectures exhibit inherent limitations in modeling long-range temporal dependencies in visual tracking, often necessitating either complex customized modules or substantial computational costs to integrate temporal cues. Inspired by the success of the state space model, we propose a novel temporal modeling paradigm for visual tracking, termed State-aware Mamba Tracker (SMTrack), providing a neat pipeline for training and tracking without needing customized modules or substantial computational costs to build long-range temporal dependencies. It enjoys several merits. First, we propose a novel selective state-aware space model with state-wise parameters to capture more diverse temporal cues for robust tracking. Second, SMTrack facilitates long-range temporal interactions with linear computational complexity during training. Third, SMTrack enables each frame to interact with previously tracked frames via hidden state propagation and updating, which releases computational costs of handling temporal cues during tracking. Extensive experimental results demonstrate that SMTrack achieves promising performance with low computational costs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01677
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
Ma, Yinchao
Yang, Dengqing
He, Zhangyu
Yang, Wenfei
Zhang, Tianzhu
Computer Vision and Pattern Recognition
Visual tracking aims to automatically estimate the state of a target object in a video sequence, which is challenging especially in dynamic scenarios. Thus, numerous methods are proposed to introduce temporal cues to enhance tracking robustness. However, conventional CNN and Transformer architectures exhibit inherent limitations in modeling long-range temporal dependencies in visual tracking, often necessitating either complex customized modules or substantial computational costs to integrate temporal cues. Inspired by the success of the state space model, we propose a novel temporal modeling paradigm for visual tracking, termed State-aware Mamba Tracker (SMTrack), providing a neat pipeline for training and tracking without needing customized modules or substantial computational costs to build long-range temporal dependencies. It enjoys several merits. First, we propose a novel selective state-aware space model with state-wise parameters to capture more diverse temporal cues for robust tracking. Second, SMTrack facilitates long-range temporal interactions with linear computational complexity during training. Third, SMTrack enables each frame to interact with previously tracked frames via hidden state propagation and updating, which releases computational costs of handling temporal cues during tracking. Extensive experimental results demonstrate that SMTrack achieves promising performance with low computational costs.
title SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.01677