Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Shiao, Huang, Ju, Ma, Qingchuan, Gao, Jinfeng, Xu, Chunyi, Wang, Xiao, Chen, Lan, Jiang, Bo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909667125362688
author Wang, Shiao
Huang, Ju
Ma, Qingchuan
Gao, Jinfeng
Xu, Chunyi
Wang, Xiao
Chen, Lan
Jiang, Bo
author_facet Wang, Shiao
Huang, Ju
Ma, Qingchuan
Gao, Jinfeng
Xu, Chunyi
Wang, Xiao
Chen, Lan
Jiang, Bo
contents Combining traditional RGB cameras with bio-inspired event cameras for robust object tracking has garnered increasing attention in recent years. However, most existing multimodal tracking algorithms depend heavily on high-complexity Vision Transformer architectures for feature extraction and fusion across modalities. This not only leads to substantial computational overhead but also limits the effectiveness of cross-modal interactions. In this paper, we propose an efficient RGB-Event object tracking framework based on the linear-complexity Vision Mamba network, termed Mamba-FETrack V2. Specifically, we first design a lightweight Prompt Generator that utilizes embedded features from each modality, together with a shared prompt pool, to dynamically generate modality-specific learnable prompt vectors. These prompts, along with the modality-specific embedded features, are then fed into a Vision Mamba-based FEMamba backbone, which facilitates prompt-guided feature extraction, cross-modal interaction, and fusion in a unified manner. Finally, the fused representations are passed to the tracking head for accurate target localization. Extensive experimental evaluations on multiple RGB-Event tracking benchmarks, including short-term COESOT dataset and long-term datasets, i.e., FE108 and FELT V2, demonstrate the superior performance and efficiency of the proposed tracking framework. The source code and pre-trained models will be released on https://github.com/Event-AHU/Mamba_FETrack
format Preprint
id arxiv_https___arxiv_org_abs_2506_23783
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking
Wang, Shiao
Huang, Ju
Ma, Qingchuan
Gao, Jinfeng
Xu, Chunyi
Wang, Xiao
Chen, Lan
Jiang, Bo
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Combining traditional RGB cameras with bio-inspired event cameras for robust object tracking has garnered increasing attention in recent years. However, most existing multimodal tracking algorithms depend heavily on high-complexity Vision Transformer architectures for feature extraction and fusion across modalities. This not only leads to substantial computational overhead but also limits the effectiveness of cross-modal interactions. In this paper, we propose an efficient RGB-Event object tracking framework based on the linear-complexity Vision Mamba network, termed Mamba-FETrack V2. Specifically, we first design a lightweight Prompt Generator that utilizes embedded features from each modality, together with a shared prompt pool, to dynamically generate modality-specific learnable prompt vectors. These prompts, along with the modality-specific embedded features, are then fed into a Vision Mamba-based FEMamba backbone, which facilitates prompt-guided feature extraction, cross-modal interaction, and fusion in a unified manner. Finally, the fused representations are passed to the tracking head for accurate target localization. Extensive experimental evaluations on multiple RGB-Event tracking benchmarks, including short-term COESOT dataset and long-term datasets, i.e., FE108 and FELT V2, demonstrate the superior performance and efficiency of the proposed tracking framework. The source code and pre-trained models will be released on https://github.com/Event-AHU/Mamba_FETrack
title Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.23783