UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liang, Qihua, Chen, Liang, Zheng, Yaozong, Nong, Jian, Mo, Zhiyi, Zhong, Bineng
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914269345349632
author Liang, Qihua
Chen, Liang
Zheng, Yaozong
Nong, Jian
Mo, Zhiyi
Zhong, Bineng
author_facet Liang, Qihua
Chen, Liang
Zheng, Yaozong
Nong, Jian
Mo, Zhiyi
Zhong, Bineng
contents Multi-modal object tracking has attracted considerable attention by integrating multiple complementary inputs (e.g., thermal, depth, and event data) to achieve outstanding performance. Although current general-purpose multi-modal trackers primarily unify various modal tracking tasks (i.e., RGB-Thermal infrared, RGB-Depth or RGB-Event tracking) through prompt learning, they still overlook the effective capture of spatio-temporal cues. In this work, we introduce a novel multi-modal tracking framework based on a mamba-style state space model, termed UBATrack. Our UBATrack comprises two simple yet effective modules: a Spatio-temporal Mamba Adapter (STMA) and a Dynamic Multi-modal Feature Mixer. The former leverages Mamba's long-sequence modeling capability to jointly model cross-modal dependencies and spatio-temporal visual cues in an adapter-tuning manner. The latter further enhances multi-modal representation capacity across multiple feature dimensions to improve tracking robustness. In this way, UBATrack eliminates the need for costly full-parameter fine-tuning, thereby improving the training efficiency of multi-modal tracking algorithms. Experiments show that UBATrack outperforms state-of-the-art methods on RGB-T, RGB-D, and RGB-E tracking benchmarks, achieving outstanding results on the LasHeR, RGBT234, RGBT210, DepthTrack, VOT-RGBD22, and VisEvent datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14799
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
Liang, Qihua
Chen, Liang
Zheng, Yaozong
Nong, Jian
Mo, Zhiyi
Zhong, Bineng
Computer Vision and Pattern Recognition
Multi-modal object tracking has attracted considerable attention by integrating multiple complementary inputs (e.g., thermal, depth, and event data) to achieve outstanding performance. Although current general-purpose multi-modal trackers primarily unify various modal tracking tasks (i.e., RGB-Thermal infrared, RGB-Depth or RGB-Event tracking) through prompt learning, they still overlook the effective capture of spatio-temporal cues. In this work, we introduce a novel multi-modal tracking framework based on a mamba-style state space model, termed UBATrack. Our UBATrack comprises two simple yet effective modules: a Spatio-temporal Mamba Adapter (STMA) and a Dynamic Multi-modal Feature Mixer. The former leverages Mamba's long-sequence modeling capability to jointly model cross-modal dependencies and spatio-temporal visual cues in an adapter-tuning manner. The latter further enhances multi-modal representation capacity across multiple feature dimensions to improve tracking robustness. In this way, UBATrack eliminates the need for costly full-parameter fine-tuning, thereby improving the training efficiency of multi-modal tracking algorithms. Experiments show that UBATrack outperforms state-of-the-art methods on RGB-T, RGB-D, and RGB-E tracking benchmarks, achieving outstanding results on the LasHeR, RGBT234, RGBT210, DepthTrack, VOT-RGBD22, and VisEvent datasets.
title UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.14799