Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Xiangyang, Zeng, Dan, Wang, Xucheng, Wu, You, Ye, Hengzhou, Zhao, Qijun, Li, Shuiwang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914853990432768
author Yang, Xiangyang
Zeng, Dan
Wang, Xucheng
Wu, You
Ye, Hengzhou
Zhao, Qijun
Li, Shuiwang
author_facet Yang, Xiangyang
Zeng, Dan
Wang, Xucheng
Wu, You
Ye, Hengzhou
Zhao, Qijun
Li, Shuiwang
contents Empowered by transformer-based models, visual tracking has advanced significantly. However, the slow speed of current trackers limits their applicability on devices with constrained computational resources. To address this challenge, we introduce ABTrack, an adaptive computation framework that adaptively bypassing transformer blocks for efficient visual tracking. The rationale behind ABTrack is rooted in the observation that semantic features or relations do not uniformly impact the tracking task across all abstraction levels. Instead, this impact varies based on the characteristics of the target and the scene it occupies. Consequently, disregarding insignificant semantic features or relations at certain abstraction levels may not significantly affect the tracking accuracy. We propose a Bypass Decision Module (BDM) to determine if a transformer block should be bypassed, which adaptively simplifies the architecture of ViTs and thus speeds up the inference process. To counteract the time cost incurred by the BDMs and further enhance the efficiency of ViTs, we introduce a novel ViT pruning method to reduce the dimension of the latent representation of tokens in each transformer block. Extensive experiments on multiple tracking benchmarks validate the effectiveness and generality of the proposed method and show that it achieves state-of-the-art performance. Code is released at: https://github.com/xyyang317/ABTrack.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08037
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
Yang, Xiangyang
Zeng, Dan
Wang, Xucheng
Wu, You
Ye, Hengzhou
Zhao, Qijun
Li, Shuiwang
Computer Vision and Pattern Recognition
Empowered by transformer-based models, visual tracking has advanced significantly. However, the slow speed of current trackers limits their applicability on devices with constrained computational resources. To address this challenge, we introduce ABTrack, an adaptive computation framework that adaptively bypassing transformer blocks for efficient visual tracking. The rationale behind ABTrack is rooted in the observation that semantic features or relations do not uniformly impact the tracking task across all abstraction levels. Instead, this impact varies based on the characteristics of the target and the scene it occupies. Consequently, disregarding insignificant semantic features or relations at certain abstraction levels may not significantly affect the tracking accuracy. We propose a Bypass Decision Module (BDM) to determine if a transformer block should be bypassed, which adaptively simplifies the architecture of ViTs and thus speeds up the inference process. To counteract the time cost incurred by the BDMs and further enhance the efficiency of ViTs, we introduce a novel ViT pruning method to reduce the dimension of the latent representation of tokens in each transformer block. Extensive experiments on multiple tracking benchmarks validate the effectiveness and generality of the proposed method and show that it achieves state-of-the-art performance. Code is released at: https://github.com/xyyang317/ABTrack.
title Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.08037