Context-Aware Token Selection and Packing for Enhanced Vision Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tianyi, Li, Baoxin, Seo, Jae-sun, Cao, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912102202998784
author Zhang, Tianyi
Li, Baoxin
Seo, Jae-sun
Cao, Yu
author_facet Zhang, Tianyi
Li, Baoxin
Seo, Jae-sun
Cao, Yu
contents In recent years, the long-range attention mechanism of vision transformers has driven significant performance breakthroughs across various computer vision tasks. However, the traditional self-attention mechanism, which processes both informative and non-informative tokens, suffers from inefficiency and inaccuracies. While sparse attention mechanisms have been introduced to mitigate these issues by pruning tokens involved in attention, they often lack context-awareness and intelligence. These mechanisms frequently apply a uniform token selection strategy across different inputs for batch training or optimize efficiency only for the inference stage. To overcome these challenges, we propose a novel algorithm: Select and Pack Attention (SPA). SPA dynamically selects informative tokens using a low-cost gating layer supervised by selection labels and packs these tokens into new batches, enabling a variable number of tokens to be used in parallelized GPU batch training and inference. Extensive experiments across diverse datasets and computer vision tasks demonstrate that SPA delivers superior performance and efficiency, including a 0.6 mAP improvement in object detection and a 16.4% reduction in computational costs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Context-Aware Token Selection and Packing for Enhanced Vision Transformer
Zhang, Tianyi
Li, Baoxin
Seo, Jae-sun
Cao, Yu
Computer Vision and Pattern Recognition
In recent years, the long-range attention mechanism of vision transformers has driven significant performance breakthroughs across various computer vision tasks. However, the traditional self-attention mechanism, which processes both informative and non-informative tokens, suffers from inefficiency and inaccuracies. While sparse attention mechanisms have been introduced to mitigate these issues by pruning tokens involved in attention, they often lack context-awareness and intelligence. These mechanisms frequently apply a uniform token selection strategy across different inputs for batch training or optimize efficiency only for the inference stage. To overcome these challenges, we propose a novel algorithm: Select and Pack Attention (SPA). SPA dynamically selects informative tokens using a low-cost gating layer supervised by selection labels and packs these tokens into new batches, enabling a variable number of tokens to be used in parallelized GPU batch training and inference. Extensive experiments across diverse datasets and computer vision tasks demonstrate that SPA delivers superior performance and efficiency, including a 0.6 mAP improvement in object detection and a 16.4% reduction in computational costs.
title Context-Aware Token Selection and Packing for Enhanced Vision Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.23608