Training-Free Acceleration of ViTs with Delayed Spatial Merging

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Heo, Jung Hwan, Azizi, Seyedarmin, Fayyazi, Arash, Pedram, Massoud
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914853524865024
author Heo, Jung Hwan
Azizi, Seyedarmin
Fayyazi, Arash
Pedram, Massoud
author_facet Heo, Jung Hwan
Azizi, Seyedarmin
Fayyazi, Arash
Pedram, Massoud
contents Token merging has emerged as a new paradigm that can accelerate the inference of Vision Transformers (ViTs) without any retraining or fine-tuning. To push the frontier of training-free acceleration in ViTs, we improve token merging by adding the perspectives of 1) activation outliers and 2) hierarchical representations. Through a careful analysis of the attention behavior in ViTs, we characterize a delayed onset of the convergent attention phenomenon, which makes token merging undesirable in the bottom blocks of ViTs. Moreover, we augment token merging with a hierarchical processing scheme to capture multi-scale redundancy between visual tokens. Combining these two insights, we build a unified inference framework called DSM: Delayed Spatial Merging. We extensively evaluate DSM on various ViT model scales (Tiny to Huge) and tasks (ImageNet-1k and transfer learning), achieving up to 1.8$\times$ FLOP reduction and 1.6$\times$ throughput speedup at a negligible loss while being two orders of magnitude faster than existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2303_02331
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Training-Free Acceleration of ViTs with Delayed Spatial Merging
Heo, Jung Hwan
Azizi, Seyedarmin
Fayyazi, Arash
Pedram, Massoud
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Token merging has emerged as a new paradigm that can accelerate the inference of Vision Transformers (ViTs) without any retraining or fine-tuning. To push the frontier of training-free acceleration in ViTs, we improve token merging by adding the perspectives of 1) activation outliers and 2) hierarchical representations. Through a careful analysis of the attention behavior in ViTs, we characterize a delayed onset of the convergent attention phenomenon, which makes token merging undesirable in the bottom blocks of ViTs. Moreover, we augment token merging with a hierarchical processing scheme to capture multi-scale redundancy between visual tokens. Combining these two insights, we build a unified inference framework called DSM: Delayed Spatial Merging. We extensively evaluate DSM on various ViT model scales (Tiny to Huge) and tasks (ImageNet-1k and transfer learning), achieving up to 1.8$\times$ FLOP reduction and 1.6$\times$ throughput speedup at a negligible loss while being two orders of magnitude faster than existing methods.
title Training-Free Acceleration of ViTs with Delayed Spatial Merging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2303.02331