Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gee, Leonidas, Li, Wing Yan, Sharmanska, Viktoriia, Quadrianto, Novi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908677913444352
author Gee, Leonidas
Li, Wing Yan
Sharmanska, Viktoriia
Quadrianto, Novi
author_facet Gee, Leonidas
Li, Wing Yan
Sharmanska, Viktoriia
Quadrianto, Novi
contents The cost of deploying vision transformers increasingly represents a barrier to wider industrial adoption. Existing compression techniques require additional end-to-end fine-tuning or incur a significant drawback to energy efficiency, making them ill-suited for online (real-time) inference, where a prediction is made on any new input as it comes in. We introduce the $\textbf{Visual-Word Tokenizer}$ (VWT), a training-free method for reducing energy costs while retaining performance. The VWT groups visual subwords (image patches) that are frequently used into visual words, while infrequent ones remain intact. To do so, $\textit{intra}$-image or $\textit{inter}$-image statistics are leveraged to identify similar visual concepts for sequence compression. Experimentally, we demonstrate a reduction in energy consumed of up to 47%. Comparative approaches of 8-bit quantization and token merging can lead to significantly increased energy costs (up to 500% or more). Our results indicate that VWTs are well-suited for efficient online inference with a marginal compromise on performance. The experimental code for our paper is also made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15397
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
Gee, Leonidas
Li, Wing Yan
Sharmanska, Viktoriia
Quadrianto, Novi
Computer Vision and Pattern Recognition
The cost of deploying vision transformers increasingly represents a barrier to wider industrial adoption. Existing compression techniques require additional end-to-end fine-tuning or incur a significant drawback to energy efficiency, making them ill-suited for online (real-time) inference, where a prediction is made on any new input as it comes in. We introduce the $\textbf{Visual-Word Tokenizer}$ (VWT), a training-free method for reducing energy costs while retaining performance. The VWT groups visual subwords (image patches) that are frequently used into visual words, while infrequent ones remain intact. To do so, $\textit{intra}$-image or $\textit{inter}$-image statistics are leveraged to identify similar visual concepts for sequence compression. Experimentally, we demonstrate a reduction in energy consumed of up to 47%. Comparative approaches of 8-bit quantization and token merging can lead to significantly increased energy costs (up to 500% or more). Our results indicate that VWTs are well-suited for efficient online inference with a marginal compromise on performance. The experimental code for our paper is also made publicly available.
title Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15397