ConcatPlexer: Additional Dim1 Batching for Faster ViTs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Han, Donghoon, Seo, Seunghyeon, Jeon, Donghyeon, Jang, Jiho, Kong, Chaerin, Kwak, Nojun
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910311743750144
author Han, Donghoon
Seo, Seunghyeon
Jeon, Donghyeon
Jang, Jiho
Kong, Chaerin
Kwak, Nojun
author_facet Han, Donghoon
Seo, Seunghyeon
Jeon, Donghyeon
Jang, Jiho
Kong, Chaerin
Kwak, Nojun
contents Transformers have demonstrated tremendous success not only in the natural language processing (NLP) domain but also the field of computer vision, igniting various creative approaches and applications. Yet, the superior performance and modeling flexibility of transformers came with a severe increase in computation costs, and hence several works have proposed methods to reduce this burden. Inspired by a cost-cutting method originally proposed for language models, Data Multiplexing (DataMUX), we propose a novel approach for efficient visual recognition that employs additional dim1 batching (i.e., concatenation) that greatly improves the throughput with little compromise in the accuracy. We first introduce a naive adaptation of DataMux for vision models, Image Multiplexer, and devise novel components to overcome its weaknesses, rendering our final model, ConcatPlexer, at the sweet spot between inference speed and accuracy. The ConcatPlexer was trained on ImageNet1K and CIFAR100 dataset and it achieved 23.5% less GFLOPs than ViT-B/16 with 69.5% and 83.4% validation accuracy, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2308_11199
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ConcatPlexer: Additional Dim1 Batching for Faster ViTs
Han, Donghoon
Seo, Seunghyeon
Jeon, Donghyeon
Jang, Jiho
Kong, Chaerin
Kwak, Nojun
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Transformers have demonstrated tremendous success not only in the natural language processing (NLP) domain but also the field of computer vision, igniting various creative approaches and applications. Yet, the superior performance and modeling flexibility of transformers came with a severe increase in computation costs, and hence several works have proposed methods to reduce this burden. Inspired by a cost-cutting method originally proposed for language models, Data Multiplexing (DataMUX), we propose a novel approach for efficient visual recognition that employs additional dim1 batching (i.e., concatenation) that greatly improves the throughput with little compromise in the accuracy. We first introduce a naive adaptation of DataMux for vision models, Image Multiplexer, and devise novel components to overcome its weaknesses, rendering our final model, ConcatPlexer, at the sweet spot between inference speed and accuracy. The ConcatPlexer was trained on ImageNet1K and CIFAR100 dataset and it achieved 23.5% less GFLOPs than ViT-B/16 with 69.5% and 83.4% validation accuracy, respectively.
title ConcatPlexer: Additional Dim1 Batching for Faster ViTs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2308.11199