SparseSwin: Swin Transformer with Sparse Transformer Block

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pinasthika, Krisna, Laksono, Blessius Sheldo Putra, Irsal, Riyandi Banovbi Putera, Shabiyya, Syifa Hukma, Yudistira, Novanto
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916149821702144
author Pinasthika, Krisna
Laksono, Blessius Sheldo Putra
Irsal, Riyandi Banovbi Putera
Shabiyya, Syifa Hukma
Yudistira, Novanto
author_facet Pinasthika, Krisna
Laksono, Blessius Sheldo Putra
Irsal, Riyandi Banovbi Putera
Shabiyya, Syifa Hukma
Yudistira, Novanto
contents Advancements in computer vision research have put transformer architecture as the state of the art in computer vision tasks. One of the known drawbacks of the transformer architecture is the high number of parameters, this can lead to a more complex and inefficient algorithm. This paper aims to reduce the number of parameters and in turn, made the transformer more efficient. We present Sparse Transformer (SparTa) Block, a modified transformer block with an addition of a sparse token converter that reduces the number of tokens used. We use the SparTa Block inside the Swin T architecture (SparseSwin) to leverage Swin capability to downsample its input and reduce the number of initial tokens to be calculated. The proposed SparseSwin model outperforms other state of the art models in image classification with an accuracy of 86.96%, 97.43%, and 85.35% on the ImageNet100, CIFAR10, and CIFAR100 datasets respectively. Despite its fewer parameters, the result highlights the potential of a transformer architecture using a sparse token converter with a limited number of tokens to optimize the use of the transformer and improve its performance.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05224
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SparseSwin: Swin Transformer with Sparse Transformer Block
Pinasthika, Krisna
Laksono, Blessius Sheldo Putra
Irsal, Riyandi Banovbi Putera
Shabiyya, Syifa Hukma
Yudistira, Novanto
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Advancements in computer vision research have put transformer architecture as the state of the art in computer vision tasks. One of the known drawbacks of the transformer architecture is the high number of parameters, this can lead to a more complex and inefficient algorithm. This paper aims to reduce the number of parameters and in turn, made the transformer more efficient. We present Sparse Transformer (SparTa) Block, a modified transformer block with an addition of a sparse token converter that reduces the number of tokens used. We use the SparTa Block inside the Swin T architecture (SparseSwin) to leverage Swin capability to downsample its input and reduce the number of initial tokens to be calculated. The proposed SparseSwin model outperforms other state of the art models in image classification with an accuracy of 86.96%, 97.43%, and 85.35% on the ImageNet100, CIFAR10, and CIFAR100 datasets respectively. Despite its fewer parameters, the result highlights the potential of a transformer architecture using a sparse token converter with a limited number of tokens to optimize the use of the transformer and improve its performance.
title SparseSwin: Swin Transformer with Sparse Transformer Block
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2309.05224