A survey of the Vision Transformers and their CNN-Transformer based Variants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khan, Asifullah, Rauf, Zunaira, Sohail, Anabia, Rehman, Abdul, Asif, Hifsa, Asif, Aqsa, Farooq, Umair
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929438733631488
author Khan, Asifullah
Rauf, Zunaira
Sohail, Anabia
Rehman, Abdul
Asif, Hifsa
Asif, Aqsa
Farooq, Umair
author_facet Khan, Asifullah
Rauf, Zunaira
Sohail, Anabia
Rehman, Abdul
Asif, Hifsa
Asif, Aqsa
Farooq, Umair
contents Vision transformers have become popular as a possible substitute to convolutional neural networks (CNNs) for a variety of computer vision applications. These transformers, with their ability to focus on global relationships in images, offer large learning capacity. However, they may suffer from limited generalization as they do not tend to model local correlation in images. Recently, in vision transformers hybridization of both the convolution operation and self-attention mechanism has emerged, to exploit both the local and global image representations. These hybrid vision transformers, also referred to as CNN-Transformer architectures, have demonstrated remarkable results in vision applications. Given the rapidly growing number of hybrid vision transformers, it has become necessary to provide a taxonomy and explanation of these hybrid architectures. This survey presents a taxonomy of the recent vision transformer architectures and more specifically that of the hybrid vision transformers. Additionally, the key features of these architectures such as the attention mechanisms, positional embeddings, multi-scale processing, and convolution are also discussed. In contrast to the previous survey papers that are primarily focused on individual vision transformer architectures or CNNs, this survey uniquely emphasizes the emerging trend of hybrid vision transformers. By showcasing the potential of hybrid vision transformers to deliver exceptional performance across a range of computer vision tasks, this survey sheds light on the future directions of this rapidly evolving architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2305_09880
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A survey of the Vision Transformers and their CNN-Transformer based Variants
Khan, Asifullah
Rauf, Zunaira
Sohail, Anabia
Rehman, Abdul
Asif, Hifsa
Asif, Aqsa
Farooq, Umair
Computer Vision and Pattern Recognition
Vision transformers have become popular as a possible substitute to convolutional neural networks (CNNs) for a variety of computer vision applications. These transformers, with their ability to focus on global relationships in images, offer large learning capacity. However, they may suffer from limited generalization as they do not tend to model local correlation in images. Recently, in vision transformers hybridization of both the convolution operation and self-attention mechanism has emerged, to exploit both the local and global image representations. These hybrid vision transformers, also referred to as CNN-Transformer architectures, have demonstrated remarkable results in vision applications. Given the rapidly growing number of hybrid vision transformers, it has become necessary to provide a taxonomy and explanation of these hybrid architectures. This survey presents a taxonomy of the recent vision transformer architectures and more specifically that of the hybrid vision transformers. Additionally, the key features of these architectures such as the attention mechanisms, positional embeddings, multi-scale processing, and convolution are also discussed. In contrast to the previous survey papers that are primarily focused on individual vision transformer architectures or CNNs, this survey uniquely emphasizes the emerging trend of hybrid vision transformers. By showcasing the potential of hybrid vision transformers to deliver exceptional performance across a range of computer vision tasks, this survey sheds light on the future directions of this rapidly evolving architecture.
title A survey of the Vision Transformers and their CNN-Transformer based Variants
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.09880