Enhancing Efficiency in Vision Transformer Networks: Design Techniques and Insights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heidari, Moein, Azad, Reza, Kolahi, Sina Ghorbani, Arimond, René, Niggemeier, Leon, Sulaiman, Alaa, Bozorgpour, Afshin, Aghdam, Ehsan Khodapanah, Kazerouni, Amirhossein, Hacihaliloglu, Ilker, Merhof, Dorit
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917625259360256
author Heidari, Moein
Azad, Reza
Kolahi, Sina Ghorbani
Arimond, René
Niggemeier, Leon
Sulaiman, Alaa
Bozorgpour, Afshin
Aghdam, Ehsan Khodapanah
Kazerouni, Amirhossein
Hacihaliloglu, Ilker
Merhof, Dorit
author_facet Heidari, Moein
Azad, Reza
Kolahi, Sina Ghorbani
Arimond, René
Niggemeier, Leon
Sulaiman, Alaa
Bozorgpour, Afshin
Aghdam, Ehsan Khodapanah
Kazerouni, Amirhossein
Hacihaliloglu, Ilker
Merhof, Dorit
contents Intrigued by the inherent ability of the human visual system to identify salient regions in complex scenes, attention mechanisms have been seamlessly integrated into various Computer Vision (CV) tasks. Building upon this paradigm, Vision Transformer (ViT) networks exploit attention mechanisms for improved efficiency. This review navigates the landscape of redesigned attention mechanisms within ViTs, aiming to enhance their performance. This paper provides a comprehensive exploration of techniques and insights for designing attention mechanisms, systematically reviewing recent literature in the field of CV. This survey begins with an introduction to the theoretical foundations and fundamental concepts underlying attention mechanisms. We then present a systematic taxonomy of various attention mechanisms within ViTs, employing redesigned approaches. A multi-perspective categorization is proposed based on their application, objectives, and the type of attention applied. The analysis includes an exploration of the novelty, strengths, weaknesses, and an in-depth evaluation of the different proposed strategies. This culminates in the development of taxonomies that highlight key properties and contributions. Finally, we gather the reviewed studies along with their available open-source implementations at our \href{https://github.com/mindflow-institue/Awesome-Attention-Mechanism-in-Medical-Imaging}{GitHub}\footnote{\url{https://github.com/xmindflow/Awesome-Attention-Mechanism-in-Medical-Imaging}}. We aim to regularly update it with the most recent relevant papers.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19882
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Efficiency in Vision Transformer Networks: Design Techniques and Insights
Heidari, Moein
Azad, Reza
Kolahi, Sina Ghorbani
Arimond, René
Niggemeier, Leon
Sulaiman, Alaa
Bozorgpour, Afshin
Aghdam, Ehsan Khodapanah
Kazerouni, Amirhossein
Hacihaliloglu, Ilker
Merhof, Dorit
Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
Intrigued by the inherent ability of the human visual system to identify salient regions in complex scenes, attention mechanisms have been seamlessly integrated into various Computer Vision (CV) tasks. Building upon this paradigm, Vision Transformer (ViT) networks exploit attention mechanisms for improved efficiency. This review navigates the landscape of redesigned attention mechanisms within ViTs, aiming to enhance their performance. This paper provides a comprehensive exploration of techniques and insights for designing attention mechanisms, systematically reviewing recent literature in the field of CV. This survey begins with an introduction to the theoretical foundations and fundamental concepts underlying attention mechanisms. We then present a systematic taxonomy of various attention mechanisms within ViTs, employing redesigned approaches. A multi-perspective categorization is proposed based on their application, objectives, and the type of attention applied. The analysis includes an exploration of the novelty, strengths, weaknesses, and an in-depth evaluation of the different proposed strategies. This culminates in the development of taxonomies that highlight key properties and contributions. Finally, we gather the reviewed studies along with their available open-source implementations at our \href{https://github.com/mindflow-institue/Awesome-Attention-Mechanism-in-Medical-Imaging}{GitHub}\footnote{\url{https://github.com/xmindflow/Awesome-Attention-Mechanism-in-Medical-Imaging}}. We aim to regularly update it with the most recent relevant papers.
title Enhancing Efficiency in Vision Transformer Networks: Design Techniques and Insights
topic Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.19882