Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yunusa, Haruna, Qin, Shiyin, Chukkol, Abdulrahman Hamman Adama, Yusuf, Abdulganiyu Abdu, Bello, Isah, Lawan, Adamu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916973666893824
author Yunusa, Haruna
Qin, Shiyin
Chukkol, Abdulrahman Hamman Adama
Yusuf, Abdulganiyu Abdu
Bello, Isah
Lawan, Adamu
author_facet Yunusa, Haruna
Qin, Shiyin
Chukkol, Abdulrahman Hamman Adama
Yusuf, Abdulganiyu Abdu
Bello, Isah
Lawan, Adamu
contents The hybrid of Convolutional Neural Network (CNN) and Vision Transformers (ViT) architectures has emerged as a groundbreaking approach, pushing the boundaries of computer vision (CV). This comprehensive review provides a thorough examination of the literature on state-of-the-art hybrid CNN-ViT architectures, exploring the synergies between these two approaches. The main content of this survey includes: (1) a background on the vanilla CNN and ViT, (2) systematic review of various taxonomic hybrid designs to explore the synergy achieved through merging CNNs and ViTs models, (3) comparative analysis and application task-specific synergy between different hybrid architectures, (4) challenges and future directions for hybrid models, (5) lastly, the survey concludes with a summary of key findings and recommendations. Through this exploration of hybrid CV architectures, the survey aims to serve as a guiding resource, fostering a deeper understanding of the intricate dynamics between CNNs and ViTs and their collective impact on shaping the future of CV architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02941
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
Yunusa, Haruna
Qin, Shiyin
Chukkol, Abdulrahman Hamman Adama
Yusuf, Abdulganiyu Abdu
Bello, Isah
Lawan, Adamu
Computer Vision and Pattern Recognition
Machine Learning
The hybrid of Convolutional Neural Network (CNN) and Vision Transformers (ViT) architectures has emerged as a groundbreaking approach, pushing the boundaries of computer vision (CV). This comprehensive review provides a thorough examination of the literature on state-of-the-art hybrid CNN-ViT architectures, exploring the synergies between these two approaches. The main content of this survey includes: (1) a background on the vanilla CNN and ViT, (2) systematic review of various taxonomic hybrid designs to explore the synergy achieved through merging CNNs and ViTs models, (3) comparative analysis and application task-specific synergy between different hybrid architectures, (4) challenges and future directions for hybrid models, (5) lastly, the survey concludes with a summary of key findings and recommendations. Through this exploration of hybrid CV architectures, the survey aims to serve as a guiding resource, fostering a deeper understanding of the intricate dynamics between CNNs and ViTs and their collective impact on shaping the future of CV architectures.
title Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2402.02941