A survey of the Vision Transformers and their CNN-Transformer based Variants
Fuente:
arXiv
Saved in:
| Main Authors: | Khan, Asifullah, Rauf, Zunaira, Sohail, Anabia, Rehman, Abdul, Asif, Hifsa, Asif, Aqsa, Farooq, Umair |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Channel Boosted CNN-Transformer-based Multi-Level and Multi-Scale Nuclei Segmentation
by: Rauf, Zunaira, et al.
Published: (2024)
by: Rauf, Zunaira, et al.
Published: (2024)
MaxViT-UNet: Multi-Axis Attention for Medical Image Segmentation
by: Khan, Abdul Rehman, et al.
Published: (2023)
by: Khan, Abdul Rehman, et al.
Published: (2023)
A Survey on Self-supervised Contrastive Learning for Multimodal Text-Image Analysis
by: Khan, Asifullah, et al.
Published: (2025)
by: Khan, Asifullah, et al.
Published: (2025)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
by: Khan, Asifullah, et al.
Published: (2024)
by: Khan, Asifullah, et al.
Published: (2024)
InfiltrNet: Dual-Branch CNN-Transformer Architecture for Brain Tumor Infiltration Risk Prediction
by: Hossain, S M Asif, et al.
Published: (2026)
by: Hossain, S M Asif, et al.
Published: (2026)
Event Transformer
by: Jiang, Bin, et al.
Published: (2022)
by: Jiang, Bin, et al.
Published: (2022)
Transform-Dependent Adversarial Attacks
by: Tan, Yaoteng, et al.
Published: (2024)
by: Tan, Yaoteng, et al.
Published: (2024)
Comparative Study of CNN Architectures for Binary Classification of Horses and Motorcycles in the VOC 2008 Dataset
by: Shaikh, Muhammad Annas, et al.
Published: (2025)
by: Shaikh, Muhammad Annas, et al.
Published: (2025)
VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
by: Phukan, Arpan, et al.
Published: (2025)
by: Phukan, Arpan, et al.
Published: (2025)
Dental Panoramic Radiograph Analysis Using YOLO26 From Tooth Detection to Disease Diagnosis
by: Asif, Khawaja Azfar, et al.
Published: (2026)
by: Asif, Khawaja Azfar, et al.
Published: (2026)
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation
by: Reza, Md Kaykobad, et al.
Published: (2023)
by: Reza, Md Kaykobad, et al.
Published: (2023)
GlobalWasteData: A Large-Scale, Integrated Dataset for Robust Waste Classification and Environmental Monitoring
by: Ijaz, Misbah, et al.
Published: (2026)
by: Ijaz, Misbah, et al.
Published: (2026)
Adversarial robustness of VAEs through the lens of local geometry
by: Khan, Asif, et al.
Published: (2022)
by: Khan, Asif, et al.
Published: (2022)
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
by: Albastaki, Shahad, et al.
Published: (2025)
by: Albastaki, Shahad, et al.
Published: (2025)
Global-Local Image Perceptual Score (GLIPS): Evaluating Photorealistic Quality of AI-Generated Images
by: Aziz, Memoona, et al.
Published: (2024)
by: Aziz, Memoona, et al.
Published: (2024)
Object Depth and Size Estimation using Stereo-vision and Integration with SLAM
by: Hamad, Layth, et al.
Published: (2024)
by: Hamad, Layth, et al.
Published: (2024)
Quality Detection of Stored Potatoes via Transfer Learning: A CNN and Vision Transformer Approach
by: Kapse, Shrikant, et al.
Published: (2026)
by: Kapse, Shrikant, et al.
Published: (2026)
UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer
by: Liu, Ji, et al.
Published: (2024)
by: Liu, Ji, et al.
Published: (2024)
fMRI-Diffusion: Generating fMRI Time Series Via a Temporal Transformer Diffusion Model for Major Depressive Disorder Diagnosis
by: Hasan, Muhammad Asif, et al.
Published: (2026)
by: Hasan, Muhammad Asif, et al.
Published: (2026)
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection
by: Khan, Sohail Ahmed, et al.
Published: (2024)
by: Khan, Sohail Ahmed, et al.
Published: (2024)
Camera Calibration through Geometric Constraints from Rotation and Projection Matrices
by: Waleed, Muhammad, et al.
Published: (2024)
by: Waleed, Muhammad, et al.
Published: (2024)
TractoTransformer: Diffusion MRI Streamline Tractography using CNN and Transformer Networks
by: Waizman, Itzik, et al.
Published: (2025)
by: Waizman, Itzik, et al.
Published: (2025)
Curriculum for Crowd Counting -- Is it Worthy?
by: Khan, Muhammad Asif, et al.
Published: (2024)
by: Khan, Muhammad Asif, et al.
Published: (2024)
Accelerating Deep Learning with Fixed Time Budget
by: Khan, Muhammad Asif, et al.
Published: (2024)
by: Khan, Muhammad Asif, et al.
Published: (2024)
Multimodal Crowd Counting with Pix2Pix GANs
by: Khan, Muhammad Asif, et al.
Published: (2024)
by: Khan, Muhammad Asif, et al.
Published: (2024)
Crowd Scene Analysis using Deep Learning Techniques
by: Asif, Muhammad Junaid
Published: (2025)
by: Asif, Muhammad Junaid
Published: (2025)
Anatomy-Guided Representation Learning Using a Transformer-Based Network for Thyroid Nodule Segmentation in Ultrasound Images
by: Farooq, Muhammad Umar, et al.
Published: (2025)
by: Farooq, Muhammad Umar, et al.
Published: (2025)
Braille to Text Translation for Bengali Language: A Geometric Approach
by: Kamal, Minhas, et al.
Published: (2020)
by: Kamal, Minhas, et al.
Published: (2020)
An analysis of vision-language models for fabric retrieval
by: Giuliari, Francesco, et al.
Published: (2025)
by: Giuliari, Francesco, et al.
Published: (2025)
Dynamic Weight Adjustment for Knowledge Distillation: Leveraging Vision Transformer for High-Accuracy Lung Cancer Detection and Real-Time Deployment
by: Khan, Saif Ur Rehman, et al.
Published: (2025)
by: Khan, Saif Ur Rehman, et al.
Published: (2025)
A Computer Vision Hybrid Approach: CNN and Transformer Models for Accurate Alzheimer's Detection from Brain MRI Scans
by: Hoque, Md Mahmudul, et al.
Published: (2026)
by: Hoque, Md Mahmudul, et al.
Published: (2026)
Translating Imaging to Genomics: Leveraging Transformers for Predictive Modeling
by: Farooq, Aiman, et al.
Published: (2024)
by: Farooq, Aiman, et al.
Published: (2024)
Dual-Encoder Transformer-Based Multimodal Learning for Ischemic Stroke Lesion Segmentation Using Diffusion MRI
by: Usman, Muhammad, et al.
Published: (2025)
by: Usman, Muhammad, et al.
Published: (2025)
Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models
by: Monon, Mashrafi, et al.
Published: (2026)
by: Monon, Mashrafi, et al.
Published: (2026)
An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images
by: Newaz, Asif, et al.
Published: (2025)
by: Newaz, Asif, et al.
Published: (2025)
Data Efficient Contrastive Learning in Histopathology using Active Sampling
by: Reasat, Tahsin, et al.
Published: (2023)
by: Reasat, Tahsin, et al.
Published: (2023)
CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation
by: Wu, Lanhu, et al.
Published: (2024)
by: Wu, Lanhu, et al.
Published: (2024)
AutoGen Driven Multi Agent Framework for Iterative Crime Data Analysis and Prediction
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
by: Fatima, Syeda Kisaa, et al.
Published: (2025)
TESL-Net: A Transformer-Enhanced CNN for Accurate Skin Lesion Segmentation
by: Iqbal, Shahzaib, et al.
Published: (2024)
by: Iqbal, Shahzaib, et al.
Published: (2024)
Benchmarking CNN-based Models against Transformer-based Models for Abdominal Multi-Organ Segmentation on the RATIC Dataset
by: Bayer, Lukas, et al.
Published: (2026)
by: Bayer, Lukas, et al.
Published: (2026)
Similar Items
-
Channel Boosted CNN-Transformer-based Multi-Level and Multi-Scale Nuclei Segmentation
by: Rauf, Zunaira, et al.
Published: (2024) -
MaxViT-UNet: Multi-Axis Attention for Medical Image Segmentation
by: Khan, Abdul Rehman, et al.
Published: (2023) -
A Survey on Self-supervised Contrastive Learning for Multimodal Text-Image Analysis
by: Khan, Asifullah, et al.
Published: (2025) -
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
by: Khan, Asifullah, et al.
Published: (2024) -
InfiltrNet: Dual-Branch CNN-Transformer Architecture for Brain Tumor Infiltration Risk Prediction
by: Hossain, S M Asif, et al.
Published: (2026)