CNN-ViT Fusion with Adaptive Attention Gate for Brain Tumor MRI Classification: A Hybrid Deep Learning Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hasnain, Syed Ibad, Faris, Muhammad, Tirmizi, Hafiza Syeda Yusra, Khowaja, Rabail, Israr, Hafsa
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911622610550784
author Hasnain, Syed Ibad
Faris, Muhammad
Tirmizi, Hafiza Syeda Yusra
Khowaja, Rabail
Israr, Hafsa
author_facet Hasnain, Syed Ibad
Faris, Muhammad
Tirmizi, Hafiza Syeda Yusra
Khowaja, Rabail
Israr, Hafsa
contents Early detection and classifying brain tumors using Magnetic Resonance Imaging (MRI) images is highly important but difficult to extract in medical images. Convolutional Neural Networks (CNNs) are good at capturing both local texture and spatial information whereas Vision Transformers (ViTs) are good at capturing long-range global dependencies. We propose a new hybrid architecture that combines a SqueezeNet-style CNN branch with a MobileViT-style global transformer branch, through an Adaptive Attention Gate mechanism, in this paper. The gate learns dynamically per-sample, per-feature weights to weight the contribution of each branch, allowing context-sensitive merging of local and global representations. The proposed model has a test accuracy of 97.60, a precision of 97.30, a recall of 97.50, an F1-score of 97.40, and a macro-average area under the curve (AUC) of 0.9946 with a trained and evaluated on the Brain Tumor MRI Dataset (Kaggle). These scores are higher than single CNN and ViT baselines, and current competitive fusion methods, showing that dynamic feature weighting is an effective way to classify medical images.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23137
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CNN-ViT Fusion with Adaptive Attention Gate for Brain Tumor MRI Classification: A Hybrid Deep Learning Model
Hasnain, Syed Ibad
Faris, Muhammad
Tirmizi, Hafiza Syeda Yusra
Khowaja, Rabail
Israr, Hafsa
Computer Vision and Pattern Recognition
Artificial Intelligence
Quantitative Methods
Early detection and classifying brain tumors using Magnetic Resonance Imaging (MRI) images is highly important but difficult to extract in medical images. Convolutional Neural Networks (CNNs) are good at capturing both local texture and spatial information whereas Vision Transformers (ViTs) are good at capturing long-range global dependencies. We propose a new hybrid architecture that combines a SqueezeNet-style CNN branch with a MobileViT-style global transformer branch, through an Adaptive Attention Gate mechanism, in this paper. The gate learns dynamically per-sample, per-feature weights to weight the contribution of each branch, allowing context-sensitive merging of local and global representations. The proposed model has a test accuracy of 97.60, a precision of 97.30, a recall of 97.50, an F1-score of 97.40, and a macro-average area under the curve (AUC) of 0.9946 with a trained and evaluated on the Brain Tumor MRI Dataset (Kaggle). These scores are higher than single CNN and ViT baselines, and current competitive fusion methods, showing that dynamic feature weighting is an effective way to classify medical images.
title CNN-ViT Fusion with Adaptive Attention Gate for Brain Tumor MRI Classification: A Hybrid Deep Learning Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Quantitative Methods
url https://arxiv.org/abs/2604.23137