CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Basnet, Prashant Singh, Chitrakar, Roshan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911147646517248
author Basnet, Prashant Singh
Chitrakar, Roshan
author_facet Basnet, Prashant Singh
Chitrakar, Roshan
contents This research explored the hybridization of CNN and ViT within a training dataset of limited size, and introduced a distinct class imbalance. The training was made from scratch with a mere focus on theoretically and experimentally exploring the architectural strengths of the proposed hybrid model. Experiments were conducted across varied data fractions with balanced and imbalanced training datasets. Comparatively, the hybrid model, complementing the strengths of CNN and ViT, achieved the highest recall of 0.9443 (50% data fraction in balanced) and consistency in F1 score around 0.85, suggesting reliability in diagnosis. Additionally, the model was successful in outperforming CNN and ViT in imbalanced datasets. Despite its complex architecture, it required comparable training time to the transformers in all data fractions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
Basnet, Prashant Singh
Chitrakar, Roshan
Image and Video Processing
Computer Vision and Pattern Recognition
This research explored the hybridization of CNN and ViT within a training dataset of limited size, and introduced a distinct class imbalance. The training was made from scratch with a mere focus on theoretically and experimentally exploring the architectural strengths of the proposed hybrid model. Experiments were conducted across varied data fractions with balanced and imbalanced training datasets. Comparatively, the hybrid model, complementing the strengths of CNN and ViT, achieved the highest recall of 0.9443 (50% data fraction in balanced) and consistency in F1 score around 0.85, suggesting reliability in diagnosis. Additionally, the model was successful in outperforming CNN and ViT in imbalanced datasets. Despite its complex architecture, it required comparable training time to the transformers in all data fractions.
title CNN-ViT Hybrid for Pneumonia Detection: Theory and Empiric on Limited Data without Pretraining
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.08586