Ensemble Deep Learning for Medical Image Classification Across Diverse Modalities: A Multi-Architecture Evaluation With Uncertainty Quantification on MedMNIST

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Sanap, Hrushikesh
Format: Recurso digital
Language:English
Published: Zenodo 2026
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866902143466733568
author Sanap, Hrushikesh
author_facet Sanap, Hrushikesh
contents <p dir="ltr">Medical image classification still struggles with three major issues: limited training data, severe class imbalance, and fuzzy decision boundaries between disease categories. Deep learning models now perform as well as human experts in  many tasks, but there’s been surprisingly little work on how best to combine different architectures. In this study, I evaluate ensemble learning across four datasets from the MedMNIST v2 collection - BloodMNIST, BreastMNIST, DermaMNIST, and OrganAMNIST each representing different clinical imaging challenges. I built an ensemble using four modern architectures: ConvNeXt-Base, Vision Transformer (ViT-Base), EfficientNetV2-M, and InceptionResNetV2.</p> <p dir="ltr">The results show that modern backbones consistently beat the official ResNet baselines on every task. More interestingly, I discovered what I call “Validation Starvation”, a critical threshold that determines which ensemble method works best. When there’s enough validation data, Rigorous Stacking (a meta-learning approach) wins by learning to fix systematic errors between models. This delivered state-of-the-art accuracy on BloodMNIST (99.33%) and BreastMNIST (93.59%, which is +3.5% over baseline). But when classes are extremely imbalanced or rare, simple Soft Voting actually works better. it achieved state-of-the-art on DermaMNIST (91.97%), a massive 15.2% jump over previous benchmarks. I also validated that these ensembles are safe for clinical use through calibration analysis and entropy-based uncertainty quantification, showing they can reliably flag ambiguous cases for human review. These findings give us a practical, reproducible strategy for deploying high-performance diagnostic AI in resource-limited medical settings.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18813204
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Ensemble Deep Learning for Medical Image Classification Across Diverse Modalities: A Multi-Architecture Evaluation With Uncertainty Quantification on MedMNIST
Sanap, Hrushikesh
<p dir="ltr">Medical image classification still struggles with three major issues: limited training data, severe class imbalance, and fuzzy decision boundaries between disease categories. Deep learning models now perform as well as human experts in  many tasks, but there’s been surprisingly little work on how best to combine different architectures. In this study, I evaluate ensemble learning across four datasets from the MedMNIST v2 collection - BloodMNIST, BreastMNIST, DermaMNIST, and OrganAMNIST each representing different clinical imaging challenges. I built an ensemble using four modern architectures: ConvNeXt-Base, Vision Transformer (ViT-Base), EfficientNetV2-M, and InceptionResNetV2.</p> <p dir="ltr">The results show that modern backbones consistently beat the official ResNet baselines on every task. More interestingly, I discovered what I call “Validation Starvation”, a critical threshold that determines which ensemble method works best. When there’s enough validation data, Rigorous Stacking (a meta-learning approach) wins by learning to fix systematic errors between models. This delivered state-of-the-art accuracy on BloodMNIST (99.33%) and BreastMNIST (93.59%, which is +3.5% over baseline). But when classes are extremely imbalanced or rare, simple Soft Voting actually works better. it achieved state-of-the-art on DermaMNIST (91.97%), a massive 15.2% jump over previous benchmarks. I also validated that these ensembles are safe for clinical use through calibration analysis and entropy-based uncertainty quantification, showing they can reliably flag ambiguous cases for human review. These findings give us a practical, reproducible strategy for deploying high-performance diagnostic AI in resource-limited medical settings.</p>
title Ensemble Deep Learning for Medical Image Classification Across Diverse Modalities: A Multi-Architecture Evaluation With Uncertainty Quantification on MedMNIST
url https://doi.org/10.5281/zenodo.18813204