AutoML-Med: A Framework for Automated Machine Learning in Medical Tabular Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Francia, Riccardo, Leone, Maurizio, Leonardi, Giorgio, Montani, Stefania, Pennisi, Marzio, Striani, Manuel, D'Alfonso, Sandra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912518973161472
author Francia, Riccardo
Leone, Maurizio
Leonardi, Giorgio
Montani, Stefania
Pennisi, Marzio
Striani, Manuel
D'Alfonso, Sandra
author_facet Francia, Riccardo
Leone, Maurizio
Leonardi, Giorgio
Montani, Stefania
Pennisi, Marzio
Striani, Manuel
D'Alfonso, Sandra
contents Medical datasets are typically affected by issues such as missing values, class imbalance, a heterogeneous feature types, and a high number of features versus a relatively small number of samples, preventing machine learning models from obtaining proper results in classification and regression tasks. This paper introduces AutoML-Med, an Automated Machine Learning tool specifically designed to address these challenges, minimizing user intervention and identifying the optimal combination of preprocessing techniques and predictive models. AutoML-Med's architecture incorporates Latin Hypercube Sampling (LHS) for exploring preprocessing methods, trains models using selected metrics, and utilizes Partial Rank Correlation Coefficient (PRCC) for fine-tuned optimization of the most influential preprocessing steps. Experimental results demonstrate AutoML-Med's effectiveness in two different clinical settings, achieving higher balanced accuracy and sensitivity, which are crucial for identifying at-risk patients, compared to other state-of-the-art tools. AutoML-Med's ability to improve prediction results, especially in medical datasets with sparse data and class imbalance, highlights its potential to streamline Machine Learning applications in healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoML-Med: A Framework for Automated Machine Learning in Medical Tabular Data
Francia, Riccardo
Leone, Maurizio
Leonardi, Giorgio
Montani, Stefania
Pennisi, Marzio
Striani, Manuel
D'Alfonso, Sandra
Machine Learning
Artificial Intelligence
Medical datasets are typically affected by issues such as missing values, class imbalance, a heterogeneous feature types, and a high number of features versus a relatively small number of samples, preventing machine learning models from obtaining proper results in classification and regression tasks. This paper introduces AutoML-Med, an Automated Machine Learning tool specifically designed to address these challenges, minimizing user intervention and identifying the optimal combination of preprocessing techniques and predictive models. AutoML-Med's architecture incorporates Latin Hypercube Sampling (LHS) for exploring preprocessing methods, trains models using selected metrics, and utilizes Partial Rank Correlation Coefficient (PRCC) for fine-tuned optimization of the most influential preprocessing steps. Experimental results demonstrate AutoML-Med's effectiveness in two different clinical settings, achieving higher balanced accuracy and sensitivity, which are crucial for identifying at-risk patients, compared to other state-of-the-art tools. AutoML-Med's ability to improve prediction results, especially in medical datasets with sparse data and class imbalance, highlights its potential to streamline Machine Learning applications in healthcare.
title AutoML-Med: A Framework for Automated Machine Learning in Medical Tabular Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.02625