Diabetes Prediction Modeling Using XGBoost Ensemble Learning with SHAP Explainability: A Machine Learning Approach on the Pima Indians Diabetes Dataset

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Khan Gulrez Shagufa Fazal Ahmed
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902299848212480
author Khan Gulrez Shagufa Fazal Ahmed
author_facet Khan Gulrez Shagufa Fazal Ahmed
contents <p>Diabetes mellitus is a chronic metabolic disorder affecting over 537 million adults worldwide. This study presents a complete end-to-end machine learning pipeline for binary classification of diabetes status using the Pima Indians Diabetes Dataset (n=768). The pipeline integrates systematic data cleaning, group-median imputation, IQR-based outlier clipping, and six engineered interaction features. An XGBoost classifier was trained with 300 estimators, class-weighted loss, and L1/L2 regularization. Cross-validation was performed using a scikit-learn Pipeline to prevent data leakage. The model achieved accuracy of 87.0%, recall of 85.2%, precision of 79.3%, F1 score of 82.1%, and ROC-AUC of 94.7%. Five-fold CV AUC was 0.944 (SD=0.014). SHAP analysis identified Glucose, Glucose x BMI interaction, and BMI as the three most impactful predictors. Source code, trained model artifacts, and figures are publicly available on GitHub (https://github.com/randomthingsonlineatsk-cloud/diabetes-xgboost-prediction) and archived on Zenodo (DOI: 10.5281/zenodo.20332710).</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20336854
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Diabetes Prediction Modeling Using XGBoost Ensemble Learning with SHAP Explainability: A Machine Learning Approach on the Pima Indians Diabetes Dataset
Khan Gulrez Shagufa Fazal Ahmed
Diabetes prediction
XGBoost
Ensemble learning
SHAP explainability
Feature engineering
Machine learning
Machine Learning
Machine learning
Healthcare AI
Binary classification
Pima Indians dataset
Clinical decision support
Pharmacoinformatics
Precision Medicine
Artificial intelligence
Latent Autoimmune Diabetes in Adults/diagnosis
<p>Diabetes mellitus is a chronic metabolic disorder affecting over 537 million adults worldwide. This study presents a complete end-to-end machine learning pipeline for binary classification of diabetes status using the Pima Indians Diabetes Dataset (n=768). The pipeline integrates systematic data cleaning, group-median imputation, IQR-based outlier clipping, and six engineered interaction features. An XGBoost classifier was trained with 300 estimators, class-weighted loss, and L1/L2 regularization. Cross-validation was performed using a scikit-learn Pipeline to prevent data leakage. The model achieved accuracy of 87.0%, recall of 85.2%, precision of 79.3%, F1 score of 82.1%, and ROC-AUC of 94.7%. Five-fold CV AUC was 0.944 (SD=0.014). SHAP analysis identified Glucose, Glucose x BMI interaction, and BMI as the three most impactful predictors. Source code, trained model artifacts, and figures are publicly available on GitHub (https://github.com/randomthingsonlineatsk-cloud/diabetes-xgboost-prediction) and archived on Zenodo (DOI: 10.5281/zenodo.20332710).</p>
title Diabetes Prediction Modeling Using XGBoost Ensemble Learning with SHAP Explainability: A Machine Learning Approach on the Pima Indians Diabetes Dataset
topic Diabetes prediction
XGBoost
Ensemble learning
SHAP explainability
Feature engineering
Machine learning
Machine Learning
Machine learning
Healthcare AI
Binary classification
Pima Indians dataset
Clinical decision support
Pharmacoinformatics
Precision Medicine
Artificial intelligence
Latent Autoimmune Diabetes in Adults/diagnosis
url https://doi.org/10.5281/zenodo.20336854