Diabetes Prediction Modeling Using XGBoost Ensemble Learning with SHAP Explainability: A Machine Learning Approach on the Pima Indians Diabetes Dataset
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Lingua: | inglese |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902299848212480 |
|---|---|
| author | Khan Gulrez Shagufa Fazal Ahmed |
| author_facet | Khan Gulrez Shagufa Fazal Ahmed |
| contents | <p>Diabetes mellitus is a chronic metabolic disorder affecting over 537 million adults worldwide. This study presents a complete end-to-end machine learning pipeline for binary classification of diabetes status using the Pima Indians Diabetes Dataset (n=768). The pipeline integrates systematic data cleaning, group-median imputation, IQR-based outlier clipping, and six engineered interaction features. An XGBoost classifier was trained with 300 estimators, class-weighted loss, and L1/L2 regularization. Cross-validation was performed using a scikit-learn Pipeline to prevent data leakage. The model achieved accuracy of 87.0%, recall of 85.2%, precision of 79.3%, F1 score of 82.1%, and ROC-AUC of 94.7%. Five-fold CV AUC was 0.944 (SD=0.014). SHAP analysis identified Glucose, Glucose x BMI interaction, and BMI as the three most impactful predictors. Source code, trained model artifacts, and figures are publicly available on GitHub (https://github.com/randomthingsonlineatsk-cloud/diabetes-xgboost-prediction) and archived on Zenodo (DOI: 10.5281/zenodo.20332710).</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20336854 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Diabetes Prediction Modeling Using XGBoost Ensemble Learning with SHAP Explainability: A Machine Learning Approach on the Pima Indians Diabetes Dataset Khan Gulrez Shagufa Fazal Ahmed Diabetes prediction XGBoost Ensemble learning SHAP explainability Feature engineering Machine learning Machine Learning Machine learning Healthcare AI Binary classification Pima Indians dataset Clinical decision support Pharmacoinformatics Precision Medicine Artificial intelligence Latent Autoimmune Diabetes in Adults/diagnosis <p>Diabetes mellitus is a chronic metabolic disorder affecting over 537 million adults worldwide. This study presents a complete end-to-end machine learning pipeline for binary classification of diabetes status using the Pima Indians Diabetes Dataset (n=768). The pipeline integrates systematic data cleaning, group-median imputation, IQR-based outlier clipping, and six engineered interaction features. An XGBoost classifier was trained with 300 estimators, class-weighted loss, and L1/L2 regularization. Cross-validation was performed using a scikit-learn Pipeline to prevent data leakage. The model achieved accuracy of 87.0%, recall of 85.2%, precision of 79.3%, F1 score of 82.1%, and ROC-AUC of 94.7%. Five-fold CV AUC was 0.944 (SD=0.014). SHAP analysis identified Glucose, Glucose x BMI interaction, and BMI as the three most impactful predictors. Source code, trained model artifacts, and figures are publicly available on GitHub (https://github.com/randomthingsonlineatsk-cloud/diabetes-xgboost-prediction) and archived on Zenodo (DOI: 10.5281/zenodo.20332710).</p> |
| title | Diabetes Prediction Modeling Using XGBoost Ensemble Learning with SHAP Explainability: A Machine Learning Approach on the Pima Indians Diabetes Dataset |
| topic | Diabetes prediction XGBoost Ensemble learning SHAP explainability Feature engineering Machine learning Machine Learning Machine learning Healthcare AI Binary classification Pima Indians dataset Clinical decision support Pharmacoinformatics Precision Medicine Artificial intelligence Latent Autoimmune Diabetes in Adults/diagnosis |
| url | https://doi.org/10.5281/zenodo.20336854 |