A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dwarampudi, Jagan Mohan Reddy, Purks, Jennifer L, Wong, Joshua, Hu, Renjie, Banerjee, Tania
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911417804783616
author Dwarampudi, Jagan Mohan Reddy
Purks, Jennifer L
Wong, Joshua
Hu, Renjie
Banerjee, Tania
author_facet Dwarampudi, Jagan Mohan Reddy
Purks, Jennifer L
Wong, Joshua
Hu, Renjie
Banerjee, Tania
contents We introduce a reproducible, bias-resistant machine learning framework that integrates domain-informed feature engineering, nested cross-validation, and calibrated decision-threshold optimization for small-sample neuroimaging data. Conventional cross-validation frameworks that reuse the same folds for both model selection and performance estimation yield optimistically biased results, limiting reproducibility and generalization. Demonstrated on a high-dimensional structural MRI dataset of deep brain stimulation cognitive outcomes, the framework achieved a nested-CV balanced accuracy of 0.660\,$\pm$\,0.068 using a compact, interpretable subset selected via importance-guided ranking. By combining interpretability and unbiased evaluation, this work provides a generalizable computational blueprint for reliable machine learning in data-limited biomedical domains.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02920
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data
Dwarampudi, Jagan Mohan Reddy
Purks, Jennifer L
Wong, Joshua
Hu, Renjie
Banerjee, Tania
Machine Learning
Computer Vision and Pattern Recognition
Neurons and Cognition
Quantitative Methods
We introduce a reproducible, bias-resistant machine learning framework that integrates domain-informed feature engineering, nested cross-validation, and calibrated decision-threshold optimization for small-sample neuroimaging data. Conventional cross-validation frameworks that reuse the same folds for both model selection and performance estimation yield optimistically biased results, limiting reproducibility and generalization. Demonstrated on a high-dimensional structural MRI dataset of deep brain stimulation cognitive outcomes, the framework achieved a nested-CV balanced accuracy of 0.660\,$\pm$\,0.068 using a compact, interpretable subset selected via importance-guided ranking. By combining interpretability and unbiased evaluation, this work provides a generalizable computational blueprint for reliable machine learning in data-limited biomedical domains.
title A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data
topic Machine Learning
Computer Vision and Pattern Recognition
Neurons and Cognition
Quantitative Methods
url https://arxiv.org/abs/2602.02920