A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911417804783616 |
|---|---|
| author | Dwarampudi, Jagan Mohan Reddy Purks, Jennifer L Wong, Joshua Hu, Renjie Banerjee, Tania |
| author_facet | Dwarampudi, Jagan Mohan Reddy Purks, Jennifer L Wong, Joshua Hu, Renjie Banerjee, Tania |
| contents | We introduce a reproducible, bias-resistant machine learning framework that integrates domain-informed feature engineering, nested cross-validation, and calibrated decision-threshold optimization for small-sample neuroimaging data. Conventional cross-validation frameworks that reuse the same folds for both model selection and performance estimation yield optimistically biased results, limiting reproducibility and generalization. Demonstrated on a high-dimensional structural MRI dataset of deep brain stimulation cognitive outcomes, the framework achieved a nested-CV balanced accuracy of 0.660\,$\pm$\,0.068 using a compact, interpretable subset selected via importance-guided ranking. By combining interpretability and unbiased evaluation, this work provides a generalizable computational blueprint for reliable machine learning in data-limited biomedical domains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_02920 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data Dwarampudi, Jagan Mohan Reddy Purks, Jennifer L Wong, Joshua Hu, Renjie Banerjee, Tania Machine Learning Computer Vision and Pattern Recognition Neurons and Cognition Quantitative Methods We introduce a reproducible, bias-resistant machine learning framework that integrates domain-informed feature engineering, nested cross-validation, and calibrated decision-threshold optimization for small-sample neuroimaging data. Conventional cross-validation frameworks that reuse the same folds for both model selection and performance estimation yield optimistically biased results, limiting reproducibility and generalization. Demonstrated on a high-dimensional structural MRI dataset of deep brain stimulation cognitive outcomes, the framework achieved a nested-CV balanced accuracy of 0.660\,$\pm$\,0.068 using a compact, interpretable subset selected via importance-guided ranking. By combining interpretability and unbiased evaluation, this work provides a generalizable computational blueprint for reliable machine learning in data-limited biomedical domains. |
| title | A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data |
| topic | Machine Learning Computer Vision and Pattern Recognition Neurons and Cognition Quantitative Methods |
| url | https://arxiv.org/abs/2602.02920 |