Handcrafted vs. Deep Radiomics vs. Fusion vs. Deep Learning: A Comprehensive Review of Machine Learning -Based Cancer Outcome Prediction in PET and SPECT Imaging

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Salmanpour, Mohammad R., Mehrnia, Somayeh Sadat, Ghandilu, Sajad Jabarzadeh, Falahati, Sonya, Taeb, Shahram, Mousavi, Ghazal, Maghsoudi, Mehdi, Shariftabrizi, Ahmad, Hacihaliloglu, Ilker, Rahmim, Arman
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916917224144896
author Salmanpour, Mohammad R.
Mehrnia, Somayeh Sadat
Ghandilu, Sajad Jabarzadeh
Falahati, Sonya
Taeb, Shahram
Mousavi, Ghazal
Maghsoudi, Mehdi
Shariftabrizi, Ahmad
Hacihaliloglu, Ilker
Rahmim, Arman
author_facet Salmanpour, Mohammad R.
Mehrnia, Somayeh Sadat
Ghandilu, Sajad Jabarzadeh
Falahati, Sonya
Taeb, Shahram
Mousavi, Ghazal
Maghsoudi, Mehdi
Shariftabrizi, Ahmad
Hacihaliloglu, Ilker
Rahmim, Arman
contents Machine learning (ML), including deep learning (DL) and radiomics-based methods, is increasingly used for cancer outcome prediction with PET and SPECT imaging. However, the comparative performance of handcrafted radiomics features (HRF), deep radiomics features (DRF), DL models, and hybrid fusion approaches remains inconsistent across clinical applications. This systematic review analyzed 226 studies published from 2020 to 2025 that applied ML to PET or SPECT imaging for outcome prediction. Each study was evaluated using a 59-item framework covering dataset construction, feature extraction, validation methods, interpretability, and risk of bias. We extracted key details including model type, cancer site, imaging modality, and performance metrics such as accuracy and area under the curve (AUC). PET-based studies (95%) generally outperformed those using SPECT, likely due to higher spatial resolution and sensitivity. DRF models achieved the highest mean accuracy (0.862), while fusion models yielded the highest AUC (0.861). ANOVA confirmed significant differences in performance (accuracy: p=0.0006, AUC: p=0.0027). Common limitations included inadequate handling of class imbalance (59%), missing data (29%), and low population diversity (19%). Only 48% of studies adhered to IBSI standards. These findings highlight the need for standardized pipelines, improved data quality, and explainable AI to support clinical integration.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Handcrafted vs. Deep Radiomics vs. Fusion vs. Deep Learning: A Comprehensive Review of Machine Learning -Based Cancer Outcome Prediction in PET and SPECT Imaging
Salmanpour, Mohammad R.
Mehrnia, Somayeh Sadat
Ghandilu, Sajad Jabarzadeh
Falahati, Sonya
Taeb, Shahram
Mousavi, Ghazal
Maghsoudi, Mehdi
Shariftabrizi, Ahmad
Hacihaliloglu, Ilker
Rahmim, Arman
Medical Physics
Computer Vision and Pattern Recognition
14J60 (Primary) 14F05, 14J26 (Secondary)
F.2.2; I.2.7
Machine learning (ML), including deep learning (DL) and radiomics-based methods, is increasingly used for cancer outcome prediction with PET and SPECT imaging. However, the comparative performance of handcrafted radiomics features (HRF), deep radiomics features (DRF), DL models, and hybrid fusion approaches remains inconsistent across clinical applications. This systematic review analyzed 226 studies published from 2020 to 2025 that applied ML to PET or SPECT imaging for outcome prediction. Each study was evaluated using a 59-item framework covering dataset construction, feature extraction, validation methods, interpretability, and risk of bias. We extracted key details including model type, cancer site, imaging modality, and performance metrics such as accuracy and area under the curve (AUC). PET-based studies (95%) generally outperformed those using SPECT, likely due to higher spatial resolution and sensitivity. DRF models achieved the highest mean accuracy (0.862), while fusion models yielded the highest AUC (0.861). ANOVA confirmed significant differences in performance (accuracy: p=0.0006, AUC: p=0.0027). Common limitations included inadequate handling of class imbalance (59%), missing data (29%), and low population diversity (19%). Only 48% of studies adhered to IBSI standards. These findings highlight the need for standardized pipelines, improved data quality, and explainable AI to support clinical integration.
title Handcrafted vs. Deep Radiomics vs. Fusion vs. Deep Learning: A Comprehensive Review of Machine Learning -Based Cancer Outcome Prediction in PET and SPECT Imaging
topic Medical Physics
Computer Vision and Pattern Recognition
14J60 (Primary) 14F05, 14J26 (Secondary)
F.2.2; I.2.7
url https://arxiv.org/abs/2507.16065