Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gaur, Shaurya, Vitale, Michel, Hering, Alessa, Kwisthout, Johan, Jacobs, Colin, Philipp, Lena, van der Graaf, Fennie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918264416763904
author Gaur, Shaurya
Vitale, Michel
Hering, Alessa
Kwisthout, Johan
Jacobs, Colin
Philipp, Lena
van der Graaf, Fennie
author_facet Gaur, Shaurya
Vitale, Michel
Hering, Alessa
Kwisthout, Johan
Jacobs, Colin
Philipp, Lena
van der Graaf, Fennie
contents Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p < .001). At 90% specificity, Venkadesh21 showed lower sensitivity for Black (0.39, 95% CI: 0.23, 0.59) than White participants (0.69, 95% CI: 0.65, 0.73). These differences were not explained by available clinical confounders and thus may be classified as unfair biases according to JustEFAB. Our findings highlight the importance of improving and monitoring model performance across underrepresented subgroups, and further research on algorithmic fairness, in lung cancer screening.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening
Gaur, Shaurya
Vitale, Michel
Hering, Alessa
Kwisthout, Johan
Jacobs, Colin
Philipp, Lena
van der Graaf, Fennie
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
Image and Video Processing
Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p < .001). At 90% specificity, Venkadesh21 showed lower sensitivity for Black (0.39, 95% CI: 0.23, 0.59) than White participants (0.69, 95% CI: 0.65, 0.73). These differences were not explained by available clinical confounders and thus may be classified as unfair biases according to JustEFAB. Our findings highlight the importance of improving and monitoring model performance across underrepresented subgroups, and further research on algorithmic fairness, in lung cancer screening.
title Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
Image and Video Processing
url https://arxiv.org/abs/2512.22242