3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Weicheng, Huang, Haoxu, Tang, Huanze, Musthyala, Rushabh, Yu, Boyang, Chen, Long, Vega, Emilio, O'Donnell, Thomas, Dehkharghani, Seena, Frontera, Jennifer A., Masurkar, Arjun V., Melmed, Kara, Razavian, Narges
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908984242339840
author Zhu, Weicheng
Huang, Haoxu
Tang, Huanze
Musthyala, Rushabh
Yu, Boyang
Chen, Long
Vega, Emilio
O'Donnell, Thomas
Dehkharghani, Seena
Frontera, Jennifer A.
Masurkar, Arjun V.
Melmed, Kara
Razavian, Narges
author_facet Zhu, Weicheng
Huang, Haoxu
Tang, Huanze
Musthyala, Rushabh
Yu, Boyang
Chen, Long
Vega, Emilio
O'Donnell, Thomas
Dehkharghani, Seena
Frontera, Jennifer A.
Masurkar, Arjun V.
Melmed, Kara
Razavian, Narges
contents Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing pathology of the brain, skull, and cerebrovascular system. It is commonly the first-line imaging in neurologic emergencies given its rapidity of image acquisition, safety, cost, and ubiquity. Deep learning models may facilitate detection of a wide range of diseases. However, the scarcity of high-quality labels and annotations, particularly among less common conditions, significantly hinders the development of powerful models. To address this challenge, we introduce FM-CT: a Foundation Model for Head CT for generalizable disease detection, trained using self-supervised learning. Our approach pre-trains a deep learning model on a large, diverse dataset of 361,663 non-contrast 3D head CT scans without the need for manual annotations, enabling the model to learn robust, generalizable features. To investigate the potential of self-supervised learning in head CT, we employed both discrimination with self-distillation and masked image modeling, and we construct our model in 3D rather than at the slice level (2D) to exploit the structure of head CT scans more comprehensively and efficiently. The model's downstream classification performance is evaluated using internal and three external datasets, encompassing both in-distribution (ID) and out-of-distribution (OOD) data. Our results demonstrate that the self-supervised foundation model significantly improves performance on downstream diagnostic tasks compared to models trained from scratch and previous 3D CT foundation models on scarce annotated datasets. This work highlights the effectiveness of self-supervised learning in medical imaging and sets a new benchmark for head CT image analysis in 3D, enabling broader use of artificial intelligence for head CT-based diagnosis.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02779
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography
Zhu, Weicheng
Huang, Haoxu
Tang, Huanze
Musthyala, Rushabh
Yu, Boyang
Chen, Long
Vega, Emilio
O'Donnell, Thomas
Dehkharghani, Seena
Frontera, Jennifer A.
Masurkar, Arjun V.
Melmed, Kara
Razavian, Narges
Computer Vision and Pattern Recognition
Artificial Intelligence
Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing pathology of the brain, skull, and cerebrovascular system. It is commonly the first-line imaging in neurologic emergencies given its rapidity of image acquisition, safety, cost, and ubiquity. Deep learning models may facilitate detection of a wide range of diseases. However, the scarcity of high-quality labels and annotations, particularly among less common conditions, significantly hinders the development of powerful models. To address this challenge, we introduce FM-CT: a Foundation Model for Head CT for generalizable disease detection, trained using self-supervised learning. Our approach pre-trains a deep learning model on a large, diverse dataset of 361,663 non-contrast 3D head CT scans without the need for manual annotations, enabling the model to learn robust, generalizable features. To investigate the potential of self-supervised learning in head CT, we employed both discrimination with self-distillation and masked image modeling, and we construct our model in 3D rather than at the slice level (2D) to exploit the structure of head CT scans more comprehensively and efficiently. The model's downstream classification performance is evaluated using internal and three external datasets, encompassing both in-distribution (ID) and out-of-distribution (OOD) data. Our results demonstrate that the self-supervised foundation model significantly improves performance on downstream diagnostic tasks compared to models trained from scratch and previous 3D CT foundation models on scarce annotated datasets. This work highlights the effectiveness of self-supervised learning in medical imaging and sets a new benchmark for head CT image analysis in 3D, enabling broader use of artificial intelligence for head CT-based diagnosis.
title 3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.02779