Vision Foundation Models for Computed Tomography

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pai, Suraj, Hadzic, Ibrahim, Bontempi, Dennis, Bressem, Keno, Kann, Benjamin H., Fedorov, Andriy, Mak, Raymond H., Aerts, Hugo J. W. L.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916631726260224
author Pai, Suraj
Hadzic, Ibrahim
Bontempi, Dennis
Bressem, Keno
Kann, Benjamin H.
Fedorov, Andriy
Mak, Raymond H.
Aerts, Hugo J. W. L.
author_facet Pai, Suraj
Hadzic, Ibrahim
Bontempi, Dennis
Bressem, Keno
Kann, Benjamin H.
Fedorov, Andriy
Mak, Raymond H.
Aerts, Hugo J. W. L.
contents Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for various radiological tasks. CT-FM was pre-trained using 148,000 computed tomography (CT) scans from the Imaging Data Commons through label-agnostic contrastive learning. We evaluated CT-FM across four categories of tasks, namely, whole-body and tumor segmentation, head CT triage, medical image retrieval, and semantic understanding, showing superior performance against state-of-the-art models. Beyond quantitative success, CT-FM demonstrated the ability to cluster regions anatomically and identify similar anatomical and structural concepts across scans. Furthermore, it remained robust across test-retest settings and indicated reasonable salient regions attached to its embeddings. This study demonstrates the value of large-scale medical imaging foundation models and by open-sourcing the model weights, code, and data, aims to support more adaptable, reliable, and interpretable AI solutions in radiology.
format Preprint
id arxiv_https___arxiv_org_abs_2501_09001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Foundation Models for Computed Tomography
Pai, Suraj
Hadzic, Ibrahim
Bontempi, Dennis
Bressem, Keno
Kann, Benjamin H.
Fedorov, Andriy
Mak, Raymond H.
Aerts, Hugo J. W. L.
Image and Video Processing
Computer Vision and Pattern Recognition
Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for various radiological tasks. CT-FM was pre-trained using 148,000 computed tomography (CT) scans from the Imaging Data Commons through label-agnostic contrastive learning. We evaluated CT-FM across four categories of tasks, namely, whole-body and tumor segmentation, head CT triage, medical image retrieval, and semantic understanding, showing superior performance against state-of-the-art models. Beyond quantitative success, CT-FM demonstrated the ability to cluster regions anatomically and identify similar anatomical and structural concepts across scans. Furthermore, it remained robust across test-retest settings and indicated reasonable salient regions attached to its embeddings. This study demonstrates the value of large-scale medical imaging foundation models and by open-sourcing the model weights, code, and data, aims to support more adaptable, reliable, and interpretable AI solutions in radiology.
title Vision Foundation Models for Computed Tomography
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.09001