Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yutao, Zheng, Ying, Miao, Shumei, Zhang, Xiaolei, Xia, Jiahao, Qi, Yaolei, Zhang, Yiyang, He, Yuting, Chen, Qian, Ye, Jing, Qiao, Hongyan, Hu, Xiuhua, Xu, Lei, Zhang, Jiayin, Liu, Hui, Zheng, Minwen, Wang, Yining, Zhang, Daimin, Zhang, Ji, Shao, Wenqi, Liu, Yun, Zhang, Longjiang, Yang, Guanyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912509049438208
author Hu, Yutao
Zheng, Ying
Miao, Shumei
Zhang, Xiaolei
Xia, Jiahao
Qi, Yaolei
Zhang, Yiyang
He, Yuting
Chen, Qian
Ye, Jing
Qiao, Hongyan
Hu, Xiuhua
Xu, Lei
Zhang, Jiayin
Liu, Hui
Zheng, Minwen
Wang, Yining
Zhang, Daimin
Zhang, Ji
Shao, Wenqi
Liu, Yun
Zhang, Longjiang
Yang, Guanyu
author_facet Hu, Yutao
Zheng, Ying
Miao, Shumei
Zhang, Xiaolei
Xia, Jiahao
Qi, Yaolei
Zhang, Yiyang
He, Yuting
Chen, Qian
Ye, Jing
Qiao, Hongyan
Hu, Xiuhua
Xu, Lei
Zhang, Jiayin
Liu, Hui
Zheng, Minwen
Wang, Yining
Zhang, Daimin
Zhang, Ji
Shao, Wenqi
Liu, Yun
Zhang, Longjiang
Yang, Guanyu
contents Foundation models have demonstrated remarkable potential in medical domain. However, their application to complex cardiovascular diagnostics remains underexplored. In this paper, we present Cardiac-CLIP, a multi-modal foundation model designed for 3D cardiac CT images. Cardiac-CLIP is developed through a two-stage pre-training strategy. The first stage employs a 3D masked autoencoder (MAE) to perform self-supervised representation learning from large-scale unlabeled volumetric data, enabling the visual encoder to capture rich anatomical and contextual features. In the second stage, contrastive learning is introduced to align visual and textual representations, facilitating cross-modal understanding. To support the pre-training, we collect 16641 real clinical CT scans, supplemented by 114k publicly available data. Meanwhile, we standardize free-text radiology reports into unified templates and construct the pathology vectors according to diagnostic attributes, based on which the soft-label matrix is generated to supervise the contrastive learning process. On the other hand, to comprehensively evaluate the effectiveness of Cardiac-CLIP, we collect 6,722 real-clinical data from 12 independent institutions, along with the open-source data to construct the evaluation dataset. Specifically, Cardiac-CLIP is comprehensively evaluated across multiple tasks, including cardiovascular abnormality classification, information retrieval and clinical analysis. Experimental results demonstrate that Cardiac-CLIP achieves state-of-the-art performance across various downstream tasks in both internal and external data. Particularly, Cardiac-CLIP exhibits great effectiveness in supporting complex clinical tasks such as the prospective prediction of acute coronary syndrome, which is notoriously difficult in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
Hu, Yutao
Zheng, Ying
Miao, Shumei
Zhang, Xiaolei
Xia, Jiahao
Qi, Yaolei
Zhang, Yiyang
He, Yuting
Chen, Qian
Ye, Jing
Qiao, Hongyan
Hu, Xiuhua
Xu, Lei
Zhang, Jiayin
Liu, Hui
Zheng, Minwen
Wang, Yining
Zhang, Daimin
Zhang, Ji
Shao, Wenqi
Liu, Yun
Zhang, Longjiang
Yang, Guanyu
Image and Video Processing
Computer Vision and Pattern Recognition
Foundation models have demonstrated remarkable potential in medical domain. However, their application to complex cardiovascular diagnostics remains underexplored. In this paper, we present Cardiac-CLIP, a multi-modal foundation model designed for 3D cardiac CT images. Cardiac-CLIP is developed through a two-stage pre-training strategy. The first stage employs a 3D masked autoencoder (MAE) to perform self-supervised representation learning from large-scale unlabeled volumetric data, enabling the visual encoder to capture rich anatomical and contextual features. In the second stage, contrastive learning is introduced to align visual and textual representations, facilitating cross-modal understanding. To support the pre-training, we collect 16641 real clinical CT scans, supplemented by 114k publicly available data. Meanwhile, we standardize free-text radiology reports into unified templates and construct the pathology vectors according to diagnostic attributes, based on which the soft-label matrix is generated to supervise the contrastive learning process. On the other hand, to comprehensively evaluate the effectiveness of Cardiac-CLIP, we collect 6,722 real-clinical data from 12 independent institutions, along with the open-source data to construct the evaluation dataset. Specifically, Cardiac-CLIP is comprehensively evaluated across multiple tasks, including cardiovascular abnormality classification, information retrieval and clinical analysis. Experimental results demonstrate that Cardiac-CLIP achieves state-of-the-art performance across various downstream tasks in both internal and external data. Particularly, Cardiac-CLIP exhibits great effectiveness in supporting complex clinical tasks such as the prospective prediction of acute coronary syndrome, which is notoriously difficult in real-world scenarios.
title Cardiac-CLIP: A Vision-Language Foundation Model for 3D Cardiac CT Images
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.22024