SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Rongsheng, Tang, Fenghe, Yao, Qingsong, Yan, Rui, Zhang, Xu, Huang, Zhen, Lai, Haoran, He, Zhiyang, Tao, Xiaodong, Jiang, Zihang, Zhou, Shaohua Kevin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916944507043840
author Wang, Rongsheng
Tang, Fenghe
Yao, Qingsong
Yan, Rui
Zhang, Xu
Huang, Zhen
Lai, Haoran
He, Zhiyang
Tao, Xiaodong
Jiang, Zihang
Zhou, Shaohua Kevin
author_facet Wang, Rongsheng
Tang, Fenghe
Yao, Qingsong
Yan, Rui
Zhang, Xu
Huang, Zhen
Lai, Haoran
He, Zhiyang
Tao, Xiaodong
Jiang, Zihang
Zhou, Shaohua Kevin
contents Medical vision-language pre-training shows great potential in learning representative features from massive paired radiographs and reports. However, in computed tomography (CT) scans, the distribution of lesions which contain intricate structures is characterized by spatial sparsity. Besides, the complex and implicit relationships between different pathological descriptions in each sentence of the report and their corresponding sub-regions in radiographs pose additional challenges. In this paper, we propose a Similarity-Driven Cross-Granularity Pre-training (SimCroP) framework on chest CTs, which combines similarity-driven alignment and cross-granularity fusion to improve radiograph interpretation. We first leverage multi-modal masked modeling to optimize the encoder for understanding precise low-level semantics from radiographs. Then, similarity-driven alignment is designed to pre-train the encoder to adaptively select and align the correct patches corresponding to each sentence in reports. The cross-granularity fusion module integrates multimodal information across instance level and word-patch level, which helps the model better capture key pathology structures in sparse radiographs, resulting in improved performance for multi-scale downstream tasks. SimCroP is pre-trained on a large-scale paired CT-reports dataset and validated on image classification and segmentation tasks across five public datasets. Experimental results demonstrate that SimCroP outperforms both cutting-edge medical self-supervised learning methods and medical vision-language pre-training methods. Codes and models are available at https://github.com/ToniChopp/SimCroP.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training
Wang, Rongsheng
Tang, Fenghe
Yao, Qingsong
Yan, Rui
Zhang, Xu
Huang, Zhen
Lai, Haoran
He, Zhiyang
Tao, Xiaodong
Jiang, Zihang
Zhou, Shaohua Kevin
Computer Vision and Pattern Recognition
Medical vision-language pre-training shows great potential in learning representative features from massive paired radiographs and reports. However, in computed tomography (CT) scans, the distribution of lesions which contain intricate structures is characterized by spatial sparsity. Besides, the complex and implicit relationships between different pathological descriptions in each sentence of the report and their corresponding sub-regions in radiographs pose additional challenges. In this paper, we propose a Similarity-Driven Cross-Granularity Pre-training (SimCroP) framework on chest CTs, which combines similarity-driven alignment and cross-granularity fusion to improve radiograph interpretation. We first leverage multi-modal masked modeling to optimize the encoder for understanding precise low-level semantics from radiographs. Then, similarity-driven alignment is designed to pre-train the encoder to adaptively select and align the correct patches corresponding to each sentence in reports. The cross-granularity fusion module integrates multimodal information across instance level and word-patch level, which helps the model better capture key pathology structures in sparse radiographs, resulting in improved performance for multi-scale downstream tasks. SimCroP is pre-trained on a large-scale paired CT-reports dataset and validated on image classification and segmentation tasks across five public datasets. Experimental results demonstrate that SimCroP outperforms both cutting-edge medical self-supervised learning methods and medical vision-language pre-training methods. Codes and models are available at https://github.com/ToniChopp/SimCroP.
title SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.08311