Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Tan, Tan, Zhaorui, Guo, Kaiyu, Xu, Dongli, Xu, Weidi, Jiang, Chen, Guo, Xin, Qi, Yuan, Cheng, Yuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916824377982976
author Pan, Tan
Tan, Zhaorui
Guo, Kaiyu
Xu, Dongli
Xu, Weidi
Jiang, Chen
Guo, Xin
Qi, Yuan
Cheng, Yuan
author_facet Pan, Tan
Tan, Zhaorui
Guo, Kaiyu
Xu, Dongli
Xu, Weidi
Jiang, Chen
Guo, Xin
Qi, Yuan
Cheng, Yuan
contents 3D medical image self-supervised learning (mSSL) holds great promise for medical analysis. Effectively supporting broader applications requires considering anatomical structure variations in location, scale, and morphology, which are crucial for capturing meaningful distinctions. However, previous mSSL methods partition images with fixed-size patches, often ignoring the structure variations. In this work, we introduce a novel perspective on 3D medical images with the goal of learning structure-aware representations. We assume that patches within the same structure share the same semantics (semantic consistency) while those from different structures exhibit distinct semantics (semantic discrepancy). Based on this assumption, we propose an mSSL framework named $S^2DC$, achieving Structure-aware Semantic Discrepancy and Consistency in two steps. First, $S^2DC$ enforces distinct representations for different patches to increase semantic discrepancy by leveraging an optimal transport strategy. Second, $S^2DC$ advances semantic consistency at the structural level based on neighborhood similarity distribution. By bridging patch-level and structure-level representations, $S^2DC$ achieves structure-aware representations. Thoroughly evaluated across 10 datasets, 4 tasks, and 3 modalities, our proposed method consistently outperforms the state-of-the-art methods in mSSL.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning
Pan, Tan
Tan, Zhaorui
Guo, Kaiyu
Xu, Dongli
Xu, Weidi
Jiang, Chen
Guo, Xin
Qi, Yuan
Cheng, Yuan
Computer Vision and Pattern Recognition
3D medical image self-supervised learning (mSSL) holds great promise for medical analysis. Effectively supporting broader applications requires considering anatomical structure variations in location, scale, and morphology, which are crucial for capturing meaningful distinctions. However, previous mSSL methods partition images with fixed-size patches, often ignoring the structure variations. In this work, we introduce a novel perspective on 3D medical images with the goal of learning structure-aware representations. We assume that patches within the same structure share the same semantics (semantic consistency) while those from different structures exhibit distinct semantics (semantic discrepancy). Based on this assumption, we propose an mSSL framework named $S^2DC$, achieving Structure-aware Semantic Discrepancy and Consistency in two steps. First, $S^2DC$ enforces distinct representations for different patches to increase semantic discrepancy by leveraging an optimal transport strategy. Second, $S^2DC$ advances semantic consistency at the structural level based on neighborhood similarity distribution. By bridging patch-level and structure-level representations, $S^2DC$ achieves structure-aware representations. Thoroughly evaluated across 10 datasets, 4 tasks, and 3 modalities, our proposed method consistently outperforms the state-of-the-art methods in mSSL.
title Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.02581