Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Xusheng, Zhou, Lihua, Li, Nianxin, Xu, Miao, Song, Ziyang, Yi, Dong, Wu, Jinlin, Ma, Jiawei, Liu, Hongbin, Lei, Zhen, Luo, Jiebo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911682374139904
author Liang, Xusheng
Zhou, Lihua
Li, Nianxin
Xu, Miao
Song, Ziyang
Yi, Dong
Wu, Jinlin
Ma, Jiawei
Liu, Hongbin
Lei, Zhen
Luo, Jiebo
author_facet Liang, Xusheng
Zhou, Lihua
Li, Nianxin
Xu, Miao
Song, Ziyang
Yi, Dong
Wu, Jinlin
Ma, Jiawei
Liu, Hongbin
Lei, Zhen
Luo, Jiebo
contents Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging remains challenging due to the high variability and complexity of medical data. Specifically, medical images often exhibit significant domain shifts caused by various confounders, including equipment differences, procedure artifacts, and imaging modes, which can lead to poor generalization when models are applied to unseen domains. To address this limitation, we propose Multimodal Causal-Driven Representation Learning (MCDRL), a novel framework that integrates causal inference with the VLM to tackle domain generalization in medical image segmentation. MCDRL is implemented in two steps: first, it leverages CLIP's cross-modal capabilities to identify candidate lesion regions and construct a confounder dictionary through text prompts, specifically designed to represent domain-specific variations; second, it trains a causal intervention network that utilizes this dictionary to identify and eliminate the influence of these domain-specific variations while preserving the anatomical structural information critical for segmentation tasks. Extensive experiments demonstrate that MCDRL consistently outperforms competing methods, yielding superior segmentation accuracy and exhibiting robust generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
Liang, Xusheng
Zhou, Lihua
Li, Nianxin
Xu, Miao
Song, Ziyang
Yi, Dong
Wu, Jinlin
Ma, Jiawei
Liu, Hongbin
Lei, Zhen
Luo, Jiebo
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging remains challenging due to the high variability and complexity of medical data. Specifically, medical images often exhibit significant domain shifts caused by various confounders, including equipment differences, procedure artifacts, and imaging modes, which can lead to poor generalization when models are applied to unseen domains. To address this limitation, we propose Multimodal Causal-Driven Representation Learning (MCDRL), a novel framework that integrates causal inference with the VLM to tackle domain generalization in medical image segmentation. MCDRL is implemented in two steps: first, it leverages CLIP's cross-modal capabilities to identify candidate lesion regions and construct a confounder dictionary through text prompts, specifically designed to represent domain-specific variations; second, it trains a causal intervention network that utilizes this dictionary to identify and eliminate the influence of these domain-specific variations while preserving the anatomical structural information critical for segmentation tasks. Extensive experiments demonstrate that MCDRL consistently outperforms competing methods, yielding superior segmentation accuracy and exhibiting robust generalizability.
title Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.05008