Diff4MMLiTS: Advanced Multimodal Liver Tumor Segmentation via Diffusion-Based Image Synthesis and Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Shiyun, Lin, Li, Cheng, Pujin, Jin, ZhiCheng, Chen, JianJian, Zhu, HaiDong, Wong, Kenneth K. Y., Tang, Xiaoying
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918320103489536
author Chen, Shiyun
Lin, Li
Cheng, Pujin
Jin, ZhiCheng
Chen, JianJian
Zhu, HaiDong
Wong, Kenneth K. Y.
Tang, Xiaoying
author_facet Chen, Shiyun
Lin, Li
Cheng, Pujin
Jin, ZhiCheng
Chen, JianJian
Zhu, HaiDong
Wong, Kenneth K. Y.
Tang, Xiaoying
contents Multimodal learning has been demonstrated to enhance performance across various clinical tasks, owing to the diverse perspectives offered by different modalities of data. However, existing multimodal segmentation methods rely on well-registered multimodal data, which is unrealistic for real-world clinical images, particularly for indistinct and diffuse regions such as liver tumors. In this paper, we introduce Diff4MMLiTS, a four-stage multimodal liver tumor segmentation pipeline: pre-registration of the target organs in multimodal CTs; dilation of the annotated modality's mask and followed by its use in inpainting to obtain multimodal normal CTs without tumors; synthesis of strictly aligned multimodal CTs with tumors using the latent diffusion model based on multimodal CT features and randomly generated tumor masks; and finally, training the segmentation model, thus eliminating the need for strictly aligned multimodal data. Extensive experiments on public and internal datasets demonstrate the superiority of Diff4MMLiTS over other state-of-the-art multimodal segmentation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20418
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diff4MMLiTS: Advanced Multimodal Liver Tumor Segmentation via Diffusion-Based Image Synthesis and Alignment
Chen, Shiyun
Lin, Li
Cheng, Pujin
Jin, ZhiCheng
Chen, JianJian
Zhu, HaiDong
Wong, Kenneth K. Y.
Tang, Xiaoying
Image and Video Processing
Computer Vision and Pattern Recognition
Multimodal learning has been demonstrated to enhance performance across various clinical tasks, owing to the diverse perspectives offered by different modalities of data. However, existing multimodal segmentation methods rely on well-registered multimodal data, which is unrealistic for real-world clinical images, particularly for indistinct and diffuse regions such as liver tumors. In this paper, we introduce Diff4MMLiTS, a four-stage multimodal liver tumor segmentation pipeline: pre-registration of the target organs in multimodal CTs; dilation of the annotated modality's mask and followed by its use in inpainting to obtain multimodal normal CTs without tumors; synthesis of strictly aligned multimodal CTs with tumors using the latent diffusion model based on multimodal CT features and randomly generated tumor masks; and finally, training the segmentation model, thus eliminating the need for strictly aligned multimodal data. Extensive experiments on public and internal datasets demonstrate the superiority of Diff4MMLiTS over other state-of-the-art multimodal segmentation methods.
title Diff4MMLiTS: Advanced Multimodal Liver Tumor Segmentation via Diffusion-Based Image Synthesis and Alignment
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20418