Label-efficient multi-organ segmentation with a diffusion model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Yongzhi, Xi, Fengjun, Tu, Liyun, Zhu, Jinxin, Hassan, Haseeb, Su, Liyilei, Peng, Yun, Li, Jingyu, Ma, Jun, Huang, Bingding
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913746218123264
author Huang, Yongzhi
Xi, Fengjun
Tu, Liyun
Zhu, Jinxin
Hassan, Haseeb
Su, Liyilei
Peng, Yun
Li, Jingyu
Ma, Jun
Huang, Bingding
author_facet Huang, Yongzhi
Xi, Fengjun
Tu, Liyun
Zhu, Jinxin
Hassan, Haseeb
Su, Liyilei
Peng, Yun
Li, Jingyu
Ma, Jun
Huang, Bingding
contents Accurate segmentation of multiple organs in Computed Tomography (CT) images plays a vital role in computer-aided diagnosis systems. While various supervised learning approaches have been proposed recently, these methods heavily depend on a large amount of high-quality labeled data, which are expensive to obtain in practice. To address this challenge, we propose a label-efficient framework using knowledge transfer from a pre-trained diffusion model for CT multi-organ segmentation. Specifically, we first pre-train a denoising diffusion model on 207,029 unlabeled 2D CT slices to capture anatomical patterns. Then, the model backbone is transferred to the downstream multi-organ segmentation task, followed by fine-tuning with few labeled data. In fine-tuning, two fine-tuning strategies, linear classification and fine-tuning decoder, are employed to enhance segmentation performance while preserving learned representations. Quantitative results show that the pre-trained diffusion model is capable of generating diverse and realistic 256x256 CT images (Fréchet inception distance (FID): 11.32, spatial Fréchet inception distance (sFID): 46.93, F1-score: 73.1%). Compared to state-of-the-art methods for multi-organ segmentation, our method achieves competitive performance on the FLARE 2022 dataset, particularly in limited labeled data scenarios. After fine-tuning with 1% and 10% labeled data, our method achieves dice similarity coefficients (DSCs) of 71.56% and 78.51%, respectively. Remarkably, the method achieves a DSC score of 51.81% using only four labeled CT slices. These results demonstrate the efficacy of our approach in overcoming the limitations of supervised learning approaches that is highly dependent on large-scale labeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15216
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Label-efficient multi-organ segmentation with a diffusion model
Huang, Yongzhi
Xi, Fengjun
Tu, Liyun
Zhu, Jinxin
Hassan, Haseeb
Su, Liyilei
Peng, Yun
Li, Jingyu
Ma, Jun
Huang, Bingding
Computer Vision and Pattern Recognition
Accurate segmentation of multiple organs in Computed Tomography (CT) images plays a vital role in computer-aided diagnosis systems. While various supervised learning approaches have been proposed recently, these methods heavily depend on a large amount of high-quality labeled data, which are expensive to obtain in practice. To address this challenge, we propose a label-efficient framework using knowledge transfer from a pre-trained diffusion model for CT multi-organ segmentation. Specifically, we first pre-train a denoising diffusion model on 207,029 unlabeled 2D CT slices to capture anatomical patterns. Then, the model backbone is transferred to the downstream multi-organ segmentation task, followed by fine-tuning with few labeled data. In fine-tuning, two fine-tuning strategies, linear classification and fine-tuning decoder, are employed to enhance segmentation performance while preserving learned representations. Quantitative results show that the pre-trained diffusion model is capable of generating diverse and realistic 256x256 CT images (Fréchet inception distance (FID): 11.32, spatial Fréchet inception distance (sFID): 46.93, F1-score: 73.1%). Compared to state-of-the-art methods for multi-organ segmentation, our method achieves competitive performance on the FLARE 2022 dataset, particularly in limited labeled data scenarios. After fine-tuning with 1% and 10% labeled data, our method achieves dice similarity coefficients (DSCs) of 71.56% and 78.51%, respectively. Remarkably, the method achieves a DSC score of 51.81% using only four labeled CT slices. These results demonstrate the efficacy of our approach in overcoming the limitations of supervised learning approaches that is highly dependent on large-scale labeled data.
title Label-efficient multi-organ segmentation with a diffusion model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.15216