MOSformer: Momentum encoder-based inter-slice fusion transformer for medical image segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, De-Xing, Zhou, Xiao-Hu, Gui, Mei-Jiang, Xie, Xiao-Liang, Liu, Shi-Qi, Wang, Shuang-Yi, Feng, Zhen-Qiu, Lai, Zhi-Chao, Hou, Zeng-Guang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914151644790784
author Huang, De-Xing
Zhou, Xiao-Hu
Gui, Mei-Jiang
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Feng, Zhen-Qiu
Lai, Zhi-Chao
Hou, Zeng-Guang
author_facet Huang, De-Xing
Zhou, Xiao-Hu
Gui, Mei-Jiang
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Feng, Zhen-Qiu
Lai, Zhi-Chao
Hou, Zeng-Guang
contents Medical image segmentation takes an important position in various clinical applications. 2.5D-based segmentation models bridge the computational efficiency of 2D-based models with the spatial perception capabilities of 3D-based models. However, existing 2.5D-based models primarily adopt a single encoder to extract features of target and neighborhood slices, failing to effectively fuse inter-slice information, resulting in suboptimal segmentation performance. In this study, a novel momentum encoder-based inter-slice fusion transformer (MOSformer) is proposed to overcome this issue by leveraging inter-slice information from multi-scale feature maps extracted by different encoders. Specifically, dual encoders are employed to enhance feature distinguishability among different slices. One of the encoders is moving-averaged to maintain consistent slice representations. Moreover, an inter-slice fusion transformer (IF-Trans) module is developed to fuse inter-slice multi-scale features. MOSformer is evaluated on three benchmark datasets (Synapse, ACDC, and AMOS), achieving a new state-of-the-art with 85.63%, 92.19%, and 85.43% DSC, respectively. These results demonstrate MOSformer's competitiveness in medical image segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11856
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MOSformer: Momentum encoder-based inter-slice fusion transformer for medical image segmentation
Huang, De-Xing
Zhou, Xiao-Hu
Gui, Mei-Jiang
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Feng, Zhen-Qiu
Lai, Zhi-Chao
Hou, Zeng-Guang
Image and Video Processing
Computer Vision and Pattern Recognition
Medical image segmentation takes an important position in various clinical applications. 2.5D-based segmentation models bridge the computational efficiency of 2D-based models with the spatial perception capabilities of 3D-based models. However, existing 2.5D-based models primarily adopt a single encoder to extract features of target and neighborhood slices, failing to effectively fuse inter-slice information, resulting in suboptimal segmentation performance. In this study, a novel momentum encoder-based inter-slice fusion transformer (MOSformer) is proposed to overcome this issue by leveraging inter-slice information from multi-scale feature maps extracted by different encoders. Specifically, dual encoders are employed to enhance feature distinguishability among different slices. One of the encoders is moving-averaged to maintain consistent slice representations. Moreover, an inter-slice fusion transformer (IF-Trans) module is developed to fuse inter-slice multi-scale features. MOSformer is evaluated on three benchmark datasets (Synapse, ACDC, and AMOS), achieving a new state-of-the-art with 85.63%, 92.19%, and 85.43% DSC, respectively. These results demonstrate MOSformer's competitiveness in medical image segmentation.
title MOSformer: Momentum encoder-based inter-slice fusion transformer for medical image segmentation
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.11856