CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xiao, Gao, Peng, Yu, Tao, Wang, Fei, Yuan, Ru-Yue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917779761790976
author Liu, Xiao
Gao, Peng
Yu, Tao
Wang, Fei
Yuan, Ru-Yue
author_facet Liu, Xiao
Gao, Peng
Yu, Tao
Wang, Fei
Yuan, Ru-Yue
contents Deep learning, especially convolutional neural networks (CNNs) and Transformer architectures, have become the focus of extensive research in medical image segmentation, achieving impressive results. However, CNNs come with inductive biases that limit their effectiveness in more complex, varied segmentation scenarios. Conversely, while Transformer-based methods excel at capturing global and long-range semantic details, they suffer from high computational demands. In this study, we propose CSWin-UNet, a novel U-shaped segmentation method that incorporates the CSWin self-attention mechanism into the UNet to facilitate horizontal and vertical stripes self-attention. This method significantly enhances both computational efficiency and receptive field interactions. Additionally, our innovative decoder utilizes a content-aware reassembly operator that strategically reassembles features, guided by predicted kernels, for precise image resolution restoration. Our extensive empirical evaluations on diverse datasets, including synapse multi-organ CT, cardiac MRI, and skin lesions, demonstrate that CSWin-UNet maintains low model complexity while delivering high segmentation accuracy. Codes are available at https://github.com/eatbeanss/CSWin-UNet.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18070
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation
Liu, Xiao
Gao, Peng
Yu, Tao
Wang, Fei
Yuan, Ru-Yue
Image and Video Processing
Computer Vision and Pattern Recognition
Deep learning, especially convolutional neural networks (CNNs) and Transformer architectures, have become the focus of extensive research in medical image segmentation, achieving impressive results. However, CNNs come with inductive biases that limit their effectiveness in more complex, varied segmentation scenarios. Conversely, while Transformer-based methods excel at capturing global and long-range semantic details, they suffer from high computational demands. In this study, we propose CSWin-UNet, a novel U-shaped segmentation method that incorporates the CSWin self-attention mechanism into the UNet to facilitate horizontal and vertical stripes self-attention. This method significantly enhances both computational efficiency and receptive field interactions. Additionally, our innovative decoder utilizes a content-aware reassembly operator that strategically reassembles features, guided by predicted kernels, for precise image resolution restoration. Our extensive empirical evaluations on diverse datasets, including synapse multi-organ CT, cardiac MRI, and skin lesions, demonstrate that CSWin-UNet maintains low model complexity while delivering high segmentation accuracy. Codes are available at https://github.com/eatbeanss/CSWin-UNet.
title CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.18070