STA-Unet: Rethink the semantic redundant for Medical Imaging Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vasa, Vamsi Krishna, Zhu, Wenhui, Chen, Xiwen, Qiu, Peijie, Dong, Xuanzhao, Wang, Yalin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929543689797632
author Vasa, Vamsi Krishna
Zhu, Wenhui
Chen, Xiwen
Qiu, Peijie
Dong, Xuanzhao
Wang, Yalin
author_facet Vasa, Vamsi Krishna
Zhu, Wenhui
Chen, Xiwen
Qiu, Peijie
Dong, Xuanzhao
Wang, Yalin
contents In recent years, significant progress has been made in the medical image analysis domain using convolutional neural networks (CNNs). In particular, deep neural networks based on a U-shaped architecture (UNet) with skip connections have been adopted for several medical imaging tasks, including organ segmentation. Despite their great success, CNNs are not good at learning global or semantic features. Especially ones that require human-like reasoning to understand the context. Many UNet architectures attempted to adjust with the introduction of Transformer-based self-attention mechanisms, and notable gains in performance have been noted. However, the transformers are inherently flawed with redundancy to learn at shallow layers, which often leads to an increase in the computation of attention from the nearby pixels offering limited information. The recently introduced Super Token Attention (STA) mechanism adapts the concept of superpixels from pixel space to token space, using super tokens as compact visual representations. This approach tackles the redundancy by learning efficient global representations in vision transformers, especially for the shallow layers. In this work, we introduce the STA module in the UNet architecture (STA-UNet), to limit redundancy without losing rich information. Experimental results on four publicly available datasets demonstrate the superiority of STA-UNet over existing state-of-the-art architectures in terms of Dice score and IOU for organ segmentation tasks. The code is available at \url{https://github.com/Retinal-Research/STA-UNet}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11578
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STA-Unet: Rethink the semantic redundant for Medical Imaging Segmentation
Vasa, Vamsi Krishna
Zhu, Wenhui
Chen, Xiwen
Qiu, Peijie
Dong, Xuanzhao
Wang, Yalin
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
In recent years, significant progress has been made in the medical image analysis domain using convolutional neural networks (CNNs). In particular, deep neural networks based on a U-shaped architecture (UNet) with skip connections have been adopted for several medical imaging tasks, including organ segmentation. Despite their great success, CNNs are not good at learning global or semantic features. Especially ones that require human-like reasoning to understand the context. Many UNet architectures attempted to adjust with the introduction of Transformer-based self-attention mechanisms, and notable gains in performance have been noted. However, the transformers are inherently flawed with redundancy to learn at shallow layers, which often leads to an increase in the computation of attention from the nearby pixels offering limited information. The recently introduced Super Token Attention (STA) mechanism adapts the concept of superpixels from pixel space to token space, using super tokens as compact visual representations. This approach tackles the redundancy by learning efficient global representations in vision transformers, especially for the shallow layers. In this work, we introduce the STA module in the UNet architecture (STA-UNet), to limit redundancy without losing rich information. Experimental results on four publicly available datasets demonstrate the superiority of STA-UNet over existing state-of-the-art architectures in terms of Dice score and IOU for organ segmentation tasks. The code is available at \url{https://github.com/Retinal-Research/STA-UNet}.
title STA-Unet: Rethink the semantic redundant for Medical Imaging Segmentation
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.11578