Learning Multi-axis Representation in Frequency Domain for Medical Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruan, Jiacheng, Gao, Jingsheng, Xie, Mingye, Xiang, Suncheng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914955487346688
author Ruan, Jiacheng
Gao, Jingsheng
Xie, Mingye
Xiang, Suncheng
author_facet Ruan, Jiacheng
Gao, Jingsheng
Xie, Mingye
Xiang, Suncheng
contents Recently, Visual Transformer (ViT) has been extensively used in medical image segmentation (MIS) due to applying self-attention mechanism in the spatial domain to modeling global knowledge. However, many studies have focused on improving models in the spatial domain while neglecting the importance of frequency domain information. Therefore, we propose Multi-axis External Weights UNet (MEW-UNet) based on the U-shape architecture by replacing self-attention in ViT with our Multi-axis External Weights block. Specifically, our block performs a Fourier transform on the three axes of the input features and assigns the external weight in the frequency domain, which is generated by our External Weights Generator. Then, an inverse Fourier transform is performed to change the features back to the spatial domain. We evaluate our model on four datasets, including Synapse, ACDC, ISIC17 and ISIC18 datasets, and our approach demonstrates competitive performance, owing to its effective utilization of frequency domain information.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17030
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Multi-axis Representation in Frequency Domain for Medical Image Segmentation
Ruan, Jiacheng
Gao, Jingsheng
Xie, Mingye
Xiang, Suncheng
Image and Video Processing
Computer Vision and Pattern Recognition
Recently, Visual Transformer (ViT) has been extensively used in medical image segmentation (MIS) due to applying self-attention mechanism in the spatial domain to modeling global knowledge. However, many studies have focused on improving models in the spatial domain while neglecting the importance of frequency domain information. Therefore, we propose Multi-axis External Weights UNet (MEW-UNet) based on the U-shape architecture by replacing self-attention in ViT with our Multi-axis External Weights block. Specifically, our block performs a Fourier transform on the three axes of the input features and assigns the external weight in the frequency domain, which is generated by our External Weights Generator. Then, an inverse Fourier transform is performed to change the features back to the spatial domain. We evaluate our model on four datasets, including Synapse, ACDC, ISIC17 and ISIC18 datasets, and our approach demonstrates competitive performance, owing to its effective utilization of frequency domain information.
title Learning Multi-axis Representation in Frequency Domain for Medical Image Segmentation
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.17030