ViM-UNet: Vision Mamba for Biomedical Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Archit, Anwai, Pape, Constantin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910447732523008
author Archit, Anwai
Pape, Constantin
author_facet Archit, Anwai
Pape, Constantin
contents CNNs, most notably the UNet, are the default architecture for biomedical segmentation. Transformer-based approaches, such as UNETR, have been proposed to replace them, benefiting from a global field of view, but suffering from larger runtimes and higher parameter counts. The recent Vision Mamba architecture offers a compelling alternative to transformers, also providing a global field of view, but at higher efficiency. Here, we introduce ViM-UNet, a novel segmentation architecture based on it and compare it to UNet and UNETR for two challenging microscopy instance segmentation tasks. We find that it performs similarly or better than UNet, depending on the task, and outperforms UNETR while being more efficient. Our code is open source and documented at https://github.com/constantinpape/torch-em/blob/main/vimunet.md.
format Preprint
id arxiv_https___arxiv_org_abs_2404_07705
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ViM-UNet: Vision Mamba for Biomedical Segmentation
Archit, Anwai
Pape, Constantin
Computer Vision and Pattern Recognition
CNNs, most notably the UNet, are the default architecture for biomedical segmentation. Transformer-based approaches, such as UNETR, have been proposed to replace them, benefiting from a global field of view, but suffering from larger runtimes and higher parameter counts. The recent Vision Mamba architecture offers a compelling alternative to transformers, also providing a global field of view, but at higher efficiency. Here, we introduce ViM-UNet, a novel segmentation architecture based on it and compare it to UNet and UNETR for two challenging microscopy instance segmentation tasks. We find that it performs similarly or better than UNet, depending on the task, and outperforms UNETR while being more efficient. Our code is open source and documented at https://github.com/constantinpape/torch-em/blob/main/vimunet.md.
title ViM-UNet: Vision Mamba for Biomedical Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.07705