Taming Mambas for Voxel Level 3D Medical Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lumetti, Luca, Pipoli, Vittorio, Marchesini, Kevin, Ficarra, Elisa, Grana, Costantino, Bolelli, Federico
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917810166300672
author Lumetti, Luca
Pipoli, Vittorio
Marchesini, Kevin
Ficarra, Elisa
Grana, Costantino
Bolelli, Federico
author_facet Lumetti, Luca
Pipoli, Vittorio
Marchesini, Kevin
Ficarra, Elisa
Grana, Costantino
Bolelli, Federico
contents Recently, the field of 3D medical segmentation has been dominated by deep learning models employing Convolutional Neural Networks (CNNs) and Transformer-based architectures, each with their distinctive strengths and limitations. CNNs are constrained by a local receptive field, whereas transformers are hindered by their substantial memory requirements as well as they data hungriness, making them not ideal for processing 3D medical volumes at a fine-grained level. For these reasons, fully convolutional neural networks, as nnUNet, still dominate the scene when segmenting medical structures in 3D large medical volumes. Despite numerous advancements towards developing transformer variants with subquadratic time and memory complexity, these models still fall short in content-based reasoning. A recent breakthrough is Mamba, a Recurrent Neural Network (RNN) based on State Space Models (SSMs) outperforming Transformers in many long-context tasks (million-length sequences) on famous natural language processing and genomic benchmarks while keeping a linear complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15496
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Taming Mambas for Voxel Level 3D Medical Image Segmentation
Lumetti, Luca
Pipoli, Vittorio
Marchesini, Kevin
Ficarra, Elisa
Grana, Costantino
Bolelli, Federico
Computer Vision and Pattern Recognition
Recently, the field of 3D medical segmentation has been dominated by deep learning models employing Convolutional Neural Networks (CNNs) and Transformer-based architectures, each with their distinctive strengths and limitations. CNNs are constrained by a local receptive field, whereas transformers are hindered by their substantial memory requirements as well as they data hungriness, making them not ideal for processing 3D medical volumes at a fine-grained level. For these reasons, fully convolutional neural networks, as nnUNet, still dominate the scene when segmenting medical structures in 3D large medical volumes. Despite numerous advancements towards developing transformer variants with subquadratic time and memory complexity, these models still fall short in content-based reasoning. A recent breakthrough is Mamba, a Recurrent Neural Network (RNN) based on State Space Models (SSMs) outperforming Transformers in many long-context tasks (million-length sequences) on famous natural language processing and genomic benchmarks while keeping a linear complexity.
title Taming Mambas for Voxel Level 3D Medical Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.15496