Slide-SAM: Medical SAM Meets Sliding Window

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Quan, Quan, Tang, Fenghe, Xu, Zikang, Zhu, Heqin, Zhou, S. Kevin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913316363829248
author Quan, Quan
Tang, Fenghe
Xu, Zikang
Zhu, Heqin
Zhou, S. Kevin
author_facet Quan, Quan
Tang, Fenghe
Xu, Zikang
Zhu, Heqin
Zhou, S. Kevin
contents The Segment Anything Model (SAM) has achieved a notable success in two-dimensional image segmentation in natural images. However, the substantial gap between medical and natural images hinders its direct application to medical image segmentation tasks. Particularly in 3D medical images, SAM struggles to learn contextual relationships between slices, limiting its practical applicability. Moreover, applying 2D SAM to 3D images requires prompting the entire volume, which is time- and label-consuming. To address these problems, we propose Slide-SAM, which treats a stack of three adjacent slices as a prediction window. It firstly takes three slices from a 3D volume and point- or bounding box prompts on the central slice as inputs to predict segmentation masks for all three slices. Subsequently, the masks of the top and bottom slices are then used to generate new prompts for adjacent slices. Finally, step-wise prediction can be achieved by sliding the prediction window forward or backward through the entire volume. Our model is trained on multiple public and private medical datasets and demonstrates its effectiveness through extensive 3D segmetnation experiments, with the help of minimal prompts. Code is available at \url{https://github.com/Curli-quan/Slide-SAM}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_10121
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Slide-SAM: Medical SAM Meets Sliding Window
Quan, Quan
Tang, Fenghe
Xu, Zikang
Zhu, Heqin
Zhou, S. Kevin
Computer Vision and Pattern Recognition
The Segment Anything Model (SAM) has achieved a notable success in two-dimensional image segmentation in natural images. However, the substantial gap between medical and natural images hinders its direct application to medical image segmentation tasks. Particularly in 3D medical images, SAM struggles to learn contextual relationships between slices, limiting its practical applicability. Moreover, applying 2D SAM to 3D images requires prompting the entire volume, which is time- and label-consuming. To address these problems, we propose Slide-SAM, which treats a stack of three adjacent slices as a prediction window. It firstly takes three slices from a 3D volume and point- or bounding box prompts on the central slice as inputs to predict segmentation masks for all three slices. Subsequently, the masks of the top and bottom slices are then used to generate new prompts for adjacent slices. Finally, step-wise prediction can be achieved by sliding the prediction window forward or backward through the entire volume. Our model is trained on multiple public and private medical datasets and demonstrates its effectiveness through extensive 3D segmetnation experiments, with the help of minimal prompts. Code is available at \url{https://github.com/Curli-quan/Slide-SAM}.
title Slide-SAM: Medical SAM Meets Sliding Window
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.10121