MedSAM3: Delving into Segment Anything with Medical Concepts

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Anglin, Xue, Rundong, Cao, Xu R., Shen, Yifan, Lu, Yi, Li, Xiang, Chen, Qianqian, Chen, Jintai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915635057917952
author Liu, Anglin
Xue, Rundong
Cao, Xu R.
Shen, Yifan
Lu, Yi
Li, Xiang
Chen, Qianqian
Chen, Jintai
author_facet Liu, Anglin
Xue, Rundong
Cao, Xu R.
Shen, Yifan
Lu, Yi
Li, Xiang
Chen, Qianqian
Chen, Jintai
contents Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizability and demand extensive, time-consuming manual annotation for new clinical application. Here, we propose MedSAM-3, a text promptable medical segmentation model for medical image and video segmentation. By fine-tuning the Segment Anything Model (SAM) 3 architecture on medical images paired with semantic conceptual labels, our MedSAM-3 enables medical Promptable Concept Segmentation (PCS), allowing precise targeting of anatomical structures via open-vocabulary text descriptions rather than solely geometric prompts. We further introduce the MedSAM-3 Agent, a framework that integrates Multimodal Large Language Models (MLLMs) to perform complex reasoning and iterative refinement in an agent-in-the-loop workflow. Comprehensive experiments across diverse medical imaging modalities, including X-ray, MRI, Ultrasound, CT, and video, demonstrate that our approach significantly outperforms existing specialist and foundation models. We will release our code and model at https://github.com/Joey-S-Liu/MedSAM3.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedSAM3: Delving into Segment Anything with Medical Concepts
Liu, Anglin
Xue, Rundong
Cao, Xu R.
Shen, Yifan
Lu, Yi
Li, Xiang
Chen, Qianqian
Chen, Jintai
Computer Vision and Pattern Recognition
Artificial Intelligence
Medical image segmentation is fundamental for biomedical discovery. Existing methods lack generalizability and demand extensive, time-consuming manual annotation for new clinical application. Here, we propose MedSAM-3, a text promptable medical segmentation model for medical image and video segmentation. By fine-tuning the Segment Anything Model (SAM) 3 architecture on medical images paired with semantic conceptual labels, our MedSAM-3 enables medical Promptable Concept Segmentation (PCS), allowing precise targeting of anatomical structures via open-vocabulary text descriptions rather than solely geometric prompts. We further introduce the MedSAM-3 Agent, a framework that integrates Multimodal Large Language Models (MLLMs) to perform complex reasoning and iterative refinement in an agent-in-the-loop workflow. Comprehensive experiments across diverse medical imaging modalities, including X-ray, MRI, Ultrasound, CT, and video, demonstrate that our approach significantly outperforms existing specialist and foundation models. We will release our code and model at https://github.com/Joey-S-Liu/MedSAM3.
title MedSAM3: Delving into Segment Anything with Medical Concepts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.19046