Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moglia, Andrea, Leccardi, Matteo, Cavicchioli, Matteo, Maccarini, Alice, Marcon, Marco, Mainardi, Luca, Cerveri, Pietro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918210932047872
author Moglia, Andrea
Leccardi, Matteo
Cavicchioli, Matteo
Maccarini, Alice
Marcon, Marco
Mainardi, Luca
Cerveri, Pietro
author_facet Moglia, Andrea
Leccardi, Matteo
Cavicchioli, Matteo
Maccarini, Alice
Marcon, Marco
Mainardi, Luca
Cerveri, Pietro
contents Following the successful paradigm shift of large language models, leveraging pre-training on a massive corpus of data and fine-tuning on different downstream tasks, generalist models have made their foray into computer vision. The introduction of Segment Anything Model (SAM) set a milestone on segmentation of natural images, inspiring the design of a multitude of architectures for medical image segmentation. In this survey we offer a comprehensive and in-depth investigation on generalist models for medical image segmentation. We start with an introduction on the fundamentals concepts underpinning their development. Then, we provide a taxonomy on the different declinations of SAM in terms of zero-shot, few-shot, fine-tuning, adapters, on the recent SAM 2, on other innovative models trained on images alone, and others trained on both text and images. We thoroughly analyze their performances at the level of both primary research and best-in-literature, followed by a rigorous comparison with the state-of-the-art task-specific models. We emphasize the need to address challenges in terms of compliance with regulatory frameworks, privacy and security laws, budget, and trustworthy artificial intelligence (AI). Finally, we share our perspective on future directions concerning synthetic data, early fusion, lessons learnt from generalist models in natural language processing, agentic AI and physical AI, and clinical translation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
Moglia, Andrea
Leccardi, Matteo
Cavicchioli, Matteo
Maccarini, Alice
Marcon, Marco
Mainardi, Luca
Cerveri, Pietro
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
A.1; I.2.0; I.4.6
Following the successful paradigm shift of large language models, leveraging pre-training on a massive corpus of data and fine-tuning on different downstream tasks, generalist models have made their foray into computer vision. The introduction of Segment Anything Model (SAM) set a milestone on segmentation of natural images, inspiring the design of a multitude of architectures for medical image segmentation. In this survey we offer a comprehensive and in-depth investigation on generalist models for medical image segmentation. We start with an introduction on the fundamentals concepts underpinning their development. Then, we provide a taxonomy on the different declinations of SAM in terms of zero-shot, few-shot, fine-tuning, adapters, on the recent SAM 2, on other innovative models trained on images alone, and others trained on both text and images. We thoroughly analyze their performances at the level of both primary research and best-in-literature, followed by a rigorous comparison with the state-of-the-art task-specific models. We emphasize the need to address challenges in terms of compliance with regulatory frameworks, privacy and security laws, budget, and trustworthy artificial intelligence (AI). Finally, we share our perspective on future directions concerning synthetic data, early fusion, lessons learnt from generalist models in natural language processing, agentic AI and physical AI, and clinical translation.
title Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
A.1; I.2.0; I.4.6
url https://arxiv.org/abs/2506.10825