Foundation Models in Medical Image Analysis: A Systematic Review and Meta-Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rajendran, Praveenbalaji, Safari, Mojtaba, He, Wenfeng, Hu, Mingzhe, Wang, Shansong, Zhou, Jun, Yang, Xiaofeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912659509608448
author Rajendran, Praveenbalaji
Safari, Mojtaba
He, Wenfeng
Hu, Mingzhe
Wang, Shansong
Zhou, Jun
Yang, Xiaofeng
author_facet Rajendran, Praveenbalaji
Safari, Mojtaba
He, Wenfeng
Hu, Mingzhe
Wang, Shansong
Zhou, Jun
Yang, Xiaofeng
contents Recent advancements in artificial intelligence (AI), particularly foundation models (FMs), have revolutionized medical image analysis, demonstrating strong zero- and few-shot performance across diverse medical imaging tasks, from segmentation to report generation. Unlike traditional task-specific AI models, FMs leverage large corpora of labeled and unlabeled multimodal datasets to learn generalized representations that can be adapted to various downstream clinical applications with minimal fine-tuning. However, despite the rapid proliferation of FM research in medical imaging, the field remains fragmented, lacking a unified synthesis that systematically maps the evolution of architectures, training paradigms, and clinical applications across modalities. To address this gap, this review article provides a comprehensive and structured analysis of FMs in medical image analysis. We systematically categorize studies into vision-only and vision-language FMs based on their architectural foundations, training strategies, and downstream clinical tasks. Additionally, a quantitative meta-analysis of the studies was conducted to characterize temporal trends in dataset utilization and application domains. We also critically discuss persistent challenges, including domain adaptation, efficient fine-tuning, computational constraints, and interpretability along with emerging solutions such as federated learning, knowledge distillation, and advanced prompting. Finally, we identify key future research directions aimed at enhancing the robustness, explainability, and clinical integration of FMs, thereby accelerating their translation into real-world medical practice.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16973
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Foundation Models in Medical Image Analysis: A Systematic Review and Meta-Analysis
Rajendran, Praveenbalaji
Safari, Mojtaba
He, Wenfeng
Hu, Mingzhe
Wang, Shansong
Zhou, Jun
Yang, Xiaofeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Medical Physics
Recent advancements in artificial intelligence (AI), particularly foundation models (FMs), have revolutionized medical image analysis, demonstrating strong zero- and few-shot performance across diverse medical imaging tasks, from segmentation to report generation. Unlike traditional task-specific AI models, FMs leverage large corpora of labeled and unlabeled multimodal datasets to learn generalized representations that can be adapted to various downstream clinical applications with minimal fine-tuning. However, despite the rapid proliferation of FM research in medical imaging, the field remains fragmented, lacking a unified synthesis that systematically maps the evolution of architectures, training paradigms, and clinical applications across modalities. To address this gap, this review article provides a comprehensive and structured analysis of FMs in medical image analysis. We systematically categorize studies into vision-only and vision-language FMs based on their architectural foundations, training strategies, and downstream clinical tasks. Additionally, a quantitative meta-analysis of the studies was conducted to characterize temporal trends in dataset utilization and application domains. We also critically discuss persistent challenges, including domain adaptation, efficient fine-tuning, computational constraints, and interpretability along with emerging solutions such as federated learning, knowledge distillation, and advanced prompting. Finally, we identify key future research directions aimed at enhancing the robustness, explainability, and clinical integration of FMs, thereby accelerating their translation into real-world medical practice.
title Foundation Models in Medical Image Analysis: A Systematic Review and Meta-Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Medical Physics
url https://arxiv.org/abs/2510.16973