Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhong, Yuan, Jin, Ruinan, Dou, Qi, Li, Xiaoxiao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917367151329280
author Zhong, Yuan
Jin, Ruinan
Dou, Qi
Li, Xiaoxiao
author_facet Zhong, Yuan
Jin, Ruinan
Dou, Qi
Li, Xiaoxiao
contents Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing specialist medical VLMs requires substantial computational resources and carefully curated datasets, and it remains unclear under which conditions generalist and specialist medical VLMs each perform best. This study highlights the complementary strengths of specialist medical and generalist VLMs. Specialists remain valuable in modality-aligned use cases, but we find that efficiently fine-tuned generalist VLMs can achieve comparable or even superior performance in most tasks, particularly when transferring to unseen or rare OOD medical modalities. These results suggest that generalist VLMs, rather than being constrained by their lack of specialist medical pretraining, may offer a scalable and cost-effective pathway for advancing clinical AI development.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17337
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
Zhong, Yuan
Jin, Ruinan
Dou, Qi
Li, Xiaoxiao
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing specialist medical VLMs requires substantial computational resources and carefully curated datasets, and it remains unclear under which conditions generalist and specialist medical VLMs each perform best. This study highlights the complementary strengths of specialist medical and generalist VLMs. Specialists remain valuable in modality-aligned use cases, but we find that efficiently fine-tuned generalist VLMs can achieve comparable or even superior performance in most tasks, particularly when transferring to unseen or rare OOD medical modalities. These results suggest that generalist VLMs, rather than being constrained by their lack of specialist medical pretraining, may offer a scalable and cost-effective pathway for advancing clinical AI development.
title Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.17337