A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Jie, Wang, Wenxuan, Su, Yihang, Huan, Jingyuan, Chen, Wenting, Zhang, Yudi, Li, Cheng-Yi, Chang, Kao-Jung, Xin, Xiaohan, Shen, Linlin, Lyu, Michael R.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913589262024704
author Liu, Jie
Wang, Wenxuan
Su, Yihang
Huan, Jingyuan
Chen, Wenting
Zhang, Yudi
Li, Cheng-Yi
Chang, Kao-Jung
Xin, Xiaohan
Shen, Linlin
Lyu, Michael R.
author_facet Liu, Jie
Wang, Wenxuan
Su, Yihang
Huan, Jingyuan
Chen, Wenting
Zhang, Yudi
Li, Cheng-Yi
Chang, Kao-Jung
Xin, Xiaohan
Shen, Linlin
Lyu, Michael R.
contents The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gastroenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11217
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
Liu, Jie
Wang, Wenxuan
Su, Yihang
Huan, Jingyuan
Chen, Wenting
Zhang, Yudi
Li, Cheng-Yi
Chang, Kao-Jung
Xin, Xiaohan
Shen, Linlin
Lyu, Michael R.
Computation and Language
Computer Vision and Pattern Recognition
The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gastroenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments.
title A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.11217