A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913589262024704 |
|---|---|
| author | Liu, Jie Wang, Wenxuan Su, Yihang Huan, Jingyuan Chen, Wenting Zhang, Yudi Li, Cheng-Yi Chang, Kao-Jung Xin, Xiaohan Shen, Linlin Lyu, Michael R. |
| author_facet | Liu, Jie Wang, Wenxuan Su, Yihang Huan, Jingyuan Chen, Wenting Zhang, Yudi Li, Cheng-Yi Chang, Kao-Jung Xin, Xiaohan Shen, Linlin Lyu, Michael R. |
| contents | The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gastroenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_11217 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models Liu, Jie Wang, Wenxuan Su, Yihang Huan, Jingyuan Chen, Wenting Zhang, Yudi Li, Cheng-Yi Chang, Kao-Jung Xin, Xiaohan Shen, Linlin Lyu, Michael R. Computation and Language Computer Vision and Pattern Recognition The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gastroenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments. |
| title | A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2402.11217 |