A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yingshu, Liu, Yunyi, Wang, Zhanyu, Liang, Xinyu, Wang, Lei, Liu, Lingqiao, Cui, Leyang, Tu, Zhaopeng, Wang, Longyue, Zhou, Luping
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911767918018560
author Li, Yingshu
Liu, Yunyi
Wang, Zhanyu
Liang, Xinyu
Wang, Lei
Liu, Lingqiao
Cui, Leyang
Tu, Zhaopeng
Wang, Longyue
Zhou, Luping
author_facet Li, Yingshu
Liu, Yunyi
Wang, Zhanyu
Liang, Xinyu
Wang, Lei
Liu, Lingqiao
Cui, Leyang
Tu, Zhaopeng
Wang, Longyue
Zhou, Luping
contents This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For the evaluation, a set of prompts is designed for each task to induce the corresponding capability of GPT-4V to produce sufficiently good outputs. Three evaluation ways including quantitative analysis, human evaluation, and case study are employed to achieve an in-depth and extensive evaluation. Our evaluation shows that GPT-4V excels in understanding medical images and is able to generate high-quality radiology reports and effectively answer questions about medical images. Meanwhile, it is found that its performance for medical visual grounding needs to be substantially improved. In addition, we observe the discrepancy between the evaluation outcome from quantitative analysis and that from human evaluation. This discrepancy suggests the limitations of conventional metrics in assessing the performance of large language models like GPT-4V and the necessity of developing new metrics for automatic quantitative analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2310_20381
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
Li, Yingshu
Liu, Yunyi
Wang, Zhanyu
Liang, Xinyu
Wang, Lei
Liu, Lingqiao
Cui, Leyang
Tu, Zhaopeng
Wang, Longyue
Zhou, Luping
Computer Vision and Pattern Recognition
Artificial Intelligence
This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For the evaluation, a set of prompts is designed for each task to induce the corresponding capability of GPT-4V to produce sufficiently good outputs. Three evaluation ways including quantitative analysis, human evaluation, and case study are employed to achieve an in-depth and extensive evaluation. Our evaluation shows that GPT-4V excels in understanding medical images and is able to generate high-quality radiology reports and effectively answer questions about medical images. Meanwhile, it is found that its performance for medical visual grounding needs to be substantially improved. In addition, we observe the discrepancy between the evaluation outcome from quantitative analysis and that from human evaluation. This discrepancy suggests the limitations of conventional metrics in assessing the performance of large language models like GPT-4V and the necessity of developing new metrics for automatic quantitative analysis.
title A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2310.20381