Saved in:
Bibliographic Details
Main Authors: Xie, Ronald, Palayew, Steven, Toma, Augustin, Bader, Gary, Wang, Bo
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.14567
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917647717761024
author Xie, Ronald
Palayew, Steven
Toma, Augustin
Bader, Gary
Wang, Bo
author_facet Xie, Ronald
Palayew, Steven
Toma, Augustin
Bader, Gary
Wang, Bo
contents This paper outlines our submission to the MEDIQA2024 Multilingual and Multimodal Medical Answer Generation (M3G) shared task. We report results for two standalone solutions under the English category of the task, the first involving two consecutive API calls to the Claude 3 Opus API and the second involving training an image-disease label joint embedding in the style of CLIP for image classification. These two solutions scored 1st and 2nd place respectively on the competition leaderboard, substantially outperforming the next best solution. Additionally, we discuss insights gained from post-competition experiments. While the performance of these two solutions have significant room for improvement due to the difficulty of the shared task and the challenging nature of medical visual question answering in general, we identify the multi-stage LLM approach and the CLIP image classification approach as promising avenues for further investigation.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14567
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models
Xie, Ronald
Palayew, Steven
Toma, Augustin
Bader, Gary
Wang, Bo
Computation and Language
This paper outlines our submission to the MEDIQA2024 Multilingual and Multimodal Medical Answer Generation (M3G) shared task. We report results for two standalone solutions under the English category of the task, the first involving two consecutive API calls to the Claude 3 Opus API and the second involving training an image-disease label joint embedding in the style of CLIP for image classification. These two solutions scored 1st and 2nd place respectively on the competition leaderboard, substantially outperforming the next best solution. Additionally, we discuss insights gained from post-competition experiments. While the performance of these two solutions have significant room for improvement due to the difficulty of the shared task and the challenging nature of medical visual question answering in general, we identify the multi-stage LLM approach and the CLIP image classification approach as promising avenues for further investigation.
title WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2404.14567