MMedAgent: Learning to Use Medical Tools with Multi-modal Agent

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Binxu, Yan, Tiankai, Pan, Yuanting, Luo, Jie, Ji, Ruiyang, Ding, Jiayuan, Xu, Zhe, Liu, Shilong, Dong, Haoyu, Lin, Zihao, Wang, Yixin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913533098196992
author Li, Binxu
Yan, Tiankai
Pan, Yuanting
Luo, Jie
Ji, Ruiyang
Ding, Jiayuan
Xu, Zhe
Liu, Shilong
Dong, Haoyu
Lin, Zihao
Wang, Yixin
author_facet Li, Binxu
Yan, Tiankai
Pan, Yuanting
Luo, Jie
Ji, Ruiyang
Ding, Jiayuan
Xu, Zhe
Liu, Shilong
Dong, Haoyu
Lin, Zihao
Wang, Yixin
contents Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents have been developed to address these challenges by selecting appropriate specialized models as tools based on user inputs. However, such advancements have not been extensively explored within the medical domain. To bridge this gap, this paper introduces the first agent explicitly designed for the medical field, named \textbf{M}ulti-modal \textbf{Med}ical \textbf{Agent} (MMedAgent). We curate an instruction-tuning dataset comprising six medical tools solving seven tasks across five modalities, enabling the agent to choose the most suitable tools for a given task. Comprehensive experiments demonstrate that MMedAgent achieves superior performance across a variety of medical tasks compared to state-of-the-art open-source methods and even the closed-source model, GPT-4o. Furthermore, MMedAgent exhibits efficiency in updating and integrating new medical tools. Codes and models are all available.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02483
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
Li, Binxu
Yan, Tiankai
Pan, Yuanting
Luo, Jie
Ji, Ruiyang
Ding, Jiayuan
Xu, Zhe
Liu, Shilong
Dong, Haoyu
Lin, Zihao
Wang, Yixin
Computation and Language
Artificial Intelligence
Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents have been developed to address these challenges by selecting appropriate specialized models as tools based on user inputs. However, such advancements have not been extensively explored within the medical domain. To bridge this gap, this paper introduces the first agent explicitly designed for the medical field, named \textbf{M}ulti-modal \textbf{Med}ical \textbf{Agent} (MMedAgent). We curate an instruction-tuning dataset comprising six medical tools solving seven tasks across five modalities, enabling the agent to choose the most suitable tools for a given task. Comprehensive experiments demonstrate that MMedAgent achieves superior performance across a variety of medical tasks compared to state-of-the-art open-source methods and even the closed-source model, GPT-4o. Furthermore, MMedAgent exhibits efficiency in updating and integrating new medical tools. Codes and models are all available.
title MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.02483