How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Qiaoyu, Wu, Chaoyi, Qiu, Pengcheng, Dai, Lisong, Zhang, Ya, Wang, Yanfeng, Xie, Weidi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909571041198080
author Zheng, Qiaoyu
Wu, Chaoyi
Qiu, Pengcheng
Dai, Lisong
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
author_facet Zheng, Qiaoyu
Wu, Chaoyi
Qiu, Pengcheng
Dai, Lisong
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
contents We introduce RadA-BenchPlat, an evaluation platform that benchmarks the performance of large language models (LLMs) act as agent cores in radiology environments using 2,200 radiologist-verified synthetic patient records covering six anatomical regions, five imaging modalities, and 2,200 disease scenarios, resulting in 24,200 question-answer pairs that simulate diverse clinical situations. The platform also defines ten categories of tools for agent-driven task solving and evaluates seven leading LLMs, revealing that while models like Claude-3.7-Sonnet can achieve a 67.1% task completion rate in routine settings, they still struggle with complex task understanding and tool coordination, limiting their capacity to serve as the central core of automated radiology systems. By incorporating four advanced prompt engineering strategies--where prompt-backpropagation and multi-agent collaboration contributed 16.8% and 30.7% improvements, respectively--the performance for complex tasks was enhanced by 48.2% overall. Furthermore, automated tool building was explored to improve robustness, achieving a 65.4% success rate, thereby offering promising insights for the future integration of fully automated radiology applications into clinical practice. All of our code and data are openly available at https://github.com/MAGIC-AI4Med/RadABench.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09529
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?
Zheng, Qiaoyu
Wu, Chaoyi
Qiu, Pengcheng
Dai, Lisong
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
Computer Vision and Pattern Recognition
We introduce RadA-BenchPlat, an evaluation platform that benchmarks the performance of large language models (LLMs) act as agent cores in radiology environments using 2,200 radiologist-verified synthetic patient records covering six anatomical regions, five imaging modalities, and 2,200 disease scenarios, resulting in 24,200 question-answer pairs that simulate diverse clinical situations. The platform also defines ten categories of tools for agent-driven task solving and evaluates seven leading LLMs, revealing that while models like Claude-3.7-Sonnet can achieve a 67.1% task completion rate in routine settings, they still struggle with complex task understanding and tool coordination, limiting their capacity to serve as the central core of automated radiology systems. By incorporating four advanced prompt engineering strategies--where prompt-backpropagation and multi-agent collaboration contributed 16.8% and 30.7% improvements, respectively--the performance for complex tasks was enhanced by 48.2% overall. Furthermore, automated tool building was explored to improve robustness, achieving a 65.4% success rate, thereby offering promising insights for the future integration of fully automated radiology applications into clinical practice. All of our code and data are openly available at https://github.com/MAGIC-AI4Med/RadABench.
title How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09529