CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Mingyu Derek, Ye, Chenchen, Yan, Yu, Wang, Xiaoxuan, Ping, Peipei, Chang, Timothy S, Wang, Wei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929538323185664
author Ma, Mingyu Derek
Ye, Chenchen
Yan, Yu
Wang, Xiaoxuan
Ping, Peipei
Chang, Timothy S
Wang, Wei
author_facet Ma, Mingyu Derek
Ye, Chenchen
Yan, Yu
Wang, Xiaoxuan
Ping, Peipei
Chang, Timothy S
Wang, Wei
contents The integration of Artificial Intelligence (AI), especially Large Language Models (LLMs), into the clinical diagnosis process offers significant potential to improve the efficiency and accessibility of medical care. While LLMs have shown some promise in the medical domain, their application in clinical diagnosis remains underexplored, especially in real-world clinical practice, where highly sophisticated, patient-specific decisions need to be made. Current evaluations of LLMs in this field are often narrow in scope, focusing on specific diseases or specialties and employing simplified diagnostic tasks. To bridge this gap, we introduce CliBench, a novel benchmark developed from the MIMIC IV dataset, offering a comprehensive and realistic assessment of LLMs' capabilities in clinical diagnosis. This benchmark not only covers diagnoses from a diverse range of medical cases across various specialties but also incorporates tasks of clinical significance: treatment procedure identification, lab test ordering and medication prescriptions. Supported by structured output ontologies, CliBench enables a precise and multi-granular evaluation, offering an in-depth understanding of LLM's capability on diverse clinical tasks of desired granularity. We conduct a zero-shot evaluation of leading LLMs to assess their proficiency in clinical decision-making. Our preliminary results shed light on the potential and limitations of current LLMs in clinical settings, providing valuable insights for future advancements in LLM-powered healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09923
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making
Ma, Mingyu Derek
Ye, Chenchen
Yan, Yu
Wang, Xiaoxuan
Ping, Peipei
Chang, Timothy S
Wang, Wei
Computation and Language
Artificial Intelligence
Machine Learning
The integration of Artificial Intelligence (AI), especially Large Language Models (LLMs), into the clinical diagnosis process offers significant potential to improve the efficiency and accessibility of medical care. While LLMs have shown some promise in the medical domain, their application in clinical diagnosis remains underexplored, especially in real-world clinical practice, where highly sophisticated, patient-specific decisions need to be made. Current evaluations of LLMs in this field are often narrow in scope, focusing on specific diseases or specialties and employing simplified diagnostic tasks. To bridge this gap, we introduce CliBench, a novel benchmark developed from the MIMIC IV dataset, offering a comprehensive and realistic assessment of LLMs' capabilities in clinical diagnosis. This benchmark not only covers diagnoses from a diverse range of medical cases across various specialties but also incorporates tasks of clinical significance: treatment procedure identification, lab test ordering and medication prescriptions. Supported by structured output ontologies, CliBench enables a precise and multi-granular evaluation, offering an in-depth understanding of LLM's capability on diverse clinical tasks of desired granularity. We conduct a zero-shot evaluation of leading LLMs to assess their proficiency in clinical decision-making. Our preliminary results shed light on the potential and limitations of current LLMs in clinical settings, providing valuable insights for future advancements in LLM-powered healthcare.
title CliBench: A Multifaceted and Multigranular Evaluation of Large Language Models for Clinical Decision Making
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.09923