Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Nicholas, Jin, Qiao, Chan, Joey, Xiong, Guangzhi, Applebaum, Serina, Gilson, Aidan, McMurry, Reid, Taylor, R. Andrew, Zhang, Aidong, Chen, Qingyu, Lu, Zhiyong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915208789753856
author Wan, Nicholas
Jin, Qiao
Chan, Joey
Xiong, Guangzhi
Applebaum, Serina
Gilson, Aidan
McMurry, Reid
Taylor, R. Andrew
Zhang, Aidong
Chen, Qingyu
Lu, Zhiyong
author_facet Wan, Nicholas
Jin, Qiao
Chan, Joey
Xiong, Guangzhi
Applebaum, Serina
Gilson, Aidan
McMurry, Reid
Taylor, R. Andrew
Zhang, Aidong
Chen, Qingyu
Lu, Zhiyong
contents Although large language models (LLMs) have been assessed for general medical knowledge using licensing exams, their ability to support clinical decision-making, such as selecting medical calculators, remains uncertain. We assessed nine LLMs, including open-source, proprietary, and domain-specific models, with 1,009 multiple-choice question-answer pairs across 35 clinical calculators and compared LLMs to humans on a subset of questions. While the highest-performing LLM, OpenAI o1, provided an answer accuracy of 66.0% (CI: 56.7-75.3%) on the subset of 100 questions, two human annotators nominally outperformed LLMs with an average answer accuracy of 79.5% (CI: 73.5-85.0%). Ultimately, we evaluated medical trainees and LLMs in recommending medical calculators across clinical scenarios like risk stratification and diagnosis. With error analysis showing that the highest-performing LLMs continue to make mistakes in comprehension (49.3% of errors) and calculator knowledge (7.1% of errors), our findings highlight that LLMs are not superior to humans in calculator recommendation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05897
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators
Wan, Nicholas
Jin, Qiao
Chan, Joey
Xiong, Guangzhi
Applebaum, Serina
Gilson, Aidan
McMurry, Reid
Taylor, R. Andrew
Zhang, Aidong
Chen, Qingyu
Lu, Zhiyong
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Although large language models (LLMs) have been assessed for general medical knowledge using licensing exams, their ability to support clinical decision-making, such as selecting medical calculators, remains uncertain. We assessed nine LLMs, including open-source, proprietary, and domain-specific models, with 1,009 multiple-choice question-answer pairs across 35 clinical calculators and compared LLMs to humans on a subset of questions. While the highest-performing LLM, OpenAI o1, provided an answer accuracy of 66.0% (CI: 56.7-75.3%) on the subset of 100 questions, two human annotators nominally outperformed LLMs with an average answer accuracy of 79.5% (CI: 73.5-85.0%). Ultimately, we evaluated medical trainees and LLMs in recommending medical calculators across clinical scenarios like risk stratification and diagnosis. With error analysis showing that the highest-performing LLMs continue to make mistakes in comprehension (49.3% of errors) and calculator knowledge (7.1% of errors), our findings highlight that LLMs are not superior to humans in calculator recommendation.
title Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2411.05897