Answering real-world clinical questions using large language model based systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Low, Yen Sia, Jackson, Michael L., Hyde, Rebecca J., Brown, Robert E., Sanghavi, Neil M., Baldwin, Julian D., Pike, C. William, Muralidharan, Jananee, Hui, Gavin, Alexander, Natasha, Hassan, Hadeel, Nene, Rahul V., Pike, Morgan, Pokrzywa, Courtney J., Vedak, Shivam, Yan, Adam Paul, Yao, Dong-han, Zipursky, Amy R., Dinh, Christina, Ballentine, Philip, Derieg, Dan C., Polony, Vladimir, Chawdry, Rehan N., Davies, Jordan, Hyde, Brigham B., Shah, Nigam H., Gombar, Saurabh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911938420670464
author Low, Yen Sia
Jackson, Michael L.
Hyde, Rebecca J.
Brown, Robert E.
Sanghavi, Neil M.
Baldwin, Julian D.
Pike, C. William
Muralidharan, Jananee
Hui, Gavin
Alexander, Natasha
Hassan, Hadeel
Nene, Rahul V.
Pike, Morgan
Pokrzywa, Courtney J.
Vedak, Shivam
Yan, Adam Paul
Yao, Dong-han
Zipursky, Amy R.
Dinh, Christina
Ballentine, Philip
Derieg, Dan C.
Polony, Vladimir
Chawdry, Rehan N.
Davies, Jordan
Hyde, Brigham B.
Shah, Nigam H.
Gombar, Saurabh
author_facet Low, Yen Sia
Jackson, Michael L.
Hyde, Rebecca J.
Brown, Robert E.
Sanghavi, Neil M.
Baldwin, Julian D.
Pike, C. William
Muralidharan, Jananee
Hui, Gavin
Alexander, Natasha
Hassan, Hadeel
Nene, Rahul V.
Pike, Morgan
Pokrzywa, Courtney J.
Vedak, Shivam
Yan, Adam Paul
Yao, Dong-han
Zipursky, Amy R.
Dinh, Christina
Ballentine, Philip
Derieg, Dan C.
Polony, Vladimir
Chawdry, Rehan N.
Davies, Jordan
Hyde, Brigham B.
Shah, Nigam H.
Gombar, Saurabh
contents Evidence to guide healthcare decisions is often limited by a lack of relevant and trustworthy literature as well as difficulty in contextualizing existing research for a specific patient. Large language models (LLMs) could potentially address both challenges by either summarizing published literature or generating new studies based on real-world data (RWD). We evaluated the ability of five LLM-based systems in answering 50 clinical questions and had nine independent physicians review the responses for relevance, reliability, and actionability. As it stands, general-purpose LLMs (ChatGPT-4, Claude 3 Opus, Gemini Pro 1.5) rarely produced answers that were deemed relevant and evidence-based (2% - 10%). In contrast, retrieval augmented generation (RAG)-based and agentic LLM systems produced relevant and evidence-based answers for 24% (OpenEvidence) to 58% (ChatRWD) of questions. Only the agentic ChatRWD was able to answer novel questions compared to other LLMs (65% vs. 0-9%). These results suggest that while general-purpose LLMs should not be used as-is, a purpose-built system for evidence summarization based on RAG and one for generating novel evidence working synergistically would improve availability of pertinent evidence for patient care.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00541
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Answering real-world clinical questions using large language model based systems
Low, Yen Sia
Jackson, Michael L.
Hyde, Rebecca J.
Brown, Robert E.
Sanghavi, Neil M.
Baldwin, Julian D.
Pike, C. William
Muralidharan, Jananee
Hui, Gavin
Alexander, Natasha
Hassan, Hadeel
Nene, Rahul V.
Pike, Morgan
Pokrzywa, Courtney J.
Vedak, Shivam
Yan, Adam Paul
Yao, Dong-han
Zipursky, Amy R.
Dinh, Christina
Ballentine, Philip
Derieg, Dan C.
Polony, Vladimir
Chawdry, Rehan N.
Davies, Jordan
Hyde, Brigham B.
Shah, Nigam H.
Gombar, Saurabh
Computation and Language
Artificial Intelligence
Information Retrieval
Evidence to guide healthcare decisions is often limited by a lack of relevant and trustworthy literature as well as difficulty in contextualizing existing research for a specific patient. Large language models (LLMs) could potentially address both challenges by either summarizing published literature or generating new studies based on real-world data (RWD). We evaluated the ability of five LLM-based systems in answering 50 clinical questions and had nine independent physicians review the responses for relevance, reliability, and actionability. As it stands, general-purpose LLMs (ChatGPT-4, Claude 3 Opus, Gemini Pro 1.5) rarely produced answers that were deemed relevant and evidence-based (2% - 10%). In contrast, retrieval augmented generation (RAG)-based and agentic LLM systems produced relevant and evidence-based answers for 24% (OpenEvidence) to 58% (ChatRWD) of questions. Only the agentic ChatRWD was able to answer novel questions compared to other LLMs (65% vs. 0-9%). These results suggest that while general-purpose LLMs should not be used as-is, a purpose-built system for evidence summarization based on RAG and one for generating novel evidence working synergistically would improve availability of pertinent evidence for patient care.
title Answering real-world clinical questions using large language model based systems
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2407.00541