Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Van Veen, Dave, Van Uden, Cara, Blankemeier, Louis, Delbrouck, Jean-Benoit, Aali, Asad, Bluethgen, Christian, Pareek, Anuj, Polacin, Malgorzata, Reis, Eduardo Pontes, Seehofnerova, Anna, Rohatgi, Nidhi, Hosamani, Poonam, Collins, William, Ahuja, Neera, Langlotz, Curtis P., Hom, Jason, Gatidis, Sergios, Pauly, John, Chaudhari, Akshay S.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916201917054976
author Van Veen, Dave
Van Uden, Cara
Blankemeier, Louis
Delbrouck, Jean-Benoit
Aali, Asad
Bluethgen, Christian
Pareek, Anuj
Polacin, Malgorzata
Reis, Eduardo Pontes
Seehofnerova, Anna
Rohatgi, Nidhi
Hosamani, Poonam
Collins, William
Ahuja, Neera
Langlotz, Curtis P.
Hom, Jason
Gatidis, Sergios
Pauly, John
Chaudhari, Akshay S.
author_facet Van Veen, Dave
Van Uden, Cara
Blankemeier, Louis
Delbrouck, Jean-Benoit
Aali, Asad
Bluethgen, Christian
Pareek, Anuj
Polacin, Malgorzata
Reis, Eduardo Pontes
Seehofnerova, Anna
Rohatgi, Nidhi
Hosamani, Poonam
Collins, William
Ahuja, Neera
Langlotz, Curtis P.
Hom, Jason
Gatidis, Sergios
Pauly, John
Chaudhari, Akshay S.
contents Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language processing (NLP), their effectiveness on a diverse range of clinical summarization tasks remains unproven. In this study, we apply adaptation methods to eight LLMs, spanning four distinct clinical summarization tasks: radiology reports, patient questions, progress notes, and doctor-patient dialogue. Quantitative assessments with syntactic, semantic, and conceptual NLP metrics reveal trade-offs between models and adaptation methods. A clinical reader study with ten physicians evaluates summary completeness, correctness, and conciseness; in a majority of cases, summaries from our best adapted LLMs are either equivalent (45%) or superior (36%) compared to summaries from medical experts. The ensuing safety analysis highlights challenges faced by both LLMs and medical experts, as we connect errors to potential medical harm and categorize types of fabricated information. Our research provides evidence of LLMs outperforming medical experts in clinical text summarization across multiple tasks. This suggests that integrating LLMs into clinical workflows could alleviate documentation burden, allowing clinicians to focus more on patient care.
format Preprint
id arxiv_https___arxiv_org_abs_2309_07430
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
Van Veen, Dave
Van Uden, Cara
Blankemeier, Louis
Delbrouck, Jean-Benoit
Aali, Asad
Bluethgen, Christian
Pareek, Anuj
Polacin, Malgorzata
Reis, Eduardo Pontes
Seehofnerova, Anna
Rohatgi, Nidhi
Hosamani, Poonam
Collins, William
Ahuja, Neera
Langlotz, Curtis P.
Hom, Jason
Gatidis, Sergios
Pauly, John
Chaudhari, Akshay S.
Computation and Language
Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language processing (NLP), their effectiveness on a diverse range of clinical summarization tasks remains unproven. In this study, we apply adaptation methods to eight LLMs, spanning four distinct clinical summarization tasks: radiology reports, patient questions, progress notes, and doctor-patient dialogue. Quantitative assessments with syntactic, semantic, and conceptual NLP metrics reveal trade-offs between models and adaptation methods. A clinical reader study with ten physicians evaluates summary completeness, correctness, and conciseness; in a majority of cases, summaries from our best adapted LLMs are either equivalent (45%) or superior (36%) compared to summaries from medical experts. The ensuing safety analysis highlights challenges faced by both LLMs and medical experts, as we connect errors to potential medical harm and categorize types of fabricated information. Our research provides evidence of LLMs outperforming medical experts in clinical text summarization across multiple tasks. This suggests that integrating LLMs into clinical workflows could alleviate documentation burden, allowing clinicians to focus more on patient care.
title Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
topic Computation and Language
url https://arxiv.org/abs/2309.07430