Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Hanyin, Gao, Chufan, Liu, Bolun, Xu, Qiping, Hussein, Guleid, Labban, Mohamad El, Iheasirim, Kingsley, Korsapati, Hariprasad, Outcalt, Chuck, Sun, Jimeng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918035560857600
author Wang, Hanyin
Gao, Chufan
Liu, Bolun
Xu, Qiping
Hussein, Guleid
Labban, Mohamad El
Iheasirim, Kingsley
Korsapati, Hariprasad
Outcalt, Chuck
Sun, Jimeng
author_facet Wang, Hanyin
Gao, Chufan
Liu, Bolun
Xu, Qiping
Hussein, Guleid
Labban, Mohamad El
Iheasirim, Kingsley
Korsapati, Hariprasad
Outcalt, Chuck
Sun, Jimeng
contents Proprietary Large Language Models (LLMs) such as GPT-4 and Gemini have demonstrated promising capabilities in clinical text summarization tasks. However, due to patient data privacy concerns and computational costs, many healthcare providers prefer using small, locally-hosted models over external generic LLMs. This study presents a comprehensive domain- and task-specific adaptation process for the open-source LLaMA-2 13 billion parameter model, enabling it to generate high-quality clinical notes from outpatient patient-doctor dialogues. Our process incorporates continued pretraining, supervised fine-tuning, and reinforcement learning from both AI and human feedback. We introduced a new approach, DistillDirect, for performing on-policy reinforcement learning with Gemini 1.0 Pro as the teacher model. Our resulting model, LLaMA-Clinic, can generate clinical notes comparable in quality to those authored by physicians. In a blinded physician reader study, the majority (92.8%) of individual evaluations rated the notes generated by LLaMA-Clinic as "acceptable" or higher across three criteria: real-world readiness, completeness, and accuracy. In the more challenging "Assessment and Plan" section, LLaMA-Clinic matched physician-authored notes in real-world readiness score. We highlight key considerations for future clinical note-generation tasks, emphasizing the importance of pre-defining a "best practice" note format, rather than relying on LLMs to determine this for clinical practice.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00715
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
Wang, Hanyin
Gao, Chufan
Liu, Bolun
Xu, Qiping
Hussein, Guleid
Labban, Mohamad El
Iheasirim, Kingsley
Korsapati, Hariprasad
Outcalt, Chuck
Sun, Jimeng
Computation and Language
Artificial Intelligence
Machine Learning
Proprietary Large Language Models (LLMs) such as GPT-4 and Gemini have demonstrated promising capabilities in clinical text summarization tasks. However, due to patient data privacy concerns and computational costs, many healthcare providers prefer using small, locally-hosted models over external generic LLMs. This study presents a comprehensive domain- and task-specific adaptation process for the open-source LLaMA-2 13 billion parameter model, enabling it to generate high-quality clinical notes from outpatient patient-doctor dialogues. Our process incorporates continued pretraining, supervised fine-tuning, and reinforcement learning from both AI and human feedback. We introduced a new approach, DistillDirect, for performing on-policy reinforcement learning with Gemini 1.0 Pro as the teacher model. Our resulting model, LLaMA-Clinic, can generate clinical notes comparable in quality to those authored by physicians. In a blinded physician reader study, the majority (92.8%) of individual evaluations rated the notes generated by LLaMA-Clinic as "acceptable" or higher across three criteria: real-world readiness, completeness, and accuracy. In the more challenging "Assessment and Plan" section, LLaMA-Clinic matched physician-authored notes in real-world readiness score. We highlight key considerations for future clinical note-generation tasks, emphasizing the importance of pre-defining a "best practice" note format, rather than relying on LLMs to determine this for clinical practice.
title Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.00715