Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Ran, Cui, Hejie, Yu, Yue, Kan, Xuan, Shi, Wenqi, Zhuang, Yuchen, Jin, Wei, Ho, Joyce, Yang, Carl
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909465539772416
author Xu, Ran
Cui, Hejie
Yu, Yue
Kan, Xuan
Shi, Wenqi
Zhuang, Yuchen
Jin, Wei
Ho, Joyce
Yang, Carl
author_facet Xu, Ran
Cui, Hejie
Yu, Yue
Kan, Xuan
Shi, Wenqi
Zhuang, Yuchen
Jin, Wei
Ho, Joyce
Yang, Carl
contents Clinical natural language processing requires methods that can address domain-specific challenges, such as complex medical terminology and clinical contexts. Recently, large language models (LLMs) have shown promise in this domain. Yet, their direct deployment can lead to privacy issues and are constrained by resources. To address this challenge, we delve into synthetic clinical text generation using LLMs for clinical NLP tasks. We propose an innovative, resource-efficient approach, ClinGen, which infuses knowledge into the process. Our model involves clinical knowledge extraction and context-informed LLM prompting. Both clinical topics and writing styles are drawn from external domain-specific knowledge graphs and LLMs to guide data generation. Our extensive empirical study across 7 clinical NLP tasks and 16 datasets reveals that ClinGen consistently enhances performance across various tasks, effectively aligning the distribution of real datasets and significantly enriching the diversity of generated training instances. Our code is available at \url{https://github.com/ritaranx/ClinGen}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_00287
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
Xu, Ran
Cui, Hejie
Yu, Yue
Kan, Xuan
Shi, Wenqi
Zhuang, Yuchen
Jin, Wei
Ho, Joyce
Yang, Carl
Computation and Language
Artificial Intelligence
Machine Learning
Quantitative Methods
Clinical natural language processing requires methods that can address domain-specific challenges, such as complex medical terminology and clinical contexts. Recently, large language models (LLMs) have shown promise in this domain. Yet, their direct deployment can lead to privacy issues and are constrained by resources. To address this challenge, we delve into synthetic clinical text generation using LLMs for clinical NLP tasks. We propose an innovative, resource-efficient approach, ClinGen, which infuses knowledge into the process. Our model involves clinical knowledge extraction and context-informed LLM prompting. Both clinical topics and writing styles are drawn from external domain-specific knowledge graphs and LLMs to guide data generation. Our extensive empirical study across 7 clinical NLP tasks and 16 datasets reveals that ClinGen consistently enhances performance across various tasks, effectively aligning the distribution of real datasets and significantly enriching the diversity of generated training instances. Our code is available at \url{https://github.com/ritaranx/ClinGen}.
title Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2311.00287