Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Zerui, Wu, Fang, Lu, Yingzhou, Zhang, Yuanyuan, Zhao, Yue
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912982379790336
author Xu, Zerui
Wu, Fang
Lu, Yingzhou
Zhang, Yuanyuan
Zhao, Yue
author_facet Xu, Zerui
Wu, Fang
Lu, Yingzhou
Zhang, Yuanyuan
Zhao, Yue
contents Machine learning (ML) holds great promise for clinical applications but is often hindered by limited access to high-quality data due to privacy concerns, high costs, and long timelines associated with clinical trials. While large language models (LLMs) have demonstrated strong performance in general-purpose generation tasks, their application to synthesizing realistic clinical trials remains underexplored. In this work, we propose a novel Retrieval-Reasoning framework that leverages few-shot prompting with LLMs to generate synthetic clinical trial reports annotated with binary success/failure outcomes. Our approach integrates a retrieval module to ground the generation on relevant trial data and a reasoning module to ensure domain-consistent justifications. Experiments conducted on real clinical trials from the ClinicalTrials.gov database demonstrate that the generated synthetic trials effectively augment real datasets. Fine-tuning a BioBERT classifier on synthetic data, real data, or their combination shows that hybrid fine-tuning leads to improved performance on clinical trial outcome prediction tasks. Our results suggest that LLM-based synthetic data can serve as a powerful tool for privacy-preserving data augmentation in clinical research. The code is available at https://github.com/XuZR3x/Retrieval_Reasoning_Clinical_Trial_Generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12476
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
Xu, Zerui
Wu, Fang
Lu, Yingzhou
Zhang, Yuanyuan
Zhao, Yue
Computation and Language
Machine Learning
Machine learning (ML) holds great promise for clinical applications but is often hindered by limited access to high-quality data due to privacy concerns, high costs, and long timelines associated with clinical trials. While large language models (LLMs) have demonstrated strong performance in general-purpose generation tasks, their application to synthesizing realistic clinical trials remains underexplored. In this work, we propose a novel Retrieval-Reasoning framework that leverages few-shot prompting with LLMs to generate synthetic clinical trial reports annotated with binary success/failure outcomes. Our approach integrates a retrieval module to ground the generation on relevant trial data and a reasoning module to ensure domain-consistent justifications. Experiments conducted on real clinical trials from the ClinicalTrials.gov database demonstrate that the generated synthetic trials effectively augment real datasets. Fine-tuning a BioBERT classifier on synthetic data, real data, or their combination shows that hybrid fine-tuning leads to improved performance on clinical trial outcome prediction tasks. Our results suggest that LLM-based synthetic data can serve as a powerful tool for privacy-preserving data augmentation in clinical research. The code is available at https://github.com/XuZR3x/Retrieval_Reasoning_Clinical_Trial_Generation.
title Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.12476