Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cho, Dongkyu, Zhang, Miao, Chunara, Rumi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914057229959168
author Cho, Dongkyu
Zhang, Miao
Chunara, Rumi
author_facet Cho, Dongkyu
Zhang, Miao
Chunara, Rumi
contents Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Experiments on clinical prediction tasks demonstrate that our lightweight collaboration-based approach consistently outperforms existing LLM augmentation methods while improving safety through reduced factual errors. This framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21530
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
Cho, Dongkyu
Zhang, Miao
Chunara, Rumi
Machine Learning
Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Experiments on clinical prediction tasks demonstrate that our lightweight collaboration-based approach consistently outperforms existing LLM augmentation methods while improving safety through reduced factual errors. This framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.
title Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
topic Machine Learning
url https://arxiv.org/abs/2509.21530