Improving Clinical NLP Performance through Language Model-Generated Synthetic Clinical Data
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shan, Gallifant, Jack, Guevara, Marco, Gao, Yanjun, Afshar, Majid, Miller, Timothy, Dligach, Dmitriy, Bitterman, Danielle S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
by: Myers, Skatje, et al.
Published: (2025)
by: Myers, Skatje, et al.
Published: (2025)
Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
by: Khatwani, Saksham, et al.
Published: (2025)
by: Khatwani, Saksham, et al.
Published: (2025)
Lessons Learned on Information Retrieval in Electronic Health Records: A Comparison of Embedding Models and Pooling Strategies
by: Myers, Skatje, et al.
Published: (2024)
by: Myers, Skatje, et al.
Published: (2024)
LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval
by: Cheng, He, et al.
Published: (2026)
by: Cheng, He, et al.
Published: (2026)
Leveraging Medical Knowledge Graphs Into Large Language Models for Diagnosis Prediction: Design and Application Study
by: Gao, Yanjun, et al.
Published: (2023)
by: Gao, Yanjun, et al.
Published: (2023)
CLSGen: A Dual-Head Fine-Tuning Framework for Joint Probabilistic Classification and Verbalized Explanation
by: Yoon, WonJin, et al.
Published: (2026)
by: Yoon, WonJin, et al.
Published: (2026)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
by: Xiao, Yuxin, et al.
Published: (2025)
by: Xiao, Yuxin, et al.
Published: (2025)
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Sparse Autoencoder Features for Classifications and Transferability
by: Gallifant, Jack, et al.
Published: (2025)
by: Gallifant, Jack, et al.
Published: (2025)
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks
by: Gallifant, Jack, et al.
Published: (2024)
by: Gallifant, Jack, et al.
Published: (2024)
MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
by: Chen, Shan, et al.
Published: (2025)
by: Chen, Shan, et al.
Published: (2025)
ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
by: Chen, Canyu, et al.
Published: (2024)
by: Chen, Canyu, et al.
Published: (2024)
The use of large language models to enhance cancer clinical trial educational materials
by: Gao, Mingye, et al.
Published: (2024)
by: Gao, Mingye, et al.
Published: (2024)
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
Analyzing Diversity in Healthcare LLM Research: A Scientometric Perspective
by: Restrepo, David, et al.
Published: (2024)
by: Restrepo, David, et al.
Published: (2024)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
by: Qi, Jirui, et al.
Published: (2025)
by: Qi, Jirui, et al.
Published: (2025)
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
by: Sasse, Kuleen, et al.
Published: (2024)
by: Sasse, Kuleen, et al.
Published: (2024)
Principles from Clinical Research for NLP Model Generalization
by: Elangovan, Aparna, et al.
Published: (2023)
by: Elangovan, Aparna, et al.
Published: (2023)
EHRmonize: A Framework for Medical Concept Abstraction from Electronic Health Records using Large Language Models
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
by: Seah, Natalie, et al.
Published: (2026)
by: Seah, Natalie, et al.
Published: (2026)
Enhancing Clinical Documentation with Synthetic Data: Leveraging Generative Models for Improved Accuracy
by: Biswas, Anjanava, et al.
Published: (2024)
by: Biswas, Anjanava, et al.
Published: (2024)
Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification
by: Kruse, Maya, et al.
Published: (2025)
by: Kruse, Maya, et al.
Published: (2025)
Gender Bias in Large Language Models for Healthcare: Assignment Consistency and Clinical Implications
by: Liu, Mingxuan, et al.
Published: (2025)
by: Liu, Mingxuan, et al.
Published: (2025)
Seeds of Stereotypes: A Large-Scale Textual Analysis of Race and Gender Associations with Diseases in Online Sources
by: Hansen, Lasse Hyldig, et al.
Published: (2024)
by: Hansen, Lasse Hyldig, et al.
Published: (2024)
DART: A Structured Dataset of Regulatory Drug Documents in Italian for Clinical NLP
by: Barone, Mariano, et al.
Published: (2025)
by: Barone, Mariano, et al.
Published: (2025)
Less Context, Same Performance: A RAG Framework for Resource-Efficient LLM-Based Clinical NLP
by: Cheetirala, Satya Narayana, et al.
Published: (2025)
by: Cheetirala, Satya Narayana, et al.
Published: (2025)
DualAlign: Generating Clinically Grounded Synthetic Data
by: Li, Rumeng, et al.
Published: (2025)
by: Li, Rumeng, et al.
Published: (2025)
Give me Some Hard Questions: Synthetic Data Generation for Clinical QA
by: Bai, Fan, et al.
Published: (2024)
by: Bai, Fan, et al.
Published: (2024)
NagaNLP: Bootstrapping NLP for Low-Resource Nagamese Creole with Human-in-the-Loop Synthetic Data
by: Maiti, Agniva, et al.
Published: (2025)
by: Maiti, Agniva, et al.
Published: (2025)
Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages
by: McKenna, Nick, et al.
Published: (2025)
by: McKenna, Nick, et al.
Published: (2025)
Synthetic4Health: Generating Annotated Synthetic Clinical Letters
by: Ren, Libo, et al.
Published: (2024)
by: Ren, Libo, et al.
Published: (2024)
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes
by: Kweon, Sunjun, et al.
Published: (2023)
by: Kweon, Sunjun, et al.
Published: (2023)
Synthetic Feature Augmentation Improves Generalization Performance of Language Models
by: Choudhary, Ashok, et al.
Published: (2025)
by: Choudhary, Ashok, et al.
Published: (2025)
PRISM: A Transformer-based Language Model of Structured Clinical Event Data
by: Levine, Lionel, et al.
Published: (2025)
by: Levine, Lionel, et al.
Published: (2025)
CUICurate: A GraphRAG-based Framework for Automated Clinical Concept Curation for NLP applications
by: Blake, Victoria, et al.
Published: (2026)
by: Blake, Victoria, et al.
Published: (2026)
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
by: Guevara, Marco, et al.
Published: (2023)
by: Guevara, Marco, et al.
Published: (2023)
Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
by: Xu, Zerui, et al.
Published: (2024)
by: Xu, Zerui, et al.
Published: (2024)
Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction
by: Kruse, Maya, et al.
Published: (2025)
by: Kruse, Maya, et al.
Published: (2025)
Similar Items
-
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
by: Gao, Yanjun, et al.
Published: (2024) -
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
by: Gao, Yanjun, et al.
Published: (2024) -
Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
by: Myers, Skatje, et al.
Published: (2025) -
Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
by: Khatwani, Saksham, et al.
Published: (2025) -
Lessons Learned on Information Retrieval in Electronic Health Records: A Comparison of Embedding Models and Pooling Strategies
by: Myers, Skatje, et al.
Published: (2024)