A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives
Fuente:
arXiv
Salvato in:
| Autori principali: | Rao, Hanshu, Liu, Weisi, Wang, Haohan, Huang, I-Chan, He, Zhe, Huang, Xiaolei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Time Matters: Examine Temporal Effects on Biomedical Language Models
di: Liu, Weisi, et al.
Pubblicazione: (2024)
di: Liu, Weisi, et al.
Pubblicazione: (2024)
Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
di: Han, Guangzeng, et al.
Pubblicazione: (2025)
di: Han, Guangzeng, et al.
Pubblicazione: (2025)
Model-Agnostic Meta Learning for Class Imbalance Adaptation
di: Rao, Hanshu, et al.
Pubblicazione: (2026)
di: Rao, Hanshu, et al.
Pubblicazione: (2026)
Chain-of-Interaction: Enhancing Large Language Models for Psychiatric Behavior Understanding by Dyadic Contexts
di: Han, Guangzeng, et al.
Pubblicazione: (2024)
di: Han, Guangzeng, et al.
Pubblicazione: (2024)
Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation
di: Liu, Weisi, et al.
Pubblicazione: (2026)
di: Liu, Weisi, et al.
Pubblicazione: (2026)
Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts
di: Liu, Weisi, et al.
Pubblicazione: (2025)
di: Liu, Weisi, et al.
Pubblicazione: (2025)
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
di: Han, Guangzeng, et al.
Pubblicazione: (2026)
di: Han, Guangzeng, et al.
Pubblicazione: (2026)
Examining Imbalance Effects on Performance and Demographic Fairness of Clinical Language Models
di: Jones, Precious, et al.
Pubblicazione: (2024)
di: Jones, Precious, et al.
Pubblicazione: (2024)
A Structure-aware Generative Model for Biomedical Event Extraction
di: Yuan, Haohan, et al.
Pubblicazione: (2024)
di: Yuan, Haohan, et al.
Pubblicazione: (2024)
Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection
di: You, Saukun Thika, et al.
Pubblicazione: (2026)
di: You, Saukun Thika, et al.
Pubblicazione: (2026)
Evaluating Language Models as Synthetic Data Generators
di: Kim, Seungone, et al.
Pubblicazione: (2024)
di: Kim, Seungone, et al.
Pubblicazione: (2024)
Synth-Empathy: Towards High-Quality Synthetic Empathy Data
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
DataGen: Unified Synthetic Dataset Generation via Large Language Models
di: Huang, Yue, et al.
Pubblicazione: (2024)
di: Huang, Yue, et al.
Pubblicazione: (2024)
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
di: Zhang, Jingyi, et al.
Pubblicazione: (2026)
di: Zhang, Jingyi, et al.
Pubblicazione: (2026)
WHERE and WHICH: Iterative Debate for Biomedical Synthetic Data Augmentation
di: Zhao, Zhengyi, et al.
Pubblicazione: (2025)
di: Zhao, Zhengyi, et al.
Pubblicazione: (2025)
Alleviating Distribution Shift in Synthetic Data for Machine Translation Quality Estimation
di: Geng, Xiang, et al.
Pubblicazione: (2025)
di: Geng, Xiang, et al.
Pubblicazione: (2025)
TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data
di: Huang, Xiang, et al.
Pubblicazione: (2024)
di: Huang, Xiang, et al.
Pubblicazione: (2024)
Scaling Laws of Synthetic Data for Language Models
di: Qin, Zeyu, et al.
Pubblicazione: (2025)
di: Qin, Zeyu, et al.
Pubblicazione: (2025)
Can Large Language Models Replace Data Scientists in Biomedical Research?
di: Wang, Zifeng, et al.
Pubblicazione: (2024)
di: Wang, Zifeng, et al.
Pubblicazione: (2024)
DKE-Research at SemEval-2024 Task 2: Incorporating Data Augmentation with Generative Models and Biomedical Knowledge to Enhance Inference Robustness
di: Wang, Yuqi, et al.
Pubblicazione: (2024)
di: Wang, Yuqi, et al.
Pubblicazione: (2024)
Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
di: Dhawan, Aashish, et al.
Pubblicazione: (2026)
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
di: Miranda, Lester James V., et al.
Pubblicazione: (2026)
di: Miranda, Lester James V., et al.
Pubblicazione: (2026)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
di: Sajith, Aryan, et al.
Pubblicazione: (2024)
di: Sajith, Aryan, et al.
Pubblicazione: (2024)
Towards Active Synthetic Data Generation for Finetuning Language Models
di: Kessler, Samuel, et al.
Pubblicazione: (2025)
di: Kessler, Samuel, et al.
Pubblicazione: (2025)
Generating High Quality Synthetic Data for Dutch Medical Conversations
di: Kuan, Cecilia, et al.
Pubblicazione: (2026)
di: Kuan, Cecilia, et al.
Pubblicazione: (2026)
Entropy-Based Data Selection for Language Models
di: Li, Hongming, et al.
Pubblicazione: (2026)
di: Li, Hongming, et al.
Pubblicazione: (2026)
ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis
di: Tu, Zeao, et al.
Pubblicazione: (2024)
di: Tu, Zeao, et al.
Pubblicazione: (2024)
Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection
di: Chen, Ruibo, et al.
Pubblicazione: (2024)
di: Chen, Ruibo, et al.
Pubblicazione: (2024)
DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models
di: Zhou, Ying, et al.
Pubblicazione: (2024)
di: Zhou, Ying, et al.
Pubblicazione: (2024)
SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
di: Yu, Yaoning, et al.
Pubblicazione: (2025)
di: Yu, Yaoning, et al.
Pubblicazione: (2025)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
di: Nadas, Mihai, et al.
Pubblicazione: (2025)
di: Nadas, Mihai, et al.
Pubblicazione: (2025)
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
di: Huang, Yiming, et al.
Pubblicazione: (2024)
di: Huang, Yiming, et al.
Pubblicazione: (2024)
Bhaasha, Bhasa, Zaban: A Survey for Low-Resourced Languages in South Asia -- Current Stage and Challenges
di: Poria, Sampoorna, et al.
Pubblicazione: (2025)
di: Poria, Sampoorna, et al.
Pubblicazione: (2025)
Benchmarking Retrieval-Augmented Large Language Models in Biomedical NLP: Application, Robustness, and Self-Awareness
di: Li, Mingchen, et al.
Pubblicazione: (2024)
di: Li, Mingchen, et al.
Pubblicazione: (2024)
Synthetic Data Generation for Phrase Break Prediction with Large Language Model
di: Lee, Hoyeon, et al.
Pubblicazione: (2025)
di: Lee, Hoyeon, et al.
Pubblicazione: (2025)
Synthetic Data Generation in Low-Resource Settings via Fine-Tuning of Large Language Models
di: Kaddour, Jean, et al.
Pubblicazione: (2023)
di: Kaddour, Jean, et al.
Pubblicazione: (2023)
Data Kernel Perspective Space Performance Guarantees for Synthetic Data from Transformer Models
di: Browder, Michael, et al.
Pubblicazione: (2026)
di: Browder, Michael, et al.
Pubblicazione: (2026)
Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
di: Havrilla, Alex, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Time Matters: Examine Temporal Effects on Biomedical Language Models
di: Liu, Weisi, et al.
Pubblicazione: (2024) -
Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
di: Han, Guangzeng, et al.
Pubblicazione: (2025) -
Model-Agnostic Meta Learning for Class Imbalance Adaptation
di: Rao, Hanshu, et al.
Pubblicazione: (2026) -
Chain-of-Interaction: Enhancing Large Language Models for Psychiatric Behavior Understanding by Dyadic Contexts
di: Han, Guangzeng, et al.
Pubblicazione: (2024) -
Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation
di: Liu, Weisi, et al.
Pubblicazione: (2026)