On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Long, Lin, Wang, Rui, Xiao, Ruixuan, Zhao, Junbo, Ding, Xiao, Chen, Gang, Wang, Haobo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024)
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024)
D.Va: Validate Your Demonstration First Before You Use It
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
RECOST: External Knowledge Guided Data-efficient Instruction Tuning
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
von: Zhong, Hanmeng, et al.
Veröffentlicht: (2025)
von: Zhong, Hanmeng, et al.
Veröffentlicht: (2025)
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
von: Wu, Pengzuo, et al.
Veröffentlicht: (2025)
von: Wu, Pengzuo, et al.
Veröffentlicht: (2025)
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
von: Xia, Mingxuan, et al.
Veröffentlicht: (2025)
von: Xia, Mingxuan, et al.
Veröffentlicht: (2025)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Data Swarms: Optimizable Generation of Synthetic Evaluation Data
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
von: Lupidi, Alisia, et al.
Veröffentlicht: (2024)
von: Lupidi, Alisia, et al.
Veröffentlicht: (2024)
Reasoning-Driven Synthetic Data Generation and Evaluation
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
Energy-based Automated Model Evaluation
von: Peng, Ru, et al.
Veröffentlicht: (2024)
von: Peng, Ru, et al.
Veröffentlicht: (2024)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
von: Fan, Zhiting, et al.
Veröffentlicht: (2026)
von: Fan, Zhiting, et al.
Veröffentlicht: (2026)
Evaluating Language Models as Synthetic Data Generators
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
SynCPKL: Harnessing LLMs to Generate Synthetic Data for Commonsense Persona Knowledge Linking
von: Lin, Kuan-Yen
Veröffentlicht: (2024)
von: Lin, Kuan-Yen
Veröffentlicht: (2024)
Improving Data Efficiency via Curating LLM-Driven Rating Systems
von: Pang, Jinlong, et al.
Veröffentlicht: (2024)
von: Pang, Jinlong, et al.
Veröffentlicht: (2024)
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
von: Chen, Zihong, et al.
Veröffentlicht: (2025)
von: Chen, Zihong, et al.
Veröffentlicht: (2025)
Dynamic Evaluation for Oversensitivity in LLMs
von: Pu, Sophia Xiao, et al.
Veröffentlicht: (2025)
von: Pu, Sophia Xiao, et al.
Veröffentlicht: (2025)
SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation
von: Xiao, Zhiqing, et al.
Veröffentlicht: (2025)
von: Xiao, Zhiqing, et al.
Veröffentlicht: (2025)
Generating Logically Consistent Synthetic Supply Chain Data with LLM-Driven Knowledge Graph Reasoning
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework
von: Yang, Bohao, et al.
Veröffentlicht: (2024)
von: Yang, Bohao, et al.
Veröffentlicht: (2024)
ELTEX: A Framework for Domain-Driven Synthetic Data Generation
von: Razmyslovich, Arina, et al.
Veröffentlicht: (2025)
von: Razmyslovich, Arina, et al.
Veröffentlicht: (2025)
Curating Grounded Synthetic Data with Global Perspectives for Equitable AI
von: Törnquist, Elin, et al.
Veröffentlicht: (2024)
von: Törnquist, Elin, et al.
Veröffentlicht: (2024)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
von: Son, Guijin, et al.
Veröffentlicht: (2026)
von: Son, Guijin, et al.
Veröffentlicht: (2026)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
von: Gao, Lirong, et al.
Veröffentlicht: (2026)
von: Gao, Lirong, et al.
Veröffentlicht: (2026)
DeltaMem: Towards Agentic Memory Management via Reinforcement Learning
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
von: Chen, Zhou, et al.
Veröffentlicht: (2025)
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
Embedding-Driven Diversity Sampling to Improve Few-Shot Synthetic Data Generation
von: Lopez, Ivan, et al.
Veröffentlicht: (2025)
von: Lopez, Ivan, et al.
Veröffentlicht: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026)
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
von: Hui, Zheng, et al.
Veröffentlicht: (2024)
SurveyAgent: A Conversational System for Personalized and Efficient Research Survey
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation
von: Tao, Chunliang, et al.
Veröffentlicht: (2024)
von: Tao, Chunliang, et al.
Veröffentlicht: (2024)
A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
von: Xiao, Quan, et al.
Veröffentlicht: (2025)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
von: Yan, Xiangchao, et al.
Veröffentlicht: (2025)
von: Yan, Xiangchao, et al.
Veröffentlicht: (2025)
RPDR: A Round-trip Prediction-Based Data Augmentation Framework for Long-Tail Question Answering
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
von: de Gibert, Ona, et al.
Veröffentlicht: (2025)
von: de Gibert, Ona, et al.
Veröffentlicht: (2025)
DRTriton: Large-Scale Synthetic Data Driven Reinforcement Learning for Triton Kernel Generation
von: Guo, Siqi, et al.
Veröffentlicht: (2026)
von: Guo, Siqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
von: Xiao, Ruixuan, et al.
Veröffentlicht: (2024) -
D.Va: Validate Your Demonstration First Before You Use It
von: Zhang, Qi, et al.
Veröffentlicht: (2025) -
RECOST: External Knowledge Guided Data-efficient Instruction Tuning
von: Zhang, Qi, et al.
Veröffentlicht: (2024) -
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
von: Zhong, Hanmeng, et al.
Veröffentlicht: (2025) -
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
von: Wu, Pengzuo, et al.
Veröffentlicht: (2025)