Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
Fuente:
arXiv
Saved in:
| Main Authors: | Driouich, Ilias, Cao, Hongliu, Thomas, Eoin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025)
by: Cao, Hongliu, et al.
Published: (2025)
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
by: Cao, Hongliu, et al.
Published: (2026)
by: Cao, Hongliu, et al.
Published: (2026)
Semantic Adapter for Universal Text Embeddings: Diagnosing and Mitigating Negation Blindness to Enhance Universality
by: Cao, Hongliu
Published: (2025)
by: Cao, Hongliu
Published: (2025)
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
by: Cao, Hongliu
Published: (2024)
by: Cao, Hongliu
Published: (2024)
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
by: Cao, Hongliu, et al.
Published: (2026)
by: Cao, Hongliu, et al.
Published: (2026)
Measuring Diversity in Synthetic Datasets
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
Local Model Reconstruction Attacks in Federated Learning and their Uses
by: Driouich, Ilias, et al.
Published: (2022)
by: Driouich, Ilias, et al.
Published: (2022)
Controlled Generation for Private Synthetic Text
by: Zhao, Zihao, et al.
Published: (2025)
by: Zhao, Zihao, et al.
Published: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis
by: Wu, Yuqi, et al.
Published: (2025)
by: Wu, Yuqi, et al.
Published: (2025)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
A New Pipeline For Generating Instruction Dataset via RAG and Self Fine-Tuning
by: Song, Chih-Wei, et al.
Published: (2024)
by: Song, Chih-Wei, et al.
Published: (2024)
Synthetic Dialogue Dataset Generation using LLM Agents
by: Abdullin, Yelaman, et al.
Published: (2024)
by: Abdullin, Yelaman, et al.
Published: (2024)
Generating Synthetic Datasets for Few-shot Prompt Tuning
by: Guo, Xu, et al.
Published: (2024)
by: Guo, Xu, et al.
Published: (2024)
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
by: Papadimitriou, Ioannis, et al.
Published: (2024)
by: Papadimitriou, Ioannis, et al.
Published: (2024)
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
by: Zhao, Kaiwen, et al.
Published: (2025)
by: Zhao, Kaiwen, et al.
Published: (2025)
MARS: toward more efficient multi-agent collaboration for LLM reasoning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
by: Cao, Linxiao, et al.
Published: (2025)
by: Cao, Linxiao, et al.
Published: (2025)
GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluation
by: Lee, Jeongsoo, et al.
Published: (2025)
by: Lee, Jeongsoo, et al.
Published: (2025)
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
by: Zhang, Junhao, et al.
Published: (2025)
by: Zhang, Junhao, et al.
Published: (2025)
The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
Plancraft: an evaluation dataset for planning with LLM agents
by: Dagan, Gautier, et al.
Published: (2024)
by: Dagan, Gautier, et al.
Published: (2024)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
by: Lin, Zuhong, et al.
Published: (2025)
by: Lin, Zuhong, et al.
Published: (2025)
Using Large Language Models to Generate Authentic Multi-agent Knowledge Work Datasets
by: Heim, Desiree, et al.
Published: (2024)
by: Heim, Desiree, et al.
Published: (2024)
Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
by: Liu, Xukai, et al.
Published: (2025)
by: Liu, Xukai, et al.
Published: (2025)
Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
by: Kim, Hyunwoo, et al.
Published: (2026)
by: Kim, Hyunwoo, et al.
Published: (2026)
CircuitSynth: Reliable Synthetic Data Generation
by: Cheng, Zehua, et al.
Published: (2026)
by: Cheng, Zehua, et al.
Published: (2026)
A Typology of Synthetic Datasets for Dialogue Processing in Clinical Contexts
by: Bedrick, Steven, et al.
Published: (2025)
by: Bedrick, Steven, et al.
Published: (2025)
Parameterized Synthetic Text Generation with SimpleStories
by: Finke, Lennart, et al.
Published: (2025)
by: Finke, Lennart, et al.
Published: (2025)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
Scaling DPPs for RAG: Density Meets Diversity
by: Sun, Xun, et al.
Published: (2026)
by: Sun, Xun, et al.
Published: (2026)
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
A process algebraic framework for multi-agent dynamic epistemic systems
by: Aldini, Alessandro
Published: (2024)
by: Aldini, Alessandro
Published: (2024)
Linguistic and Argument Diversity in Synthetic Data for Function-Calling Agents
by: Greenstein, Dan, et al.
Published: (2026)
by: Greenstein, Dan, et al.
Published: (2026)
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
by: Ju, Li, et al.
Published: (2026)
by: Ju, Li, et al.
Published: (2026)
DebiasRAG: A Tuning-Free Path to Fair Generation in Large Language Models through Retrieval-Augmented Generation
by: Chu, Rui, et al.
Published: (2026)
by: Chu, Rui, et al.
Published: (2026)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
DuetRAG: Collaborative Retrieval-Augmented Generation
by: Jiao, Dian, et al.
Published: (2024)
by: Jiao, Dian, et al.
Published: (2024)
Similar Items
-
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025) -
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
by: Cao, Hongliu, et al.
Published: (2026) -
Semantic Adapter for Universal Text Embeddings: Diagnosing and Mitigating Negation Blindness to Enhance Universality
by: Cao, Hongliu
Published: (2025) -
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
by: Cao, Hongliu
Published: (2024) -
When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications
by: Cao, Hongliu, et al.
Published: (2026)