Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana
Fuente:
arXiv
Saved in:
| Main Authors: | Filice, Simone, Horowitz, Guy, Carmel, David, Karnin, Zohar, Lewin-Eytan, Liane, Maarek, Yoelle |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation
by: Carmel, David, et al.
Published: (2025)
by: Carmel, David, et al.
Published: (2025)
Do RAG Systems Really Suffer From Positional Bias?
by: Cuconasu, Florin, et al.
Published: (2025)
by: Cuconasu, Florin, et al.
Published: (2025)
SIGIR 2025 -- LiveRAG Challenge Report
by: Carmel, David, et al.
Published: (2025)
by: Carmel, David, et al.
Published: (2025)
The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora
by: Amiraz, Chen, et al.
Published: (2025)
by: Amiraz, Chen, et al.
Published: (2025)
The Distracting Effect: Understanding Irrelevant Passages in RAG
by: Amiraz, Chen, et al.
Published: (2025)
by: Amiraz, Chen, et al.
Published: (2025)
Redefining Retrieval Evaluation in the Era of LLMs
by: Trappolini, Giovanni, et al.
Published: (2025)
by: Trappolini, Giovanni, et al.
Published: (2025)
The Power of Noise: Redefining Retrieval for RAG Systems
by: Cuconasu, Florin, et al.
Published: (2024)
by: Cuconasu, Florin, et al.
Published: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
by: Zhu, Kunlun, et al.
Published: (2024)
by: Zhu, Kunlun, et al.
Published: (2024)
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
by: Li, Haitao, et al.
Published: (2025)
by: Li, Haitao, et al.
Published: (2025)
Context Embeddings for Efficient Answer Generation in RAG
by: Rau, David, et al.
Published: (2024)
by: Rau, David, et al.
Published: (2024)
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores
by: Yan, Guohang, et al.
Published: (2025)
by: Yan, Guohang, et al.
Published: (2025)
Rag Performance Prediction for Question Answering
by: Dado, Or, et al.
Published: (2026)
by: Dado, Or, et al.
Published: (2026)
FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing
by: Qian, Jingbin, et al.
Published: (2026)
by: Qian, Jingbin, et al.
Published: (2026)
Evaluating Hybrid Retrieval Augmented Generation using Dynamic Test Sets: LiveRAG Challenge
by: Fensore, Chase, et al.
Published: (2025)
by: Fensore, Chase, et al.
Published: (2025)
RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
by: Zhang, Rongyang, et al.
Published: (2025)
by: Zhang, Rongyang, et al.
Published: (2025)
Loops On Retrieval Augmented Generation (LoRAG)
by: Thakur, Ayush, et al.
Published: (2024)
by: Thakur, Ayush, et al.
Published: (2024)
SRAG: RAG with Structured Data Improves Vector Retrieval
by: Shah, Shalin, et al.
Published: (2026)
by: Shah, Shalin, et al.
Published: (2026)
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
by: Li, Bryan, et al.
Published: (2026)
by: Li, Bryan, et al.
Published: (2026)
Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework
by: Rackauckas, Zackary, et al.
Published: (2024)
by: Rackauckas, Zackary, et al.
Published: (2024)
RAG-based Question Answering over Heterogeneous Data and Text
by: Christmann, Philipp, et al.
Published: (2024)
by: Christmann, Philipp, et al.
Published: (2024)
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
by: Bachyr, Omar El, et al.
Published: (2026)
by: Bachyr, Omar El, et al.
Published: (2026)
CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation
by: Wang, Nengbo, et al.
Published: (2025)
by: Wang, Nengbo, et al.
Published: (2025)
MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation
by: Chang, Chia-Yuan, et al.
Published: (2024)
by: Chang, Chia-Yuan, et al.
Published: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Causal-Counterfactual RAG: The Integration of Causal-Counterfactual Reasoning into RAG
by: Khadilkar, Harshad, et al.
Published: (2025)
by: Khadilkar, Harshad, et al.
Published: (2025)
Who Stole Your Data? A Method for Detecting Unauthorized RAG Theft
by: Liu, Peiyang, et al.
Published: (2025)
by: Liu, Peiyang, et al.
Published: (2025)
FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
by: Zhang, Zhuocheng, et al.
Published: (2025)
by: Zhang, Zhuocheng, et al.
Published: (2025)
FIT-RAG: Black-Box RAG with Factual Information and Token Reduction
by: Mao, Yuren, et al.
Published: (2024)
by: Mao, Yuren, et al.
Published: (2024)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy
by: DeMarco, Michael R.
Published: (2026)
by: DeMarco, Michael R.
Published: (2026)
U-NIAH: Unified RAG and LLM Evaluation for Long Context Needle-In-A-Haystack
by: Gao, Yunfan, et al.
Published: (2025)
by: Gao, Yunfan, et al.
Published: (2025)
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2025)
by: Pradeep, Ronak, et al.
Published: (2025)
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
by: Chen, Hung-Ting, et al.
Published: (2024)
by: Chen, Hung-Ting, et al.
Published: (2024)
MedCoT-RAG: Causal Chain-of-Thought RAG for Medical Question Answering
by: Wang, Ziyu, et al.
Published: (2025)
by: Wang, Ziyu, et al.
Published: (2025)
OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning
by: Zhou, Jiawei, et al.
Published: (2025)
by: Zhou, Jiawei, et al.
Published: (2025)
Evidence Contextualization and Counterfactual Attribution for Conversational QA over Heterogeneous Data with RAG Systems
by: Roy, Rishiraj Saha, et al.
Published: (2024)
by: Roy, Rishiraj Saha, et al.
Published: (2024)
Similar Items
-
LiveRAG: A diverse Q&A dataset with varying difficulty level for RAG evaluation
by: Carmel, David, et al.
Published: (2025) -
Do RAG Systems Really Suffer From Positional Bias?
by: Cuconasu, Florin, et al.
Published: (2025) -
SIGIR 2025 -- LiveRAG Challenge Report
by: Carmel, David, et al.
Published: (2025) -
The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora
by: Amiraz, Chen, et al.
Published: (2025) -
The Distracting Effect: Understanding Irrelevant Passages in RAG
by: Amiraz, Chen, et al.
Published: (2025)