Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
Fuente:
arXiv
Salvato in:
| Autori principali: | Brehme, Lorenz, Ströhle, Thomas, Breu, Ruth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies
di: Brehme, Lorenz, et al.
Pubblicazione: (2026)
di: Brehme, Lorenz, et al.
Pubblicazione: (2026)
Retrieval-Augmented Generation in Industry: An Interview Study on Use Cases, Requirements, Challenges, and Evaluation
di: Brehme, Lorenz, et al.
Pubblicazione: (2025)
di: Brehme, Lorenz, et al.
Pubblicazione: (2025)
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
di: Brehme, Lorenz, et al.
Pubblicazione: (2026)
di: Brehme, Lorenz, et al.
Pubblicazione: (2026)
StratRAG: A Multi-Hop Retrieval Evaluation Dataset for Retrieval-Augmented Generation Systems
di: Patodiya, Aryan
Pubblicazione: (2026)
di: Patodiya, Aryan
Pubblicazione: (2026)
AMAQA: A Metadata-based QA Dataset for RAG Systems
di: Bruni, Davide, et al.
Pubblicazione: (2025)
di: Bruni, Davide, et al.
Pubblicazione: (2025)
TrustRAG: An Information Assistant with Retrieval Augmented Generation
di: Fan, Yixing, et al.
Pubblicazione: (2025)
di: Fan, Yixing, et al.
Pubblicazione: (2025)
Engineering the RAG Stack: A Comprehensive Review of the Architecture and Trust Frameworks for Retrieval-Augmented Generation Systems
di: Wampler, Dean, et al.
Pubblicazione: (2025)
di: Wampler, Dean, et al.
Pubblicazione: (2025)
Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
di: Dietz, Laura, et al.
Pubblicazione: (2026)
di: Dietz, Laura, et al.
Pubblicazione: (2026)
Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals
di: Zeng, Linda, et al.
Pubblicazione: (2025)
di: Zeng, Linda, et al.
Pubblicazione: (2025)
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation
di: Lee, Jeongsoo, et al.
Pubblicazione: (2025)
di: Lee, Jeongsoo, et al.
Pubblicazione: (2025)
Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges
di: Liang, Jintao, et al.
Pubblicazione: (2025)
di: Liang, Jintao, et al.
Pubblicazione: (2025)
MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
di: Ning, Yingpeng, et al.
Pubblicazione: (2025)
di: Ning, Yingpeng, et al.
Pubblicazione: (2025)
Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems
di: Fan, Dongzhe, et al.
Pubblicazione: (2026)
di: Fan, Dongzhe, et al.
Pubblicazione: (2026)
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
di: Fan, Wenqi, et al.
Pubblicazione: (2024)
di: Fan, Wenqi, et al.
Pubblicazione: (2024)
How Can Recommender Systems Benefit from Large Language Models: A Survey
di: Lin, Jianghao, et al.
Pubblicazione: (2023)
di: Lin, Jianghao, et al.
Pubblicazione: (2023)
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
di: Zhang, Dake, et al.
Pubblicazione: (2026)
di: Zhang, Dake, et al.
Pubblicazione: (2026)
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition
di: Cofala, Tim, et al.
Pubblicazione: (2025)
di: Cofala, Tim, et al.
Pubblicazione: (2025)
Evaluating Ensemble Methods for News Recommender Systems
di: Gray, Alexander, et al.
Pubblicazione: (2024)
di: Gray, Alexander, et al.
Pubblicazione: (2024)
RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines
di: Cohen, Dvir, et al.
Pubblicazione: (2025)
di: Cohen, Dvir, et al.
Pubblicazione: (2025)
Generating Diverse Synthetic Datasets for Evaluation of Real-life Recommender Systems
di: Malenšek, Miha, et al.
Pubblicazione: (2024)
di: Malenšek, Miha, et al.
Pubblicazione: (2024)
Political Events using RAG with LLMs
di: Arslan, Muhammad, et al.
Pubblicazione: (2025)
di: Arslan, Muhammad, et al.
Pubblicazione: (2025)
Optimizing and Evaluating Enterprise Retrieval-Augmented Generation (RAG): A Content Design Perspective
di: Packowski, Sarah, et al.
Pubblicazione: (2024)
di: Packowski, Sarah, et al.
Pubblicazione: (2024)
DMQR-RAG: Diverse Multi-Query Rewriting for RAG
di: Li, Zhicong, et al.
Pubblicazione: (2024)
di: Li, Zhicong, et al.
Pubblicazione: (2024)
GraphRAG-Router: Learning Cost-Efficient Routing over GraphRAGs and LLMs with Reinforcement Learning
di: Fan, Dongzhe, et al.
Pubblicazione: (2026)
di: Fan, Dongzhe, et al.
Pubblicazione: (2026)
OmniBench-RAG: A Multi-Domain Evaluation Platform for Retrieval-Augmented Generation Tools
di: Liang, Jiaxuan, et al.
Pubblicazione: (2025)
di: Liang, Jiaxuan, et al.
Pubblicazione: (2025)
HugRAG: Hierarchical Causal Knowledge Graph Design for RAG
di: Wang, Nengbo, et al.
Pubblicazione: (2026)
di: Wang, Nengbo, et al.
Pubblicazione: (2026)
M-RAG: Making RAG Faster, Stronger, and More Efficient
di: Xu, Sun, et al.
Pubblicazione: (2026)
di: Xu, Sun, et al.
Pubblicazione: (2026)
Fashion-AlterEval: A Dataset for Improved Evaluation of Conversational Recommendation Systems with Alternative Relevant Items
di: Vlachou, Maria
Pubblicazione: (2025)
di: Vlachou, Maria
Pubblicazione: (2025)
HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers
di: Santra, Payel, et al.
Pubblicazione: (2025)
di: Santra, Payel, et al.
Pubblicazione: (2025)
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
di: Singh, Aditi, et al.
Pubblicazione: (2025)
di: Singh, Aditi, et al.
Pubblicazione: (2025)
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test Data
di: Yang, Yahe, et al.
Pubblicazione: (2025)
di: Yang, Yahe, et al.
Pubblicazione: (2025)
LARAG: Link-Aware Retrieval Strategy for RAG Systems in Hyperlinked Technical Documentation
di: Bolognesi, Giorgia, et al.
Pubblicazione: (2026)
di: Bolognesi, Giorgia, et al.
Pubblicazione: (2026)
Embedding in Recommender Systems: A Survey
di: Wang, Maolin, et al.
Pubblicazione: (2023)
di: Wang, Maolin, et al.
Pubblicazione: (2023)
Multimodal Recommender Systems: A Survey
di: Liu, Qidong, et al.
Pubblicazione: (2023)
di: Liu, Qidong, et al.
Pubblicazione: (2023)
A Survey of Reasoning for Substitution Relationships: Definitions, Methods, and Directions
di: Yang, Anxin, et al.
Pubblicazione: (2024)
di: Yang, Anxin, et al.
Pubblicazione: (2024)
RAG over Thinking Traces Can Improve Reasoning Tasks
di: Arabzadeh, Negar, et al.
Pubblicazione: (2026)
di: Arabzadeh, Negar, et al.
Pubblicazione: (2026)
AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
di: Ma, Bo, et al.
Pubblicazione: (2025)
di: Ma, Bo, et al.
Pubblicazione: (2025)
EcphoryRAG: Re-Imagining Knowledge-Graph RAG via Human Associative Memory
di: Liao, Zirui
Pubblicazione: (2025)
di: Liao, Zirui
Pubblicazione: (2025)
Multi-Behavior Recommender Systems: A Survey
di: Kim, Kyungho, et al.
Pubblicazione: (2025)
di: Kim, Kyungho, et al.
Pubblicazione: (2025)
A Survey on Diffusion Models for Recommender Systems
di: Lin, Jianghao, et al.
Pubblicazione: (2024)
di: Lin, Jianghao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating Multi-Hop Reasoning in RAG Systems: A Comparison of LLM-Based Retriever Evaluation Strategies
di: Brehme, Lorenz, et al.
Pubblicazione: (2026) -
Retrieval-Augmented Generation in Industry: An Interview Study on Use Cases, Requirements, Challenges, and Evaluation
di: Brehme, Lorenz, et al.
Pubblicazione: (2025) -
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
di: Brehme, Lorenz, et al.
Pubblicazione: (2026) -
StratRAG: A Multi-Hop Retrieval Evaluation Dataset for Retrieval-Augmented Generation Systems
di: Patodiya, Aryan
Pubblicazione: (2026) -
AMAQA: A Metadata-based QA Dataset for RAG Systems
di: Bruni, Davide, et al.
Pubblicazione: (2025)