Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tamber, Manveer Singh, Bao, Forrest Sheng, Xu, Chenyu, Luo, Ge, Kazi, Suleman, Bae, Minseok, Li, Miaoran, Mendelevitch, Ofer, Qu, Renyi, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024)
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024)
Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Unifying Adversarial Robustness and Training Across Text Scoring Models
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2026)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2026)
Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2024)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2024)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
Is Semantic Chunking Worth the Computational Cost?
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
The Trust Paradox: How CS Researchers Engage LLM Leaderboards
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2026)
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2026)
LEGOBench: Scientific Leaderboard Generation Benchmark
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
LegalFaith: A Faithfulness Benchmark for Legal QA
von: Singh, Rajshree
Veröffentlicht: (2026)
von: Singh, Rajshree
Veröffentlicht: (2026)
A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation
von: Oyarhoseini, Hosna, et al.
Veröffentlicht: (2026)
von: Oyarhoseini, Hosna, et al.
Veröffentlicht: (2026)
Improving LLM Leaderboards with Psychometrical Methodology
von: Federiakin, Denis
Veröffentlicht: (2025)
von: Federiakin, Denis
Veröffentlicht: (2025)
Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG
von: Triantafyllopoulos, Ilias, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Ilias, et al.
Veröffentlicht: (2025)
The Leaderboard Illusion
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
Correctness is not Faithfulness in RAG Attributions
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
von: Uemura, Kosei, et al.
Veröffentlicht: (2025)
von: Uemura, Kosei, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
From RAG to Agentic RAG for Faithful Islamic Question Answering
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
von: Mazaheri, Parsa, et al.
Veröffentlicht: (2026)
von: Mazaheri, Parsa, et al.
Veröffentlicht: (2026)
LLM Robustness Leaderboard v1 --Technical report
von: Lefebvre, Pierre Peigné -, et al.
Veröffentlicht: (2025)
von: Lefebvre, Pierre Peigné -, et al.
Veröffentlicht: (2025)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
Rembrandt – A képek színjátéka. Hermeneutikai kísérletek (online képmelléklet)
von: Rényi, András
Veröffentlicht: (2025)
von: Rényi, András
Veröffentlicht: (2025)
GUISE: Graph GaUssIan Shading watErmark
von: Yang, Renyi
Veröffentlicht: (2024)
von: Yang, Renyi
Veröffentlicht: (2024)
Rembrandt – A képek színjátéka. Hermeneutikai kísérletek
von: András, Rényi
Veröffentlicht: (2025)
von: András, Rényi
Veröffentlicht: (2025)
Pluralistic Leaderboards
von: Haghtalab, Nika, et al.
Veröffentlicht: (2026)
von: Haghtalab, Nika, et al.
Veröffentlicht: (2026)
Prompt-to-Leaderboard
von: Frick, Evan, et al.
Veröffentlicht: (2025)
von: Frick, Evan, et al.
Veröffentlicht: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
Pervasive Annotation Errors Break Text-to-SQL Benchmarks and Leaderboards
von: Jin, Tengjun, et al.
Veröffentlicht: (2026)
von: Jin, Tengjun, et al.
Veröffentlicht: (2026)
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
von: Chen, Wenting, et al.
Veröffentlicht: (2025)
Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards
von: Arzt, Varvara, et al.
Veröffentlicht: (2024)
von: Arzt, Varvara, et al.
Veröffentlicht: (2024)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
von: Cheng, Aileen, et al.
Veröffentlicht: (2025)
von: Cheng, Aileen, et al.
Veröffentlicht: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
Open FinLLM Leaderboard: Towards Financial AI Readiness
von: Lin, Shengyuan Colin, et al.
Veröffentlicht: (2025)
von: Lin, Shengyuan Colin, et al.
Veröffentlicht: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
von: Kim, Minseok, et al.
Veröffentlicht: (2025)
von: Kim, Minseok, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
von: Bao, Forrest Sheng, et al.
Veröffentlicht: (2024) -
Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025) -
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025) -
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025) -
Unifying Adversarial Robustness and Training Across Text Scoring Models
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2026)