What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Dongzhuoran, Zhu, Yuqicheng, Wang, Xiaxia, Zhou, Hongkuan, He, Yuan, Chen, Jiaoyan, Staab, Steffen, Kharlamov, Evgeny
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917195480563712
author Zhou, Dongzhuoran
Zhu, Yuqicheng
Wang, Xiaxia
Zhou, Hongkuan
He, Yuan
Chen, Jiaoyan
Staab, Steffen
Kharlamov, Evgeny
author_facet Zhou, Dongzhuoran
Zhu, Yuqicheng
Wang, Xiaxia
Zhou, Hongkuan
He, Yuan
Chen, Jiaoyan
Staab, Steffen
Kharlamov, Evgeny
contents Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) is an increasingly explored approach for combining the reasoning capabilities of large language models with the structured evidence of knowledge graphs. However, current evaluation practices fall short: existing benchmarks often include questions that can be directly answered using existing triples in KG, making it unclear whether models perform reasoning or simply retrieve answers directly. Moreover, inconsistent evaluation metrics and lenient answer matching criteria further obscure meaningful comparisons. In this work, we introduce a general method for constructing benchmarks and present BRINK (Benchmark for Reasoning under Incomplete Knowledge) to systematically assess KG-RAG methods under knowledge incompleteness. Our empirical results show that current KG-RAG methods have limited reasoning ability under missing knowledge, often rely on internal memorization, and exhibit varying degrees of generalization depending on their design.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge
Zhou, Dongzhuoran
Zhu, Yuqicheng
Wang, Xiaxia
Zhou, Hongkuan
He, Yuan
Chen, Jiaoyan
Staab, Steffen
Kharlamov, Evgeny
Artificial Intelligence
Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) is an increasingly explored approach for combining the reasoning capabilities of large language models with the structured evidence of knowledge graphs. However, current evaluation practices fall short: existing benchmarks often include questions that can be directly answered using existing triples in KG, making it unclear whether models perform reasoning or simply retrieve answers directly. Moreover, inconsistent evaluation metrics and lenient answer matching criteria further obscure meaningful comparisons. In this work, we introduce a general method for constructing benchmarks and present BRINK (Benchmark for Reasoning under Incomplete Knowledge) to systematically assess KG-RAG methods under knowledge incompleteness. Our empirical results show that current KG-RAG methods have limited reasoning ability under missing knowledge, often rely on internal memorization, and exhibit varying degrees of generalization depending on their design.
title What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge
topic Artificial Intelligence
url https://arxiv.org/abs/2508.08344