Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Liangliang, Jiang, Zhuorui, Chi, Hongliang, Chen, Haoyang, Elkoumy, Mohammed, Wang, Fali, Wu, Qiong, Zhou, Zhengyi, Pan, Shirui, Wang, Suhang, Ma, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911247766650880
author Zhang, Liangliang
Jiang, Zhuorui
Chi, Hongliang
Chen, Haoyang
Elkoumy, Mohammed
Wang, Fali
Wu, Qiong
Zhou, Zhengyi
Pan, Shirui
Wang, Suhang
Ma, Yao
author_facet Zhang, Liangliang
Jiang, Zhuorui
Chi, Hongliang
Chen, Haoyang
Elkoumy, Mohammed
Wang, Fali
Wu, Qiong
Zhou, Zhengyi
Pan, Shirui
Wang, Suhang
Ma, Yao
contents Knowledge Graph Question Answering (KGQA) systems rely on high-quality benchmarks to evaluate complex multi-hop reasoning. However, despite their widespread use, popular datasets such as WebQSP and CWQ suffer from critical quality issues, including inaccurate or incomplete ground-truth annotations, poorly constructed questions that are ambiguous, trivial, or unanswerable, and outdated or inconsistent knowledge. Through a manual audit of 16 popular KGQA datasets, including WebQSP and CWQ, we find that the average factual correctness rate is only 57 %. To address these issues, we introduce KGQAGen, an LLM-in-the-loop framework that systematically resolves these pitfalls. KGQAGen combines structured knowledge grounding, LLM-guided generation, and symbolic verification to produce challenging and verifiable QA instances. Using KGQAGen, we construct KGQAGen-10k, a ten-thousand scale benchmark grounded in Wikidata, and evaluate a diverse set of KG-RAG models. Experimental results demonstrate that even state-of-the-art systems struggle on this benchmark, highlighting its ability to expose limitations of existing models. Our findings advocate for more rigorous benchmark construction and position KGQAGen as a scalable framework for advancing KGQA evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
Zhang, Liangliang
Jiang, Zhuorui
Chi, Hongliang
Chen, Haoyang
Elkoumy, Mohammed
Wang, Fali
Wu, Qiong
Zhou, Zhengyi
Pan, Shirui
Wang, Suhang
Ma, Yao
Computation and Language
Artificial Intelligence
Machine Learning
I.2.6; I.2.7
Knowledge Graph Question Answering (KGQA) systems rely on high-quality benchmarks to evaluate complex multi-hop reasoning. However, despite their widespread use, popular datasets such as WebQSP and CWQ suffer from critical quality issues, including inaccurate or incomplete ground-truth annotations, poorly constructed questions that are ambiguous, trivial, or unanswerable, and outdated or inconsistent knowledge. Through a manual audit of 16 popular KGQA datasets, including WebQSP and CWQ, we find that the average factual correctness rate is only 57 %. To address these issues, we introduce KGQAGen, an LLM-in-the-loop framework that systematically resolves these pitfalls. KGQAGen combines structured knowledge grounding, LLM-guided generation, and symbolic verification to produce challenging and verifiable QA instances. Using KGQAGen, we construct KGQAGen-10k, a ten-thousand scale benchmark grounded in Wikidata, and evaluate a diverse set of KG-RAG models. Experimental results demonstrate that even state-of-the-art systems struggle on this benchmark, highlighting its ability to expose limitations of existing models. Our findings advocate for more rigorous benchmark construction and position KGQAGen as a scalable framework for advancing KGQA evaluation.
title Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.6; I.2.7
url https://arxiv.org/abs/2505.23495