Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cattaneo, Alberto, Luschi, Carlo, Justus, Daniel
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918230723919872
author Cattaneo, Alberto
Luschi, Carlo
Justus, Daniel
author_facet Cattaneo, Alberto
Luschi, Carlo
Justus, Daniel
contents Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs. While various solutions have been proposed, a comparison of methods is difficult due to the lack of challenging QA datasets with ground-truth targets for graph retrieval. We present SynthKGQA, an LLM-powered framework for generating high-quality Knowledge Graph Question Answering datasets from any Knowledge Graph, providing the full set of ground-truth facts in the KG to reason over questions. We show how, in addition to enabling more informative benchmarking of KG retrievers, the data produced with SynthKGQA also allows us to train better models.We apply SynthKGQA to Wikidata to generate GTSQA, a new dataset designed to test zero-shot generalization abilities of KG retrievers with respect to unseen graph structures and relation types, and benchmark popular solutions for KG-augmented LLMs on it.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
Cattaneo, Alberto
Luschi, Carlo
Justus, Daniel
Machine Learning
Artificial Intelligence
Computation and Language
Information Retrieval
Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs. While various solutions have been proposed, a comparison of methods is difficult due to the lack of challenging QA datasets with ground-truth targets for graph retrieval. We present SynthKGQA, an LLM-powered framework for generating high-quality Knowledge Graph Question Answering datasets from any Knowledge Graph, providing the full set of ground-truth facts in the KG to reason over questions. We show how, in addition to enabling more informative benchmarking of KG retrievers, the data produced with SynthKGQA also allows us to train better models.We apply SynthKGQA to Wikidata to generate GTSQA, a new dataset designed to test zero-shot generalization abilities of KG retrievers with respect to unseen graph structures and relation types, and benchmark popular solutions for KG-augmented LLMs on it.
title Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2511.04473