CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Yanlin, Papicchio, Simone, Rahman, Sajjadur
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910903956406272
author Feng, Yanlin
Papicchio, Simone
Rahman, Sajjadur
author_facet Feng, Yanlin
Papicchio, Simone
Rahman, Sajjadur
contents Retrieval from graph data is crucial for augmenting large language models (LLM) with both open-domain knowledge and private enterprise data, and it is also a key component in the recent GraphRAG system (edge et al., 2024). Despite decades of research on knowledge graphs and knowledge base question answering, leading LLM frameworks (e.g. Langchain and LlamaIndex) have only minimal support for retrieval from modern encyclopedic knowledge graphs like Wikidata. In this paper, we analyze the root cause and suggest that modern RDF knowledge graphs (e.g. Wikidata, Freebase) are less efficient for LLMs due to overly large schemas that far exceed the typical LLM context window, use of resource identifiers, overlapping relation types and lack of normalization. As a solution, we propose property graph views on top of the underlying RDF graph that can be efficiently queried by LLMs using Cypher. We instantiated this idea on Wikidata and introduced CypherBench, the first benchmark with 11 large-scale, multi-domain property graphs with 7.8 million entities and over 10,000 questions. To achieve this, we tackled several key challenges, including developing an RDF-to-property graph conversion engine, creating a systematic pipeline for text-to-Cypher task generation, and designing new evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18702
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
Feng, Yanlin
Papicchio, Simone
Rahman, Sajjadur
Computation and Language
Artificial Intelligence
Databases
Retrieval from graph data is crucial for augmenting large language models (LLM) with both open-domain knowledge and private enterprise data, and it is also a key component in the recent GraphRAG system (edge et al., 2024). Despite decades of research on knowledge graphs and knowledge base question answering, leading LLM frameworks (e.g. Langchain and LlamaIndex) have only minimal support for retrieval from modern encyclopedic knowledge graphs like Wikidata. In this paper, we analyze the root cause and suggest that modern RDF knowledge graphs (e.g. Wikidata, Freebase) are less efficient for LLMs due to overly large schemas that far exceed the typical LLM context window, use of resource identifiers, overlapping relation types and lack of normalization. As a solution, we propose property graph views on top of the underlying RDF graph that can be efficiently queried by LLMs using Cypher. We instantiated this idea on Wikidata and introduced CypherBench, the first benchmark with 11 large-scale, multi-domain property graphs with 7.8 million entities and over 10,000 questions. To achieve this, we tackled several key challenges, including developing an RDF-to-property graph conversion engine, creating a systematic pipeline for text-to-Cypher task generation, and designing new evaluation metrics.
title CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
topic Computation and Language
Artificial Intelligence
Databases
url https://arxiv.org/abs/2412.18702