Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thompson, Travis, Lim, Seung-Hwan, Liu, Paul, He, Ruoying, Xu, Dongkuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915358139482112
author Thompson, Travis
Lim, Seung-Hwan
Liu, Paul
He, Ruoying
Xu, Dongkuan
author_facet Thompson, Travis
Lim, Seung-Hwan
Liu, Paul
He, Ruoying
Xu, Dongkuan
contents Large Language Models (LLMs) have achieved impressive capabilities in language understanding and generation, yet they continue to underperform on knowledge-intensive reasoning tasks due to limited access to structured context and multi-hop information. Retrieval-Augmented Generation (RAG) partially mitigates this by grounding generation in retrieved context, but conventional RAG and GraphRAG methods often fail to capture relational structure across nodes in knowledge graphs. We introduce Inference-Scaled GraphRAG, a novel framework that enhances LLM-based graph reasoning by applying inference-time compute scaling. Our method combines sequential scaling with deep chain-of-thought graph traversal, and parallel scaling with majority voting over sampled trajectories within an interleaved reasoning-execution loop. Experiments on the GRBench benchmark demonstrate that our approach significantly improves multi-hop question answering performance, achieving substantial gains over both traditional GraphRAG and prior graph traversal baselines. These findings suggest that inference-time scaling is a practical and architecture-agnostic solution for structured knowledge reasoning with LLMs
format Preprint
id arxiv_https___arxiv_org_abs_2506_19967
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
Thompson, Travis
Lim, Seung-Hwan
Liu, Paul
He, Ruoying
Xu, Dongkuan
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have achieved impressive capabilities in language understanding and generation, yet they continue to underperform on knowledge-intensive reasoning tasks due to limited access to structured context and multi-hop information. Retrieval-Augmented Generation (RAG) partially mitigates this by grounding generation in retrieved context, but conventional RAG and GraphRAG methods often fail to capture relational structure across nodes in knowledge graphs. We introduce Inference-Scaled GraphRAG, a novel framework that enhances LLM-based graph reasoning by applying inference-time compute scaling. Our method combines sequential scaling with deep chain-of-thought graph traversal, and parallel scaling with majority voting over sampled trajectories within an interleaved reasoning-execution loop. Experiments on the GRBench benchmark demonstrate that our approach significantly improves multi-hop question answering performance, achieving substantial gains over both traditional GraphRAG and prior graph traversal baselines. These findings suggest that inference-time scaling is a practical and architecture-agnostic solution for structured knowledge reasoning with LLMs
title Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.19967