Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Shiqi, Chen, Yubo, Zhou, Ruiqi, Yao, Zhengxi, Chen, Shuai, Zhang, Tianyi, Zhang, Shijie, Zhang, Wei Qiang, Huang, Yongfeng, Duan, Haixin, Zhang, Yunqi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911467468488704
author Yan, Shiqi
Chen, Yubo
Zhou, Ruiqi
Yao, Zhengxi
Chen, Shuai
Zhang, Tianyi
Zhang, Shijie
Zhang, Wei Qiang
Huang, Yongfeng
Duan, Haixin
Zhang, Yunqi
author_facet Yan, Shiqi
Chen, Yubo
Zhou, Ruiqi
Yao, Zhengxi
Chen, Shuai
Zhang, Tianyi
Zhang, Shijie
Zhang, Wei Qiang
Huang, Yongfeng
Duan, Haixin
Zhang, Yunqi
contents The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrained LLM reasoning either by enforcing rules during generation or by imitating paths from a fixed set of demonstrations. However, they naturally confined the reasoning patterns of LLMs within the scope of prior experience or fine-tuning data, limiting their generalizability to out-of-distribution graph reasoning problems. To tackle this problem, in this paper, we propose Explore-on-Graph (EoG), a novel framework that encourages LLMs to autonomously explore a more diverse reasoning space on KGs. To incentivize exploration and discovery of novel reasoning paths, we propose to introduce reinforcement learning during training, whose reward is the correctness of the reasoning paths' final answers. To enhance the efficiency and meaningfulness of the exploration, we propose to incorporate path information as additional reward signals to refine the exploration process and reduce futile efforts. Extensive experiments on five KGQA benchmark datasets demonstrate that, to the best of our knowledge, our method achieves state-of-the-art performance, outperforming not only open-source but also even closed-source LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21728
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling
Yan, Shiqi
Chen, Yubo
Zhou, Ruiqi
Yao, Zhengxi
Chen, Shuai
Zhang, Tianyi
Zhang, Shijie
Zhang, Wei Qiang
Huang, Yongfeng
Duan, Haixin
Zhang, Yunqi
Computation and Language
The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrained LLM reasoning either by enforcing rules during generation or by imitating paths from a fixed set of demonstrations. However, they naturally confined the reasoning patterns of LLMs within the scope of prior experience or fine-tuning data, limiting their generalizability to out-of-distribution graph reasoning problems. To tackle this problem, in this paper, we propose Explore-on-Graph (EoG), a novel framework that encourages LLMs to autonomously explore a more diverse reasoning space on KGs. To incentivize exploration and discovery of novel reasoning paths, we propose to introduce reinforcement learning during training, whose reward is the correctness of the reasoning paths' final answers. To enhance the efficiency and meaningfulness of the exploration, we propose to incorporate path information as additional reward signals to refine the exploration process and reduce futile efforts. Extensive experiments on five KGQA benchmark datasets demonstrate that, to the best of our knowledge, our method achieves state-of-the-art performance, outperforming not only open-source but also even closed-source LLMs.
title Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling
topic Computation and Language
url https://arxiv.org/abs/2602.21728