Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bu, Yunhan, Zhang, Quan, Zhang, Huaping, Geng, Guotong, Gao, Chunxiao, Hamdulla, Askar, Wang, Juan, Li, Qiuchi, Zhang, Baohua, Lei, Shuai, Cao, Yunbo, Luo, Zhunchen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917473001930752
author Bu, Yunhan
Zhang, Quan
Zhang, Huaping
Geng, Guotong
Gao, Chunxiao
Hamdulla, Askar
Wang, Juan
Li, Qiuchi
Zhang, Baohua
Lei, Shuai
Cao, Yunbo
Luo, Zhunchen
author_facet Bu, Yunhan
Zhang, Quan
Zhang, Huaping
Geng, Guotong
Gao, Chunxiao
Hamdulla, Askar
Wang, Juan
Li, Qiuchi
Zhang, Baohua
Lei, Shuai
Cao, Yunbo
Luo, Zhunchen
contents Multi-Hop Fact Verification (MHFV) necessitates complex reasoning across disparate evidence, posing significant challenges for Large Language Models (LLMs) which often suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought (CoT), lack explicit modeling of the causal dependencies between evidence and claims. In this work, we introduce a novel framework that grounds reasoning in a Structural Causal Model (SCM), treating verification as a constructive causal inference process. We empirically identify an "inverted U-shaped" correlation between reasoning chain length and accuracy, revealing that excessive structural complexity degrades performance. To address this, we propose a Rule-based Reinforcement Learning strategy using Group Relative Policy Optimization (GRPO). This approach dynamically optimizes the trade-off between structural depth and conciseness. Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework significantly outperforms state-of-the-art baselines, offering a reliable and interpretable solution for complex fact verification.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01482
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization
Bu, Yunhan
Zhang, Quan
Zhang, Huaping
Geng, Guotong
Gao, Chunxiao
Hamdulla, Askar
Wang, Juan
Li, Qiuchi
Zhang, Baohua
Lei, Shuai
Cao, Yunbo
Luo, Zhunchen
Artificial Intelligence
Multi-Hop Fact Verification (MHFV) necessitates complex reasoning across disparate evidence, posing significant challenges for Large Language Models (LLMs) which often suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought (CoT), lack explicit modeling of the causal dependencies between evidence and claims. In this work, we introduce a novel framework that grounds reasoning in a Structural Causal Model (SCM), treating verification as a constructive causal inference process. We empirically identify an "inverted U-shaped" correlation between reasoning chain length and accuracy, revealing that excessive structural complexity degrades performance. To address this, we propose a Rule-based Reinforcement Learning strategy using Group Relative Policy Optimization (GRPO). This approach dynamically optimizes the trade-off between structural depth and conciseness. Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework significantly outperforms state-of-the-art baselines, offering a reliable and interpretable solution for complex fact verification.
title Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization
topic Artificial Intelligence
url https://arxiv.org/abs/2605.01482