Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Zhongxiang, Wang, Qipeng, Wang, Haoyu, Zhang, Xiao, Xu, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910953132523520
author Sun, Zhongxiang
Wang, Qipeng
Wang, Haoyu
Zhang, Xiao
Xu, Jun
author_facet Sun, Zhongxiang
Wang, Qipeng
Wang, Haoyu
Zhang, Xiao
Xu, Jun
contents Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks. However, alongside these successes, a more deceptive form of model error has emerged--Reasoning Hallucination--where logically coherent but factually incorrect reasoning traces lead to persuasive yet faulty conclusions. Unlike traditional hallucinations, these errors are embedded within structured reasoning, making them more difficult to detect and potentially more harmful. In this work, we investigate reasoning hallucinations from a mechanistic perspective. We propose the Reasoning Score, which quantifies the depth of reasoning by measuring the divergence between logits obtained from projecting late layers of LRMs to the vocabulary space, effectively distinguishing shallow pattern-matching from genuine deep reasoning. Using this score, we conduct an in-depth analysis on the ReTruthQA dataset and identify two key reasoning hallucination patterns: early-stage fluctuation in reasoning depth and incorrect backtracking to flawed prior steps. These insights motivate our Reasoning Hallucination Detection (RHD) framework, which achieves state-of-the-art performance across multiple domains. To mitigate reasoning hallucinations, we further introduce GRPO-R, an enhanced reinforcement learning algorithm that incorporates step-level deep reasoning rewards via potential-based shaping. Our theoretical analysis establishes stronger generalization guarantees, and experiments demonstrate improved reasoning quality and reduced hallucination rates.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
Sun, Zhongxiang
Wang, Qipeng
Wang, Haoyu
Zhang, Xiao
Xu, Jun
Artificial Intelligence
Computation and Language
Computers and Society
Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks. However, alongside these successes, a more deceptive form of model error has emerged--Reasoning Hallucination--where logically coherent but factually incorrect reasoning traces lead to persuasive yet faulty conclusions. Unlike traditional hallucinations, these errors are embedded within structured reasoning, making them more difficult to detect and potentially more harmful. In this work, we investigate reasoning hallucinations from a mechanistic perspective. We propose the Reasoning Score, which quantifies the depth of reasoning by measuring the divergence between logits obtained from projecting late layers of LRMs to the vocabulary space, effectively distinguishing shallow pattern-matching from genuine deep reasoning. Using this score, we conduct an in-depth analysis on the ReTruthQA dataset and identify two key reasoning hallucination patterns: early-stage fluctuation in reasoning depth and incorrect backtracking to flawed prior steps. These insights motivate our Reasoning Hallucination Detection (RHD) framework, which achieves state-of-the-art performance across multiple domains. To mitigate reasoning hallucinations, we further introduce GRPO-R, an enhanced reinforcement learning algorithm that incorporates step-level deep reasoning rewards via potential-based shaping. Our theoretical analysis establishes stronger generalization guarantees, and experiments demonstrate improved reasoning quality and reduced hallucination rates.
title Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
topic Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2505.12886