FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gui, Runquan, Li, Yafu, Qu, Xiaoye, Liu, Ziyan, Cheng, Yeqiu, Cheng, Yu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908828540338176
author Gui, Runquan
Li, Yafu
Qu, Xiaoye
Liu, Ziyan
Cheng, Yeqiu
Cheng, Yu
author_facet Gui, Runquan
Li, Yafu
Qu, Xiaoye
Liu, Ziyan
Cheng, Yeqiu
Cheng, Yu
contents Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most RLVR pipelines rely on sparse outcome-based rewards, providing little supervision over intermediate steps and thus encouraging over-confidence and spurious reasoning, which in turn increases hallucinations. To address this, we propose FaithRL, a general reinforcement learning framework that directly optimizes reasoning faithfulness. We formalize a faithfulness-maximization objective and theoretically show that optimizing it mitigates over-confidence. To instantiate this objective, we introduce a geometric reward design and a faithfulness-aware advantage modulation mechanism that assigns step-level credit by penalizing unsupported steps while preserving valid partial derivations. Across diverse backbones and benchmarks, FaithRL consistently reduces hallucination rates while maintaining (and often improving) answer correctness. Further analysis confirms that FaithRL increases step-wise reasoning faithfulness and generalizes robustly. Our code is available at https://github.com/aintdoin/FaithRL.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03507
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
Gui, Runquan
Li, Yafu
Qu, Xiaoye
Liu, Ziyan
Cheng, Yeqiu
Cheng, Yu
Computation and Language
Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most RLVR pipelines rely on sparse outcome-based rewards, providing little supervision over intermediate steps and thus encouraging over-confidence and spurious reasoning, which in turn increases hallucinations. To address this, we propose FaithRL, a general reinforcement learning framework that directly optimizes reasoning faithfulness. We formalize a faithfulness-maximization objective and theoretically show that optimizing it mitigates over-confidence. To instantiate this objective, we introduce a geometric reward design and a faithfulness-aware advantage modulation mechanism that assigns step-level credit by penalizing unsupported steps while preserving valid partial derivations. Across diverse backbones and benchmarks, FaithRL consistently reduces hallucination rates while maintaining (and often improving) answer correctness. Further analysis confirms that FaithRL increases step-wise reasoning faithfulness and generalizes robustly. Our code is available at https://github.com/aintdoin/FaithRL.
title FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
topic Computation and Language
url https://arxiv.org/abs/2602.03507