From Faithfulness to Correctness: Generative Reward Models that Think Critically

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Qiyao, Shi, Yunsheng, Tian, Hongtao, Wang, Chao, Chang, Weiming, Yao, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912615768260608
author Ma, Qiyao
Shi, Yunsheng
Tian, Hongtao
Wang, Chao
Chang, Weiming
Yao, Ting
author_facet Ma, Qiyao
Shi, Yunsheng
Tian, Hongtao
Wang, Chao
Chang, Weiming
Yao, Ting
contents Through reinforcement learning with verifiable rewards (RLVR), large language models have achieved substantial progress in domains with easily verifiable outcomes, such as mathematics and coding. However, when applied to more complex tasks like open-domain question answering, RLVR faces significant challenges due to the difficulty of verifying correctness. The nuanced and ambiguous nature of real-world knowledge makes it difficult to reliably evaluate correctness in these settings, necessitating further abilities that extend beyond mere logical consistency to encompass an understanding and assessment of both external and internal knowledge. Recent work has primarily focused on improving faithfulness, defined as semantic alignment with supporting documents, which can cause models to rely excessively on external sources and diminish their capacity for critical assessment. To address this, we propose the Thinking-supervised Reward Model (TRM), which incorporates sentence-level thinking supervision to endow reward models with critical thinking abilities. Given a query, answer, and supporting documents, TRM first assesses the faithfulness of each answer sentence to the supporting documents, and then applies a reasoning step to evaluate sentence-level correctness. By structuring reward modeling as a sequence of faithfulness, reasoning, and correctness evaluations, TRM encourages models to critically assess and leverage both external and internal knowledge. Experiments on reward signals demonstrate that TRM substantially improves the identification of incorrect sentences, and incorporating TRM into policy optimization leads to significant gains in both answer correctness and usefulness.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25409
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Faithfulness to Correctness: Generative Reward Models that Think Critically
Ma, Qiyao
Shi, Yunsheng
Tian, Hongtao
Wang, Chao
Chang, Weiming
Yao, Ting
Computation and Language
Artificial Intelligence
Machine Learning
Through reinforcement learning with verifiable rewards (RLVR), large language models have achieved substantial progress in domains with easily verifiable outcomes, such as mathematics and coding. However, when applied to more complex tasks like open-domain question answering, RLVR faces significant challenges due to the difficulty of verifying correctness. The nuanced and ambiguous nature of real-world knowledge makes it difficult to reliably evaluate correctness in these settings, necessitating further abilities that extend beyond mere logical consistency to encompass an understanding and assessment of both external and internal knowledge. Recent work has primarily focused on improving faithfulness, defined as semantic alignment with supporting documents, which can cause models to rely excessively on external sources and diminish their capacity for critical assessment. To address this, we propose the Thinking-supervised Reward Model (TRM), which incorporates sentence-level thinking supervision to endow reward models with critical thinking abilities. Given a query, answer, and supporting documents, TRM first assesses the faithfulness of each answer sentence to the supporting documents, and then applies a reasoning step to evaluate sentence-level correctness. By structuring reward modeling as a sequence of faithfulness, reasoning, and correctness evaluations, TRM encourages models to critically assess and leverage both external and internal knowledge. Experiments on reward signals demonstrate that TRM substantially improves the identification of incorrect sentences, and incorporating TRM into policy optimization leads to significant gains in both answer correctness and usefulness.
title From Faithfulness to Correctness: Generative Reward Models that Think Critically
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.25409