Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Qianxi, Ren, Qingyu, Lei, Shanzhe, Wang, Xuhong, Wang, Yingchun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911258529234944
author He, Qianxi
Ren, Qingyu
Lei, Shanzhe
Wang, Xuhong
Wang, Yingchun
author_facet He, Qianxi
Ren, Qingyu
Lei, Shanzhe
Wang, Xuhong
Wang, Yingchun
contents Recent advancements in large language models (LLMs) have shifted the post-training paradigm from traditional instruction tuning and human preference alignment toward reinforcement learning (RL) focused on reasoning capabilities. However, numerous technical reports indicate that purely rule-based reward RL frequently results in poor-quality reasoning chains or inconsistencies between reasoning processes and final answers, particularly when the base model is of smaller scale. During the RL exploration process, models might employ low-quality reasoning chains due to the lack of knowledge, occasionally producing correct answers randomly and receiving rewards based on established rule-based judges. This constrains the potential for resource-limited organizations to conduct direct reinforcement learning training on smaller-scale models. We propose a novel confidence-based reward model tailored for enhancing STEM reasoning capabilities. Unlike conventional approaches, our model penalizes not only incorrect answers but also low-confidence correct responses, thereby promoting more robust and logically consistent reasoning. We validate the effectiveness of our approach through static evaluations, Best-of-N inference tests, and PPO-based RL training. Our method outperforms several state-of-the-art open-source reward models across diverse STEM benchmarks. We release our codes and model in https://github.com/qianxiHe147/C2RM.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
He, Qianxi
Ren, Qingyu
Lei, Shanzhe
Wang, Xuhong
Wang, Yingchun
Artificial Intelligence
Machine Learning
Recent advancements in large language models (LLMs) have shifted the post-training paradigm from traditional instruction tuning and human preference alignment toward reinforcement learning (RL) focused on reasoning capabilities. However, numerous technical reports indicate that purely rule-based reward RL frequently results in poor-quality reasoning chains or inconsistencies between reasoning processes and final answers, particularly when the base model is of smaller scale. During the RL exploration process, models might employ low-quality reasoning chains due to the lack of knowledge, occasionally producing correct answers randomly and receiving rewards based on established rule-based judges. This constrains the potential for resource-limited organizations to conduct direct reinforcement learning training on smaller-scale models. We propose a novel confidence-based reward model tailored for enhancing STEM reasoning capabilities. Unlike conventional approaches, our model penalizes not only incorrect answers but also low-confidence correct responses, thereby promoting more robust and logically consistent reasoning. We validate the effectiveness of our approach through static evaluations, Best-of-N inference tests, and PPO-based RL training. Our method outperforms several state-of-the-art open-source reward models across diverse STEM benchmarks. We release our codes and model in https://github.com/qianxiHe147/C2RM.
title Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.07483