RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Qiaoyu, Xiang, Hao, Yu, Le, Yu, Bowen, Lin, Hongyu, Lu, Yaojie, Han, Xianpei, Sun, Le, Lin, Junyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911066741538816
author Tang, Qiaoyu
Xiang, Hao
Yu, Le
Yu, Bowen
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Sun, Le
Lin, Junyang
author_facet Tang, Qiaoyu
Xiang, Hao
Yu, Le
Yu, Bowen
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Sun, Le
Lin, Junyang
contents With the rapid advancement of Large Language Models (LLMs), developing effective critic modules for precise guidance has become crucial yet challenging. In this paper, we initially demonstrate that supervised fine-tuning for building critic modules (which is widely adopted in current solutions) fails to genuinely enhance models' critique abilities, producing superficial critiques with insufficient reflections and verifications. To unlock the unprecedented critique capabilities, we propose RefCritic, a long-chain-of-thought critic module based on reinforcement learning with dual rule-based rewards: (1) instance-level correctness of solution judgments and (2) refinement accuracies of the policy model based on critiques, aiming to generate high-quality evaluations with actionable feedback that effectively guides model refinement. We evaluate RefCritic on Qwen2.5-14B-Instruct and DeepSeek-R1-Distill-Qwen-14B across five benchmarks. On critique and refinement settings, RefCritic demonstrates consistent advantages across all benchmarks, e.g., 6.8\% and 7.2\% gains on AIME25 for the respective base models. Notably, under majority voting, policy models filtered by RefCritic show superior scaling with increased voting numbers. Moreover, despite training on solution-level supervision, RefCritic outperforms step-level supervised approaches on ProcessBench, a benchmark to identify erroneous steps in mathematical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
Tang, Qiaoyu
Xiang, Hao
Yu, Le
Yu, Bowen
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Sun, Le
Lin, Junyang
Computation and Language
With the rapid advancement of Large Language Models (LLMs), developing effective critic modules for precise guidance has become crucial yet challenging. In this paper, we initially demonstrate that supervised fine-tuning for building critic modules (which is widely adopted in current solutions) fails to genuinely enhance models' critique abilities, producing superficial critiques with insufficient reflections and verifications. To unlock the unprecedented critique capabilities, we propose RefCritic, a long-chain-of-thought critic module based on reinforcement learning with dual rule-based rewards: (1) instance-level correctness of solution judgments and (2) refinement accuracies of the policy model based on critiques, aiming to generate high-quality evaluations with actionable feedback that effectively guides model refinement. We evaluate RefCritic on Qwen2.5-14B-Instruct and DeepSeek-R1-Distill-Qwen-14B across five benchmarks. On critique and refinement settings, RefCritic demonstrates consistent advantages across all benchmarks, e.g., 6.8\% and 7.2\% gains on AIME25 for the respective base models. Notably, under majority voting, policy models filtered by RefCritic show superior scaling with increased voting numbers. Moreover, despite training on solution-level supervision, RefCritic outperforms step-level supervised approaches on ProcessBench, a benchmark to identify erroneous steps in mathematical reasoning.
title RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
topic Computation and Language
url https://arxiv.org/abs/2507.15024