Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yansi, Xu, Jiahao, Liang, Tian, Chen, Xingyu, He, Zhiwei, Liu, Qiuzhi, Wang, Rui, Zhang, Zhuosheng, Tu, Zhaopeng, Mi, Haitao, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913750555033600
author Li, Yansi
Xu, Jiahao
Liang, Tian
Chen, Xingyu
He, Zhiwei
Liu, Qiuzhi
Wang, Rui
Zhang, Zhuosheng
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
author_facet Li, Yansi
Xu, Jiahao
Liang, Tian
Chen, Xingyu
He, Zhiwei
Liu, Qiuzhi
Wang, Rui
Zhang, Zhuosheng
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
contents Enhancing the reasoning capabilities of large language models (LLMs), particularly for complex tasks requiring multi-step logical deductions, remains a significant challenge. Traditional inference time scaling methods utilize scalar reward signals from process reward models to evaluate candidate reasoning steps, but these scalar rewards lack the nuanced qualitative information essential for understanding and justifying each step. In this paper, we propose a novel inference-time scaling approach -- stepwise natural language self-critique (PANEL), which employs self-generated natural language critiques as feedback to guide the step-level search process. By generating rich, human-readable critiques for each candidate reasoning step, PANEL retains essential qualitative information, facilitating better-informed decision-making during inference. This approach bypasses the need for task-specific verifiers and the associated training overhead, making it broadly applicable across diverse tasks. Experimental results on challenging reasoning benchmarks, including AIME and GPQA, demonstrate that PANEL significantly enhances reasoning performance, outperforming traditional scalar reward-based methods. Our code is available at https://github.com/puddingyeah/PANEL to support and encourage future research in this promising field.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
Li, Yansi
Xu, Jiahao
Liang, Tian
Chen, Xingyu
He, Zhiwei
Liu, Qiuzhi
Wang, Rui
Zhang, Zhuosheng
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
Computation and Language
Enhancing the reasoning capabilities of large language models (LLMs), particularly for complex tasks requiring multi-step logical deductions, remains a significant challenge. Traditional inference time scaling methods utilize scalar reward signals from process reward models to evaluate candidate reasoning steps, but these scalar rewards lack the nuanced qualitative information essential for understanding and justifying each step. In this paper, we propose a novel inference-time scaling approach -- stepwise natural language self-critique (PANEL), which employs self-generated natural language critiques as feedback to guide the step-level search process. By generating rich, human-readable critiques for each candidate reasoning step, PANEL retains essential qualitative information, facilitating better-informed decision-making during inference. This approach bypasses the need for task-specific verifiers and the associated training overhead, making it broadly applicable across diverse tasks. Experimental results on challenging reasoning benchmarks, including AIME and GPQA, demonstrate that PANEL significantly enhances reasoning performance, outperforming traditional scalar reward-based methods. Our code is available at https://github.com/puddingyeah/PANEL to support and encourage future research in this promising field.
title Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
topic Computation and Language
url https://arxiv.org/abs/2503.17363