StepWiser: Stepwise Generative Judges for Wiser Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Wei, Zhao, Wenting, Yuan, Weizhe, Golovneva, Olga, Zhang, Tong, Weston, Jason, Sukhbaatar, Sainbayar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912556497502208
author Xiong, Wei
Zhao, Wenting
Yuan, Weizhe
Golovneva, Olga
Zhang, Tong
Weston, Jason
Sukhbaatar, Sainbayar
author_facet Xiong, Wei
Zhao, Wenting
Yuan, Weizhe
Golovneva, Olga
Zhang, Tong
Weston, Jason
Sukhbaatar, Sainbayar
contents As models increasingly leverage multi-step reasoning strategies to solve complex problems, supervising the logical validity of these intermediate steps has become a critical research challenge. Process reward models address this by providing step-by-step feedback, but current approaches have two major drawbacks: they typically function as classifiers without providing explanations, and their reliance on supervised fine-tuning with static datasets limits generalization. Inspired by recent advances, we reframe stepwise reward modeling from a classification task to a reasoning task itself. We thus propose a generative judge that reasons about the policy model's reasoning steps (i.e., meta-reasons), outputting thinking tokens before delivering a final verdict. Our model, StepWiser, is trained by reinforcement learning using relative outcomes of rollouts. We show it provides (i) better judgment accuracy on intermediate steps than existing methods; (ii) can be used to improve the policy model at training time; and (iii) improves inference-time search.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StepWiser: Stepwise Generative Judges for Wiser Reasoning
Xiong, Wei
Zhao, Wenting
Yuan, Weizhe
Golovneva, Olga
Zhang, Tong
Weston, Jason
Sukhbaatar, Sainbayar
Artificial Intelligence
Computation and Language
As models increasingly leverage multi-step reasoning strategies to solve complex problems, supervising the logical validity of these intermediate steps has become a critical research challenge. Process reward models address this by providing step-by-step feedback, but current approaches have two major drawbacks: they typically function as classifiers without providing explanations, and their reliance on supervised fine-tuning with static datasets limits generalization. Inspired by recent advances, we reframe stepwise reward modeling from a classification task to a reasoning task itself. We thus propose a generative judge that reasons about the policy model's reasoning steps (i.e., meta-reasons), outputting thinking tokens before delivering a final verdict. Our model, StepWiser, is trained by reinforcement learning using relative outcomes of rollouts. We show it provides (i) better judgment accuracy on intermediate steps than existing methods; (ii) can be used to improve the policy model at training time; and (iii) improves inference-time search.
title StepWiser: Stepwise Generative Judges for Wiser Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.19229