StepWiser: Stepwise Generative Judges for Wiser Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912556497502208 |
|---|---|
| author | Xiong, Wei Zhao, Wenting Yuan, Weizhe Golovneva, Olga Zhang, Tong Weston, Jason Sukhbaatar, Sainbayar |
| author_facet | Xiong, Wei Zhao, Wenting Yuan, Weizhe Golovneva, Olga Zhang, Tong Weston, Jason Sukhbaatar, Sainbayar |
| contents | As models increasingly leverage multi-step reasoning strategies to solve complex problems, supervising the logical validity of these intermediate steps has become a critical research challenge. Process reward models address this by providing step-by-step feedback, but current approaches have two major drawbacks: they typically function as classifiers without providing explanations, and their reliance on supervised fine-tuning with static datasets limits generalization. Inspired by recent advances, we reframe stepwise reward modeling from a classification task to a reasoning task itself. We thus propose a generative judge that reasons about the policy model's reasoning steps (i.e., meta-reasons), outputting thinking tokens before delivering a final verdict. Our model, StepWiser, is trained by reinforcement learning using relative outcomes of rollouts. We show it provides (i) better judgment accuracy on intermediate steps than existing methods; (ii) can be used to improve the policy model at training time; and (iii) improves inference-time search. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_19229 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | StepWiser: Stepwise Generative Judges for Wiser Reasoning Xiong, Wei Zhao, Wenting Yuan, Weizhe Golovneva, Olga Zhang, Tong Weston, Jason Sukhbaatar, Sainbayar Artificial Intelligence Computation and Language As models increasingly leverage multi-step reasoning strategies to solve complex problems, supervising the logical validity of these intermediate steps has become a critical research challenge. Process reward models address this by providing step-by-step feedback, but current approaches have two major drawbacks: they typically function as classifiers without providing explanations, and their reliance on supervised fine-tuning with static datasets limits generalization. Inspired by recent advances, we reframe stepwise reward modeling from a classification task to a reasoning task itself. We thus propose a generative judge that reasons about the policy model's reasoning steps (i.e., meta-reasons), outputting thinking tokens before delivering a final verdict. Our model, StepWiser, is trained by reinforcement learning using relative outcomes of rollouts. We show it provides (i) better judgment accuracy on intermediate steps than existing methods; (ii) can be used to improve the policy model at training time; and (iii) improves inference-time search. |
| title | StepWiser: Stepwise Generative Judges for Wiser Reasoning |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2508.19229 |