Credit Assignment with Resets in Language Model Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Samanta, Ankur, Magesh, Akshayaa, Jain, Ayush, Yu, Youliang, Jiang, Daniel, Asadi, Kavosh, Hassani, Kaveh, Sajda, Paul, Bhandari, Jalaj, Efroni, Yonathan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913163916607488
author Samanta, Ankur
Magesh, Akshayaa
Jain, Ayush
Yu, Youliang
Jiang, Daniel
Asadi, Kavosh
Hassani, Kaveh
Sajda, Paul
Bhandari, Jalaj
Efroni, Yonathan
author_facet Samanta, Ankur
Magesh, Akshayaa
Jain, Ayush
Yu, Youliang
Jiang, Daniel
Asadi, Kavosh
Hassani, Kaveh
Sajda, Paul
Bhandari, Jalaj
Efroni, Yonathan
contents Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all tokens in a trajectory. Such uniform assignment ignores which steps contributed to success or failure. Improving credit assignment can address this limitation by enabling targeted refinement of faulty reasoning steps, rather than updating entire trajectories uniformly. Resets are one such simple mechanism, enabling more precise credit assignment by returning to an intermediate state and resampling counterfactual continuations, so that outcome differences can be attributed to decisions made at that point. We propose two such methods: Random-Reset Policy Optimization (RRPO), where reset states are drawn randomly from reasoning steps, and Self-Reset Policy Optimization (SRPO), where the model self-localizes the erroneous step in an incorrect trajectory and resets there. We analyze these methods within the Conservative Policy Iteration (CPI) framework. Extending CPI with a credit-assignment oracle that targets improvable states yields provable improvements over random resets. Across models and reasoning benchmarks, SRPO consistently outperforms standard GRPO and RRPO by sampling multiple suffix continuations at a self-localized reset and learning from their rewards, using only the model itself with no external supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25507
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Credit Assignment with Resets in Language Model Reasoning
Samanta, Ankur
Magesh, Akshayaa
Jain, Ayush
Yu, Youliang
Jiang, Daniel
Asadi, Kavosh
Hassani, Kaveh
Sajda, Paul
Bhandari, Jalaj
Efroni, Yonathan
Artificial Intelligence
Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all tokens in a trajectory. Such uniform assignment ignores which steps contributed to success or failure. Improving credit assignment can address this limitation by enabling targeted refinement of faulty reasoning steps, rather than updating entire trajectories uniformly. Resets are one such simple mechanism, enabling more precise credit assignment by returning to an intermediate state and resampling counterfactual continuations, so that outcome differences can be attributed to decisions made at that point. We propose two such methods: Random-Reset Policy Optimization (RRPO), where reset states are drawn randomly from reasoning steps, and Self-Reset Policy Optimization (SRPO), where the model self-localizes the erroneous step in an incorrect trajectory and resets there. We analyze these methods within the Conservative Policy Iteration (CPI) framework. Extending CPI with a credit-assignment oracle that targets improvable states yields provable improvements over random resets. Across models and reasoning benchmarks, SRPO consistently outperforms standard GRPO and RRPO by sampling multiple suffix continuations at a self-localized reset and learning from their rewards, using only the model itself with no external supervision.
title Credit Assignment with Resets in Language Model Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2605.25507