Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yuning, Wang, Ke, Chen, Devin, Wei, Kai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917383513309184
author Wu, Yuning
Wang, Ke
Chen, Devin
Wei, Kai
author_facet Wu, Yuning
Wang, Ke
Chen, Devin
Wei, Kai
contents Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Policy Optimization (GRPO) face a critical dilemma in sparse-reward settings: pure Reinforcement Learning (RL) suffers from advantage collapse and high-variance gradient estimation, while mixed-policy optimization introduces persistent distributional bias. To resolve this dilemma, we introduce Hindsight-Anchored Policy Optimization (HAPO). HAPO employs the Synthetic Success Injection (SSI) operator, a hindsight mechanism that selectively anchors optimization to teacher demonstrations during failure. This injection is governed by a Thompson sampling-inspired gating mechanism, creating an autonomous, self-paced curriculum. Theoretically, we demonstrate that HAPO achieves \textit{asymptotic consistency}: by naturally annealing the teacher signal as the policy improves, HAPO recovers the unbiased on-policy gradient. This ensures off-policy guidance acts as a temporary scaffold rather than a persistent ceiling, enabling the model to surpass the limitations of static teacher forcing.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
Wu, Yuning
Wang, Ke
Chen, Devin
Wei, Kai
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Policy Optimization (GRPO) face a critical dilemma in sparse-reward settings: pure Reinforcement Learning (RL) suffers from advantage collapse and high-variance gradient estimation, while mixed-policy optimization introduces persistent distributional bias. To resolve this dilemma, we introduce Hindsight-Anchored Policy Optimization (HAPO). HAPO employs the Synthetic Success Injection (SSI) operator, a hindsight mechanism that selectively anchors optimization to teacher demonstrations during failure. This injection is governed by a Thompson sampling-inspired gating mechanism, creating an autonomous, self-paced curriculum. Theoretically, we demonstrate that HAPO achieves \textit{asymptotic consistency}: by naturally annealing the teacher signal as the policy improves, HAPO recovers the unbiased on-policy gradient. This ensures off-policy guidance acts as a temporary scaffold rather than a persistent ceiling, enabling the model to surpass the limitations of static teacher forcing.
title Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.11321