Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Tianyi, Li, Long, Guo, Hongcan, Chen, Yibiao, Li, Yixia, Wang, Yong, Chen, Yun, Chen, Guanhua
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2602.05717
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911424515670016
author Wang, Tianyi
Li, Long
Guo, Hongcan
Chen, Yibiao
Li, Yixia
Wang, Yong
Chen, Yun
Chen, Guanhua
author_facet Wang, Tianyi
Li, Long
Guo, Hongcan
Chen, Yibiao
Li, Yixia
Wang, Yong
Chen, Yun
Chen, Guanhua
contents Reinforcement Learning with Verifiable Rewards (RLVR) is increasingly viewed as a tree pruning mechanism. However, we identify a systemic pathology termed Recursive Space Contraction (RSC), an irreversible collapse driven by the combined dynamics of positive sharpening and negative squeezing, where the sampling probability of valid alternatives vanishes. While Kullback-Leibler (KL) regularization aims to mitigate this, it imposes a rigid Shape Matching constraint that forces the policy to mimic the reference model's full density, creating a gradient conflict with the sharpening required for correctness. We propose Anchored Policy Optimization (APO), shifting the paradigm from global Shape Matching to Support Coverage. By defining a Safe Manifold based on the reference model's high-confidence support, APO permits aggressive sharpening for efficiency while selectively invoking a restorative force during error correction to prevent collapse. We theoretically derive that APO serves as a gradient-aligned mechanism to maximize support coverage, enabling an Elastic Recovery that re-inflates valid branches. Empirical evaluations on mathematical benchmarks demonstrate that APO breaks the accuracy-diversity trade-off, significantly improving Pass@1 while restoring the Pass@K diversity typically lost by standard policy gradient methods.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05717
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification
Wang, Tianyi
Li, Long
Guo, Hongcan
Chen, Yibiao
Li, Yixia
Wang, Yong
Chen, Yun
Chen, Guanhua
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) is increasingly viewed as a tree pruning mechanism. However, we identify a systemic pathology termed Recursive Space Contraction (RSC), an irreversible collapse driven by the combined dynamics of positive sharpening and negative squeezing, where the sampling probability of valid alternatives vanishes. While Kullback-Leibler (KL) regularization aims to mitigate this, it imposes a rigid Shape Matching constraint that forces the policy to mimic the reference model's full density, creating a gradient conflict with the sharpening required for correctness. We propose Anchored Policy Optimization (APO), shifting the paradigm from global Shape Matching to Support Coverage. By defining a Safe Manifold based on the reference model's high-confidence support, APO permits aggressive sharpening for efficiency while selectively invoking a restorative force during error correction to prevent collapse. We theoretically derive that APO serves as a gradient-aligned mechanism to maximize support coverage, enabling an Elastic Recovery that re-inflates valid branches. Empirical evaluations on mathematical benchmarks demonstrate that APO breaks the accuracy-diversity trade-off, significantly improving Pass@1 while restoring the Pass@K diversity typically lost by standard policy gradient methods.
title Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification
topic Artificial Intelligence
url https://arxiv.org/abs/2602.05717