SPAR: Support-Preserving Action Rectification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jiaxin, Pan, Weihang, Liang, Xun, Lin, Binbin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911723522359296
author Zhao, Jiaxin
Pan, Weihang
Liang, Xun
Lin, Binbin
author_facet Zhao, Jiaxin
Pan, Weihang
Liang, Xun
Lin, Binbin
contents Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression is stable, it suffers from over-conservatism that suppresses high-value actions in the distribution tail; conversely, gradient-based approaches often exhibit a fitting-optimization conflict of gradients, which drives the policy off the data manifold. To address this, we propose Support-Preserving Action Rectification (SPAR), which reframes global learning as a local residual rectification anchored to a frozen pure behavior cloning policy. This framework performs fine-grained fitting and local policy improvement in the residual space, thereby contracting the search space. We further introduce Latent Self-Imitation, utilizing a latent-sampling weighted-regression mechanism to address fitting-improvement gradient conflict in the residual space. Theoretically, we prove this mechanism eliminates the manifold-normal drift of standard value gradients, while extensive D4RL experiments show SPAR extracts significant gains from suboptimal baselines to achieve state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27877
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SPAR: Support-Preserving Action Rectification
Zhao, Jiaxin
Pan, Weihang
Liang, Xun
Lin, Binbin
Machine Learning
Artificial Intelligence
Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression is stable, it suffers from over-conservatism that suppresses high-value actions in the distribution tail; conversely, gradient-based approaches often exhibit a fitting-optimization conflict of gradients, which drives the policy off the data manifold. To address this, we propose Support-Preserving Action Rectification (SPAR), which reframes global learning as a local residual rectification anchored to a frozen pure behavior cloning policy. This framework performs fine-grained fitting and local policy improvement in the residual space, thereby contracting the search space. We further introduce Latent Self-Imitation, utilizing a latent-sampling weighted-regression mechanism to address fitting-improvement gradient conflict in the residual space. Theoretically, we prove this mechanism eliminates the manifold-normal drift of standard value gradients, while extensive D4RL experiments show SPAR extracts significant gains from suboptimal baselines to achieve state-of-the-art performance.
title SPAR: Support-Preserving Action Rectification
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.27877