Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dong, Jinzong, Huang, Wei, Zhang, Jianshu, Chen, Zhuo, Yuan, Xinzhe, Gu, Qinying, Jiang, Zhaohui, Ye, Nanyang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911685107777536
author Dong, Jinzong
Huang, Wei
Zhang, Jianshu
Chen, Zhuo
Yuan, Xinzhe
Gu, Qinying
Jiang, Zhaohui
Ye, Nanyang
author_facet Dong, Jinzong
Huang, Wei
Zhang, Jianshu
Chen, Zhuo
Yuan, Xinzhe
Gu, Qinying
Jiang, Zhaohui
Ye, Nanyang
contents Offline reinforcement learning (RL), which optimizes policies using a previously collected static dataset, is an important branch of RL. A popular and promising approach is to regularize actor-critic methods with behavior cloning (BC), which quickly yields realistic policies and mitigates bias from out-of-distribution actions, but it can impose an often-overlooked performance ceiling: when dataset actions are suboptimal, indiscriminate imitation structurally prevents the actor from fully exploiting better actions suggested by the value function, especially in later training when imitation is already dominant. We formally analyzed this limitation by investigating convergence properties of BC-regularized actor-critic optimization and verified it on a controlled continuous bandit task. To break this ceiling, we propose proximal action replacement (PAR), an easy-to-use plug-and-play training sample replacer. PAR substitutes suboptimal dataset actions with better actions generated by a stable target policy, guided by the action-value function's local ascent direction and bounded by value uncertainty to ensure training stability. PAR is compatible with multiple BC regularization paradigms. Extensive experiments across offline RL benchmarks show that PAR consistently improves performance, and approaches state-of-the-art results simply by being combined with the basic TD3+BC.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07441
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning
Dong, Jinzong
Huang, Wei
Zhang, Jianshu
Chen, Zhuo
Yuan, Xinzhe
Gu, Qinying
Jiang, Zhaohui
Ye, Nanyang
Machine Learning
Artificial Intelligence
Offline reinforcement learning (RL), which optimizes policies using a previously collected static dataset, is an important branch of RL. A popular and promising approach is to regularize actor-critic methods with behavior cloning (BC), which quickly yields realistic policies and mitigates bias from out-of-distribution actions, but it can impose an often-overlooked performance ceiling: when dataset actions are suboptimal, indiscriminate imitation structurally prevents the actor from fully exploiting better actions suggested by the value function, especially in later training when imitation is already dominant. We formally analyzed this limitation by investigating convergence properties of BC-regularized actor-critic optimization and verified it on a controlled continuous bandit task. To break this ceiling, we propose proximal action replacement (PAR), an easy-to-use plug-and-play training sample replacer. PAR substitutes suboptimal dataset actions with better actions generated by a stable target policy, guided by the action-value function's local ascent direction and bounded by value uncertainty to ensure training stability. PAR is compatible with multiple BC regularization paradigms. Extensive experiments across offline RL benchmarks show that PAR consistently improves performance, and approaches state-of-the-art results simply by being combined with the basic TD3+BC.
title Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.07441