Saved in:
Bibliographic Details
Main Authors: Zhang, Fan, Huang, Baoru, Zhang, Xin
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.23974
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908867879763968
author Zhang, Fan
Huang, Baoru
Zhang, Xin
author_facet Zhang, Fan
Huang, Baoru
Zhang, Xin
contents Offline reinforcement learning aims to learn an agent from pre-collected datasets, avoiding unsafe and inefficient real-time interaction. However, inevitable access to out-ofdistribution actions during the learning process introduces approximation errors, causing the error accumulation and considerable overestimation. In this paper, we construct a new pessimistic auxiliary policy for sampling reliable actions. Specifically, we develop a pessimistic auxiliary strategy by maximizing the lower confidence bound of the Q-function. The pessimistic auxiliary strategy exhibits a relatively high value and low uncertainty in the vicinity of the learned policy, avoiding the learned policy sampling high-value actions with potentially high errors during the learning process. Less approximation error introduced by sampled action from pessimistic auxiliary strategy leads to the alleviation of error accumulation. Extensive experiments on offline reinforcement learning benchmarks reveal that utilizing the pessimistic auxiliary strategy can effectively improve the efficacy of other offline RL approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2602_23974
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pessimistic Auxiliary Policy for Offline Reinforcement Learning
Zhang, Fan
Huang, Baoru
Zhang, Xin
Artificial Intelligence
Offline reinforcement learning aims to learn an agent from pre-collected datasets, avoiding unsafe and inefficient real-time interaction. However, inevitable access to out-ofdistribution actions during the learning process introduces approximation errors, causing the error accumulation and considerable overestimation. In this paper, we construct a new pessimistic auxiliary policy for sampling reliable actions. Specifically, we develop a pessimistic auxiliary strategy by maximizing the lower confidence bound of the Q-function. The pessimistic auxiliary strategy exhibits a relatively high value and low uncertainty in the vicinity of the learned policy, avoiding the learned policy sampling high-value actions with potentially high errors during the learning process. Less approximation error introduced by sampled action from pessimistic auxiliary strategy leads to the alleviation of error accumulation. Extensive experiments on offline reinforcement learning benchmarks reveal that utilizing the pessimistic auxiliary strategy can effectively improve the efficacy of other offline RL approaches.
title Pessimistic Auxiliary Policy for Offline Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2602.23974