Semi-pessimistic Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Jin, Zhou, Xin, Yao, Jiaang, Aminian, Gholamali, Rivasplata, Omar, Little, Simon, Li, Lexin, Shi, Chengchun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915303988920320
author Zhu, Jin
Zhou, Xin
Yao, Jiaang
Aminian, Gholamali
Rivasplata, Omar
Little, Simon
Li, Lexin
Shi, Chengchun
author_facet Zhu, Jin
Zhou, Xin
Yao, Jiaang
Aminian, Gholamali
Rivasplata, Omar
Little, Simon
Li, Lexin
Shi, Chengchun
contents Offline reinforcement learning (RL) aims to learn an optimal policy from pre-collected data. However, it faces challenges of distributional shift, where the learned policy may encounter unseen scenarios not covered in the offline data. Additionally, numerous applications suffer from a scarcity of labeled reward data. Relying on labeled data alone often leads to a narrow state-action distribution, further amplifying the distributional shift, and resulting in suboptimal policy learning. To address these issues, we first recognize that the volume of unlabeled data is typically substantially larger than that of labeled data. We then propose a semi-pessimistic RL method to effectively leverage abundant unlabeled data. Our approach offers several advantages. It considerably simplifies the learning process, as it seeks a lower bound of the reward function, rather than that of the Q-function or state transition function. It is highly flexible, and can be integrated with a range of model-free and model-based RL algorithms. It enjoys the guaranteed improvement when utilizing vast unlabeled data, but requires much less restrictive conditions. We compare our method with a number of alternative solutions, both analytically and numerically, and demonstrate its clear competitiveness. We further illustrate with an application to adaptive deep brain stimulation for Parkinson's disease.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19002
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semi-pessimistic Reinforcement Learning
Zhu, Jin
Zhou, Xin
Yao, Jiaang
Aminian, Gholamali
Rivasplata, Omar
Little, Simon
Li, Lexin
Shi, Chengchun
Machine Learning
Artificial Intelligence
Offline reinforcement learning (RL) aims to learn an optimal policy from pre-collected data. However, it faces challenges of distributional shift, where the learned policy may encounter unseen scenarios not covered in the offline data. Additionally, numerous applications suffer from a scarcity of labeled reward data. Relying on labeled data alone often leads to a narrow state-action distribution, further amplifying the distributional shift, and resulting in suboptimal policy learning. To address these issues, we first recognize that the volume of unlabeled data is typically substantially larger than that of labeled data. We then propose a semi-pessimistic RL method to effectively leverage abundant unlabeled data. Our approach offers several advantages. It considerably simplifies the learning process, as it seeks a lower bound of the reward function, rather than that of the Q-function or state transition function. It is highly flexible, and can be integrated with a range of model-free and model-based RL algorithms. It enjoys the guaranteed improvement when utilizing vast unlabeled data, but requires much less restrictive conditions. We compare our method with a number of alternative solutions, both analytically and numerically, and demonstrate its clear competitiveness. We further illustrate with an application to adaptive deep brain stimulation for Parkinson's disease.
title Semi-pessimistic Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.19002