Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bai, Fengshuo, Zhao, Rui, Zhang, Hongming, Cui, Sijia, Wen, Ying, Yang, Yaodong, Xu, Bo, Han, Lei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909212667281408
author Bai, Fengshuo
Zhao, Rui
Zhang, Hongming
Cui, Sijia
Wen, Ying
Yang, Yaodong
Xu, Bo
Han, Lei
author_facet Bai, Fengshuo
Zhao, Rui
Zhang, Hongming
Cui, Sijia
Wen, Ying
Yang, Yaodong
Xu, Bo
Han, Lei
contents Preference-based reinforcement learning (PbRL) has shown impressive capabilities in training agents without reward engineering. However, a notable limitation of PbRL is its dependency on substantial human feedback. This dependency stems from the learning loop, which entails accurate reward learning compounded with value/policy learning, necessitating a considerable number of samples. To boost the learning loop, we propose SEER, an efficient PbRL method that integrates label smoothing and policy regularization techniques. Label smoothing reduces overfitting of the reward model by smoothing human preference labels. Additionally, we bootstrap a conservative estimate $\widehat{Q}$ using well-supported state-action pairs from the current replay memory to mitigate overestimation bias and utilize it for policy learning regularization. Our experimental results across a variety of complex tasks, both in online and offline settings, demonstrate that our approach improves feedback efficiency, outperforming state-of-the-art methods by a large margin. Ablation studies further reveal that SEER achieves a more accurate Q-function compared to prior work.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18688
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation
Bai, Fengshuo
Zhao, Rui
Zhang, Hongming
Cui, Sijia
Wen, Ying
Yang, Yaodong
Xu, Bo
Han, Lei
Machine Learning
Artificial Intelligence
Computation and Language
Preference-based reinforcement learning (PbRL) has shown impressive capabilities in training agents without reward engineering. However, a notable limitation of PbRL is its dependency on substantial human feedback. This dependency stems from the learning loop, which entails accurate reward learning compounded with value/policy learning, necessitating a considerable number of samples. To boost the learning loop, we propose SEER, an efficient PbRL method that integrates label smoothing and policy regularization techniques. Label smoothing reduces overfitting of the reward model by smoothing human preference labels. Additionally, we bootstrap a conservative estimate $\widehat{Q}$ using well-supported state-action pairs from the current replay memory to mitigate overestimation bias and utilize it for policy learning regularization. Our experimental results across a variety of complex tasks, both in online and offline settings, demonstrate that our approach improves feedback efficiency, outperforming state-of-the-art methods by a large margin. Ablation studies further reveal that SEER achieves a more accurate Q-function compared to prior work.
title Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.18688