OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yiqin, Hu, Hao, Mao, Yihuan, Zhang, Jin, Wu, Chengjie, Jiang, Yuhua, Yang, Xu, Xie, Runpeng, Fan, Yi, Liu, Bo, Gao, Yang, Xu, Bo, Zhang, Chongjie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914441625337856
author Yang, Yiqin
Hu, Hao
Mao, Yihuan
Zhang, Jin
Wu, Chengjie
Jiang, Yuhua
Yang, Xu
Xie, Runpeng
Fan, Yi
Liu, Bo
Gao, Yang
Xu, Bo
Zhang, Chongjie
author_facet Yang, Yiqin
Hu, Hao
Mao, Yihuan
Zhang, Jin
Wu, Chengjie
Jiang, Yuhua
Yang, Xu
Xie, Runpeng
Fan, Yi
Liu, Bo
Gao, Yang
Xu, Bo
Zhang, Chongjie
contents Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02349
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration
Yang, Yiqin
Hu, Hao
Mao, Yihuan
Zhang, Jin
Wu, Chengjie
Jiang, Yuhua
Yang, Xu
Xie, Runpeng
Fan, Yi
Liu, Bo
Gao, Yang
Xu, Bo
Zhang, Chongjie
Machine Learning
Artificial Intelligence
Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.
title OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.02349