ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yi, Qiu, Ruihong, Liu, Jiajun, Wang, Sen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909606663421952
author Zhang, Yi
Qiu, Ruihong
Liu, Jiajun
Wang, Sen
author_facet Zhang, Yi
Qiu, Ruihong
Liu, Jiajun
Wang, Sen
contents Offline reinforcement learning (RL) is an effective tool for real-world recommender systems with its capacity to model the dynamic interest of users and its interactive nature. Most existing offline RL recommender systems focus on model-based RL through learning a world model from offline data and building the recommendation policy by interacting with this model. Although these methods have made progress in the recommendation performance, the effectiveness of model-based offline RL methods is often constrained by the accuracy of the estimation of the reward model and the model uncertainties, primarily due to the extreme discrepancy between offline logged data and real-world data in user interactions with online platforms. To fill this gap, a more accurate reward model and uncertainty estimation are needed for the model-based RL methods. In this paper, a novel model-based Reward Shaping in Offline Reinforcement Learning for Recommender Systems, ROLeR, is proposed for reward and uncertainty estimation in recommendation systems. Specifically, a non-parametric reward shaping method is designed to refine the reward model. In addition, a flexible and more representative uncertainty penalty is designed to fit the needs of recommendation systems. Extensive experiments conducted on four benchmark datasets showcase that ROLeR achieves state-of-the-art performance compared with existing baselines. The source code can be downloaded at https://github.com/ArronDZhang/ROLeR.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13163
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender Systems
Zhang, Yi
Qiu, Ruihong
Liu, Jiajun
Wang, Sen
Information Retrieval
Artificial Intelligence
Offline reinforcement learning (RL) is an effective tool for real-world recommender systems with its capacity to model the dynamic interest of users and its interactive nature. Most existing offline RL recommender systems focus on model-based RL through learning a world model from offline data and building the recommendation policy by interacting with this model. Although these methods have made progress in the recommendation performance, the effectiveness of model-based offline RL methods is often constrained by the accuracy of the estimation of the reward model and the model uncertainties, primarily due to the extreme discrepancy between offline logged data and real-world data in user interactions with online platforms. To fill this gap, a more accurate reward model and uncertainty estimation are needed for the model-based RL methods. In this paper, a novel model-based Reward Shaping in Offline Reinforcement Learning for Recommender Systems, ROLeR, is proposed for reward and uncertainty estimation in recommendation systems. Specifically, a non-parametric reward shaping method is designed to refine the reward model. In addition, a flexible and more representative uncertainty penalty is designed to fit the needs of recommendation systems. Extensive experiments conducted on four benchmark datasets showcase that ROLeR achieves state-of-the-art performance compared with existing baselines. The source code can be downloaded at https://github.com/ArronDZhang/ROLeR.
title ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender Systems
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2407.13163