Future Impact Decomposition in Request-level Recommendations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaobei, Liu, Shuchang, Wang, Xueliang, Cai, Qingpeng, Hu, Lantao, Li, Han, Jiang, Peng, Gai, Kun, Xie, Guangming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929390752890880
author Wang, Xiaobei
Liu, Shuchang
Wang, Xueliang
Cai, Qingpeng
Hu, Lantao
Li, Han
Jiang, Peng
Gai, Kun
Xie, Guangming
author_facet Wang, Xiaobei
Liu, Shuchang
Wang, Xueliang
Cai, Qingpeng
Hu, Lantao
Li, Han
Jiang, Peng
Gai, Kun
Xie, Guangming
contents In recommender systems, reinforcement learning solutions have shown promising results in optimizing the interaction sequence between users and the system over the long-term performance. For practical reasons, the policy's actions are typically designed as recommending a list of items to handle users' frequent and continuous browsing requests more efficiently. In this list-wise recommendation scenario, the user state is updated upon every request in the corresponding MDP formulation. However, this request-level formulation is essentially inconsistent with the user's item-level behavior. In this study, we demonstrate that an item-level optimization approach can better utilize item characteristics and optimize the policy's performance even under the request-level MDP. We support this claim by comparing the performance of standard request-level methods with the proposed item-level actor-critic framework in both simulation and online experiments. Furthermore, we show that a reward-based future decomposition strategy can better express the item-wise future impact and improve the recommendation accuracy in the long term. To achieve a more thorough understanding of the decomposition strategy, we propose a model-based re-weighting framework with adversarial learning that further boost the performance and investigate its correlation with the reward-based strategy.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16108
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Future Impact Decomposition in Request-level Recommendations
Wang, Xiaobei
Liu, Shuchang
Wang, Xueliang
Cai, Qingpeng
Hu, Lantao
Li, Han
Jiang, Peng
Gai, Kun
Xie, Guangming
Information Retrieval
H.3.3
In recommender systems, reinforcement learning solutions have shown promising results in optimizing the interaction sequence between users and the system over the long-term performance. For practical reasons, the policy's actions are typically designed as recommending a list of items to handle users' frequent and continuous browsing requests more efficiently. In this list-wise recommendation scenario, the user state is updated upon every request in the corresponding MDP formulation. However, this request-level formulation is essentially inconsistent with the user's item-level behavior. In this study, we demonstrate that an item-level optimization approach can better utilize item characteristics and optimize the policy's performance even under the request-level MDP. We support this claim by comparing the performance of standard request-level methods with the proposed item-level actor-critic framework in both simulation and online experiments. Furthermore, we show that a reward-based future decomposition strategy can better express the item-wise future impact and improve the recommendation accuracy in the long term. To achieve a more thorough understanding of the decomposition strategy, we propose a model-based re-weighting framework with adversarial learning that further boost the performance and investigate its correlation with the reward-based strategy.
title Future Impact Decomposition in Request-level Recommendations
topic Information Retrieval
H.3.3
url https://arxiv.org/abs/2401.16108