Reformulating Conversational Recommender Systems as Tri-Phase Offline Policy Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Gangyi, Gao, Chongming, Pan, Hang, Teng, Runzhe, Li, Ruizhe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910593170014208
author Zhang, Gangyi
Gao, Chongming
Pan, Hang
Teng, Runzhe
Li, Ruizhe
author_facet Zhang, Gangyi
Gao, Chongming
Pan, Hang
Teng, Runzhe
Li, Ruizhe
contents Existing Conversational Recommender Systems (CRS) predominantly utilize user simulators for training and evaluating recommendation policies. These simulators often oversimplify the complexity of user interactions by focusing solely on static item attributes, neglecting the rich, evolving preferences that characterize real-world user behavior. This limitation frequently leads to models that perform well in simulated environments but falter in actual deployment. Addressing these challenges, this paper introduces the Tri-Phase Offline Policy Learning-based Conversational Recommender System (TCRS), which significantly reduces dependency on real-time interactions and mitigates overfitting issues prevalent in traditional approaches. TCRS integrates a model-based offline learning strategy with a controllable user simulation that dynamically aligns with both personalized and evolving user preferences. Through comprehensive experiments, TCRS demonstrates enhanced robustness, adaptability, and accuracy in recommendations, outperforming traditional CRS models in diverse user scenarios. This approach not only provides a more realistic evaluation environment but also facilitates a deeper understanding of user behavior dynamics, thereby refining the recommendation process.
format Preprint
id arxiv_https___arxiv_org_abs_2408_06809
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reformulating Conversational Recommender Systems as Tri-Phase Offline Policy Learning
Zhang, Gangyi
Gao, Chongming
Pan, Hang
Teng, Runzhe
Li, Ruizhe
Information Retrieval
Existing Conversational Recommender Systems (CRS) predominantly utilize user simulators for training and evaluating recommendation policies. These simulators often oversimplify the complexity of user interactions by focusing solely on static item attributes, neglecting the rich, evolving preferences that characterize real-world user behavior. This limitation frequently leads to models that perform well in simulated environments but falter in actual deployment. Addressing these challenges, this paper introduces the Tri-Phase Offline Policy Learning-based Conversational Recommender System (TCRS), which significantly reduces dependency on real-time interactions and mitigates overfitting issues prevalent in traditional approaches. TCRS integrates a model-based offline learning strategy with a controllable user simulation that dynamically aligns with both personalized and evolving user preferences. Through comprehensive experiments, TCRS demonstrates enhanced robustness, adaptability, and accuracy in recommendations, outperforming traditional CRS models in diverse user scenarios. This approach not only provides a more realistic evaluation environment but also facilitates a deeper understanding of user behavior dynamics, thereby refining the recommendation process.
title Reformulating Conversational Recommender Systems as Tri-Phase Offline Policy Learning
topic Information Retrieval
url https://arxiv.org/abs/2408.06809