Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Xueyang, Zhang, Jingsen, Tang, Jiakai, Li, Wei, Cai, Guohao, Chen, Xu, Dai, Quanyu, Zhu, Yue, Dong, Zhenhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915348323762176
author Feng, Xueyang
Zhang, Jingsen
Tang, Jiakai
Li, Wei
Cai, Guohao
Chen, Xu
Dai, Quanyu
Zhu, Yue
Dong, Zhenhua
author_facet Feng, Xueyang
Zhang, Jingsen
Tang, Jiakai
Li, Wei
Cai, Guohao
Chen, Xu
Dai, Quanyu
Zhu, Yue
Dong, Zhenhua
contents Recent advancements in Large Language Models (LLMs) have significantly propelled the development of Conversational Recommendation Agents (CRAs). However, these agents often generate short-sighted responses that fail to sustain user guidance and meet expectations. Although preference optimization has proven effective in aligning LLMs with user expectations, it remains costly and performs poorly in multi-turn dialogue. To address this challenge, we introduce a novel multi-turn preference optimization (MTPO) paradigm ECPO, which leverages Expectation Confirmation Theory to explicitly model the evolution of user satisfaction throughout multi-turn dialogues, uncovering the underlying causes of dissatisfaction. These causes can be utilized to support targeted optimization of unsatisfactory responses, thereby achieving turn-level preference optimization. ECPO ingeniously eliminates the significant sampling overhead of existing MTPO methods while ensuring the optimization process drives meaningful improvements. To support ECPO, we introduce an LLM-based user simulator, AILO, to simulate user feedback and perform expectation confirmation during conversational recommendations. Experimental results show that ECPO significantly enhances CRA's interaction capabilities, delivering notable improvements in both efficiency and effectiveness over existing MTPO methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
Feng, Xueyang
Zhang, Jingsen
Tang, Jiakai
Li, Wei
Cai, Guohao
Chen, Xu
Dai, Quanyu
Zhu, Yue
Dong, Zhenhua
Computation and Language
Recent advancements in Large Language Models (LLMs) have significantly propelled the development of Conversational Recommendation Agents (CRAs). However, these agents often generate short-sighted responses that fail to sustain user guidance and meet expectations. Although preference optimization has proven effective in aligning LLMs with user expectations, it remains costly and performs poorly in multi-turn dialogue. To address this challenge, we introduce a novel multi-turn preference optimization (MTPO) paradigm ECPO, which leverages Expectation Confirmation Theory to explicitly model the evolution of user satisfaction throughout multi-turn dialogues, uncovering the underlying causes of dissatisfaction. These causes can be utilized to support targeted optimization of unsatisfactory responses, thereby achieving turn-level preference optimization. ECPO ingeniously eliminates the significant sampling overhead of existing MTPO methods while ensuring the optimization process drives meaningful improvements. To support ECPO, we introduce an LLM-based user simulator, AILO, to simulate user feedback and perform expectation confirmation during conversational recommendations. Experimental results show that ECPO significantly enhances CRA's interaction capabilities, delivering notable improvements in both efficiency and effectiveness over existing MTPO methods.
title Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
topic Computation and Language
url https://arxiv.org/abs/2506.14302