MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yibiao, Zou, Jie, Guo, Weikang, Wang, Guoqing, Xu, Xing, Yang, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913807265169408
author Wei, Yibiao
Zou, Jie
Guo, Weikang
Wang, Guoqing
Xu, Xing
Yang, Yang
author_facet Wei, Yibiao
Zou, Jie
Guo, Weikang
Wang, Guoqing
Xu, Xing
Yang, Yang
contents Conversational Recommender Systems (CRSs) aim to provide personalized recommendations by interacting with users through conversations. Most existing studies of CRS focus on extracting user preferences from conversational contexts. However, due to the short and sparse nature of conversational contexts, it is difficult to fully capture user preferences by conversational contexts only. We argue that multi-modal semantic information can enrich user preference expressions from diverse dimensions (e.g., a user preference for a certain movie may stem from its magnificent visual effects and compelling storyline). In this paper, we propose a multi-modal semantic graph prompt learning framework for CRS, named MSCRS. First, we extract textual and image features of items mentioned in the conversational contexts. Second, we capture higher-order semantic associations within different semantic modalities (collaborative, textual, and image) by constructing modality-specific graph structures. Finally, we propose an innovative integration of multi-modal semantic graphs with prompt learning, harnessing the power of large language models to comprehensively explore high-dimensional semantic relationships. Experimental results demonstrate that our proposed method significantly improves accuracy in item recommendation, as well as generates more natural and contextually relevant content in response generation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems
Wei, Yibiao
Zou, Jie
Guo, Weikang
Wang, Guoqing
Xu, Xing
Yang, Yang
Information Retrieval
Conversational Recommender Systems (CRSs) aim to provide personalized recommendations by interacting with users through conversations. Most existing studies of CRS focus on extracting user preferences from conversational contexts. However, due to the short and sparse nature of conversational contexts, it is difficult to fully capture user preferences by conversational contexts only. We argue that multi-modal semantic information can enrich user preference expressions from diverse dimensions (e.g., a user preference for a certain movie may stem from its magnificent visual effects and compelling storyline). In this paper, we propose a multi-modal semantic graph prompt learning framework for CRS, named MSCRS. First, we extract textual and image features of items mentioned in the conversational contexts. Second, we capture higher-order semantic associations within different semantic modalities (collaborative, textual, and image) by constructing modality-specific graph structures. Finally, we propose an innovative integration of multi-modal semantic graphs with prompt learning, harnessing the power of large language models to comprehensively explore high-dimensional semantic relationships. Experimental results demonstrate that our proposed method significantly improves accuracy in item recommendation, as well as generates more natural and contextually relevant content in response generation.
title MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems
topic Information Retrieval
url https://arxiv.org/abs/2504.10921