Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zihan, Yang, Xiaocui, Liu, Yongkang, Feng, Shi, Wang, Daling, Zhang, Yifei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912327910031360
author Wang, Zihan
Yang, Xiaocui
Liu, Yongkang
Feng, Shi
Wang, Daling
Zhang, Yifei
author_facet Wang, Zihan
Yang, Xiaocui
Liu, Yongkang
Feng, Shi
Wang, Daling
Zhang, Yifei
contents Current conversational recommendation systems focus predominantly on text. However, real-world recommendation settings are generally multimodal, causing a significant gap between existing research and practical applications. To address this issue, we propose Muse, the first multimodal conversational recommendation dataset. Muse comprises 83,148 utterances from 7,000 conversations centered around the Clothing domain. Each conversation contains comprehensive multimodal interactions, rich elements, and natural dialogues. Data in Muse are automatically synthesized by a multi-agent framework powered by multimodal large language models (MLLMs). It innovatively derives user profiles from real-world scenarios rather than depending on manual design and history data for better scalability, and then it fulfills conversation simulation and optimization. Both human and LLM evaluations demonstrate the high quality of conversations in Muse. Additionally, fine-tuning experiments on three MLLMs demonstrate Muse's learnable patterns for recommendations and responses, confirming its value for multimodal conversational recommendation. Our dataset and codes are available at https://anonymous.4open.science/r/Muse-0086.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18416
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
Wang, Zihan
Yang, Xiaocui
Liu, Yongkang
Feng, Shi
Wang, Daling
Zhang, Yifei
Multimedia
Current conversational recommendation systems focus predominantly on text. However, real-world recommendation settings are generally multimodal, causing a significant gap between existing research and practical applications. To address this issue, we propose Muse, the first multimodal conversational recommendation dataset. Muse comprises 83,148 utterances from 7,000 conversations centered around the Clothing domain. Each conversation contains comprehensive multimodal interactions, rich elements, and natural dialogues. Data in Muse are automatically synthesized by a multi-agent framework powered by multimodal large language models (MLLMs). It innovatively derives user profiles from real-world scenarios rather than depending on manual design and history data for better scalability, and then it fulfills conversation simulation and optimization. Both human and LLM evaluations demonstrate the high quality of conversations in Muse. Additionally, fine-tuning experiments on three MLLMs demonstrate Muse's learnable patterns for recommendations and responses, confirming its value for multimodal conversational recommendation. Our dataset and codes are available at https://anonymous.4open.science/r/Muse-0086.
title Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
topic Multimedia
url https://arxiv.org/abs/2412.18416