MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dang, Yuzhuo, Zhang, Xin, Pan, Zhiqiang, Duan, Yuxiao, Chen, Wanyu, Cai, Fei, Chen, Honghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912844861145088
author Dang, Yuzhuo
Zhang, Xin
Pan, Zhiqiang
Duan, Yuxiao
Chen, Wanyu
Cai, Fei
Chen, Honghui
author_facet Dang, Yuzhuo
Zhang, Xin
Pan, Zhiqiang
Duan, Yuxiao
Chen, Wanyu
Cai, Fei
Chen, Honghui
contents Multimodal recommendation combines the user historical behaviors with the modal features of items to capture the tangible user preferences, presenting superior performance compared to the conventional ID-based recommender systems. However, existing methods still encounter two key problems in the representation learning of users and items, respectively: (1) the initialization of multimodal user representations is either agnostic to historical behaviors or contaminated by irrelevant modal noise, and (2) the widely used KNN-based item-item graph contains noisy edges with low similarities and lacks audience co-occurrence relationships. To address such issues, we propose MLLMRec, a novel preference reasoning paradigm with graph refinement for multimodal recommendation. Specifically, on the one hand, the item images are first converted into high-quality semantic descriptions using a multimodal large language model (MLLM), thereby bridging the semantic gap between visual and textual modalities. Then, we construct a behavioral description list for each user and feed it into the MLLM to reason about the purified user preference profiles that contain the latent interaction intents. On the other hand, we develop the threshold-controlled denoising and topology-aware enhancement strategies to refine the suboptimal item-item graph, thereby improving the accuracy of item representation learning. Extensive experiments on three publicly available datasets demonstrate that MLLMRec achieves the state-of-the-art performance with an average improvement of 21.48% over the optimal baselines. The source code is provided at https://github.com/Yuzhuo-Dang/MLLMRec.git.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation
Dang, Yuzhuo
Zhang, Xin
Pan, Zhiqiang
Duan, Yuxiao
Chen, Wanyu
Cai, Fei
Chen, Honghui
Information Retrieval
Multimodal recommendation combines the user historical behaviors with the modal features of items to capture the tangible user preferences, presenting superior performance compared to the conventional ID-based recommender systems. However, existing methods still encounter two key problems in the representation learning of users and items, respectively: (1) the initialization of multimodal user representations is either agnostic to historical behaviors or contaminated by irrelevant modal noise, and (2) the widely used KNN-based item-item graph contains noisy edges with low similarities and lacks audience co-occurrence relationships. To address such issues, we propose MLLMRec, a novel preference reasoning paradigm with graph refinement for multimodal recommendation. Specifically, on the one hand, the item images are first converted into high-quality semantic descriptions using a multimodal large language model (MLLM), thereby bridging the semantic gap between visual and textual modalities. Then, we construct a behavioral description list for each user and feed it into the MLLM to reason about the purified user preference profiles that contain the latent interaction intents. On the other hand, we develop the threshold-controlled denoising and topology-aware enhancement strategies to refine the suboptimal item-item graph, thereby improving the accuracy of item representation learning. Extensive experiments on three publicly available datasets demonstrate that MLLMRec achieves the state-of-the-art performance with an average improvement of 21.48% over the optimal baselines. The source code is provided at https://github.com/Yuzhuo-Dang/MLLMRec.git.
title MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2508.15304