DMESR: Dual-view MLLM-based Enhancing Framework for Multimodal Sequential Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Mingyao, Liu, Qidong, Yang, Wenxuan, Wang, Moranxin, Sun, Yuqi, Zhu, Haiping, Tian, Feng, Chen, Yan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917274703626240
author Huang, Mingyao
Liu, Qidong
Yang, Wenxuan
Wang, Moranxin
Sun, Yuqi
Zhu, Haiping
Tian, Feng
Chen, Yan
author_facet Huang, Mingyao
Liu, Qidong
Yang, Wenxuan
Wang, Moranxin
Sun, Yuqi
Zhu, Haiping
Tian, Feng
Chen, Yan
contents Sequential Recommender Systems (SRS) aim to predict users' next interaction based on their historical behaviors, while still facing the challenge of data sparsity. With the rapid advancement of Multimodal Large Language Models (MLLMs), leveraging their multimodal understanding capabilities to enrich item semantic representation has emerged as an effective enhancement strategy for SRS. However, existing MLLM-enhanced recommendation methods still suffer from two key limitations. First, they struggle to effectively align multimodal representations, leading to suboptimal utilization of semantic information across modalities. Second, they often overly rely on MLLM-generated content while overlooking the fine-grained semantic cues contained in the original textual data of items. To address these issues, we propose a Dual-view MLLM-based Enhancing framework for multimodal Sequential Recommendation (DMESR). For the misalignment issue, we employ a contrastive learning mechanism to align the cross-modal semantic representations generated by MLLMs. For the loss of fine-grained semantics, we introduce a cross-attention fusion module that integrates the coarse-grained semantic knowledge obtained from MLLMs with the fine-grained original textual semantics. Finally, these two fused representations can be seamlessly integrated into the downstream sequential recommendation models. Extensive experiments conducted on three real-world datasets and three popular sequential recommendation architectures demonstrate the superior effectiveness and generalizability of our proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13715
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DMESR: Dual-view MLLM-based Enhancing Framework for Multimodal Sequential Recommendation
Huang, Mingyao
Liu, Qidong
Yang, Wenxuan
Wang, Moranxin
Sun, Yuqi
Zhu, Haiping
Tian, Feng
Chen, Yan
Information Retrieval
Sequential Recommender Systems (SRS) aim to predict users' next interaction based on their historical behaviors, while still facing the challenge of data sparsity. With the rapid advancement of Multimodal Large Language Models (MLLMs), leveraging their multimodal understanding capabilities to enrich item semantic representation has emerged as an effective enhancement strategy for SRS. However, existing MLLM-enhanced recommendation methods still suffer from two key limitations. First, they struggle to effectively align multimodal representations, leading to suboptimal utilization of semantic information across modalities. Second, they often overly rely on MLLM-generated content while overlooking the fine-grained semantic cues contained in the original textual data of items. To address these issues, we propose a Dual-view MLLM-based Enhancing framework for multimodal Sequential Recommendation (DMESR). For the misalignment issue, we employ a contrastive learning mechanism to align the cross-modal semantic representations generated by MLLMs. For the loss of fine-grained semantics, we introduce a cross-attention fusion module that integrates the coarse-grained semantic knowledge obtained from MLLMs with the fine-grained original textual semantics. Finally, these two fused representations can be seamlessly integrated into the downstream sequential recommendation models. Extensive experiments conducted on three real-world datasets and three popular sequential recommendation architectures demonstrate the superior effectiveness and generalizability of our proposed approach.
title DMESR: Dual-view MLLM-based Enhancing Framework for Multimodal Sequential Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2602.13715