Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hui, Qin, Jinghui, Wen, Wushao, Li, Qingling, Zhong, Shanshan, Huang, Zhongzhan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909804912443392
author Wang, Hui
Qin, Jinghui
Wen, Wushao
Li, Qingling
Zhong, Shanshan
Huang, Zhongzhan
author_facet Wang, Hui
Qin, Jinghui
Wen, Wushao
Li, Qingling
Zhong, Shanshan
Huang, Zhongzhan
contents Multimodal data has significantly advanced recommendation systems by integrating diverse information sources to model user preferences and item characteristics. However, these systems often struggle with redundant and irrelevant information, which can degrade performance. Most existing methods either fuse multimodal information directly or use rigid architectural separation for disentanglement, failing to adequately filter noise and model the complex interplay between modalities. To address these challenges, we propose a novel framework, the Multimodal Representation-disentangled Information Bottleneck (MRdIB). Concretely, we first employ a Multimodal Information Bottleneck to compress the input representations, effectively filtering out task-irrelevant noise while preserving rich semantic information. Then, we decompose the information based on its relationship with the recommendation target into unique, redundant, and synergistic components. We achieve this decomposition with a series of constraints: a unique information learning objective to preserve modality-unique signals, a redundant information learning objective to minimize overlap, and a synergistic information learning objective to capture emergent information. By optimizing these objectives, MRdIB guides a model to learn more powerful and disentangled representations. Extensive experiments on several competitive models and three benchmark datasets demonstrate the effectiveness and versatility of our MRdIB in enhancing multimodal recommendation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation
Wang, Hui
Qin, Jinghui
Wen, Wushao
Li, Qingling
Zhong, Shanshan
Huang, Zhongzhan
Information Retrieval
Artificial Intelligence
Multimodal data has significantly advanced recommendation systems by integrating diverse information sources to model user preferences and item characteristics. However, these systems often struggle with redundant and irrelevant information, which can degrade performance. Most existing methods either fuse multimodal information directly or use rigid architectural separation for disentanglement, failing to adequately filter noise and model the complex interplay between modalities. To address these challenges, we propose a novel framework, the Multimodal Representation-disentangled Information Bottleneck (MRdIB). Concretely, we first employ a Multimodal Information Bottleneck to compress the input representations, effectively filtering out task-irrelevant noise while preserving rich semantic information. Then, we decompose the information based on its relationship with the recommendation target into unique, redundant, and synergistic components. We achieve this decomposition with a series of constraints: a unique information learning objective to preserve modality-unique signals, a redundant information learning objective to minimize overlap, and a synergistic information learning objective to capture emergent information. By optimizing these objectives, MRdIB guides a model to learn more powerful and disentangled representations. Extensive experiments on several competitive models and three benchmark datasets demonstrate the effectiveness and versatility of our MRdIB in enhancing multimodal recommendation.
title Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2509.20225