Saved in:
Bibliographic Details
Main Authors: Zhang, Wenyu, Luo, Jie, Zhang, Xinming, Fang, Yuan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.15542
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929724927770624
author Zhang, Wenyu
Luo, Jie
Zhang, Xinming
Fang, Yuan
author_facet Zhang, Wenyu
Luo, Jie
Zhang, Xinming
Fang, Yuan
contents With the explosive growth of multimodal content online, pre-trained visual-language models have shown great potential for multimodal recommendation. However, while these models achieve decent performance when applied in a frozen manner, surprisingly, due to significant domain gaps (e.g., feature distribution discrepancy and task objective misalignment) between pre-training and personalized recommendation, adopting a joint training approach instead leads to performance worse than baseline. Existing approaches either rely on simple feature extraction or require computationally expensive full model fine-tuning, struggling to balance effectiveness and efficiency. To tackle these challenges, we propose \textbf{P}arameter-efficient \textbf{T}uning for \textbf{M}ultimodal \textbf{Rec}ommendation (\textbf{PTMRec}), a novel framework that bridges the domain gap between pre-trained models and recommendation systems through a knowledge-guided dual-stage parameter-efficient training strategy. This framework not only eliminates the need for costly additional pre-training but also flexibly accommodates various parameter-efficient tuning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2502_15542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging Domain Gaps between Pretrained Multimodal Models and Recommendations
Zhang, Wenyu
Luo, Jie
Zhang, Xinming
Fang, Yuan
Information Retrieval
Artificial Intelligence
With the explosive growth of multimodal content online, pre-trained visual-language models have shown great potential for multimodal recommendation. However, while these models achieve decent performance when applied in a frozen manner, surprisingly, due to significant domain gaps (e.g., feature distribution discrepancy and task objective misalignment) between pre-training and personalized recommendation, adopting a joint training approach instead leads to performance worse than baseline. Existing approaches either rely on simple feature extraction or require computationally expensive full model fine-tuning, struggling to balance effectiveness and efficiency. To tackle these challenges, we propose \textbf{P}arameter-efficient \textbf{T}uning for \textbf{M}ultimodal \textbf{Rec}ommendation (\textbf{PTMRec}), a novel framework that bridges the domain gap between pre-trained models and recommendation systems through a knowledge-guided dual-stage parameter-efficient training strategy. This framework not only eliminates the need for costly additional pre-training but also flexibly accommodates various parameter-efficient tuning methods.
title Bridging Domain Gaps between Pretrained Multimodal Models and Recommendations
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2502.15542