A Metric for MLLM Alignment in Large-scale Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yubin, Huang, Yanhua, Xu, Haiming, Qi, Mingliang, Wang, Chang, Jin, Jiarui, Ren, Xiangyuan, Wang, Xiaodan, Xu, Ruiwen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915433231155200
author Zhang, Yubin
Huang, Yanhua
Xu, Haiming
Qi, Mingliang
Wang, Chang
Jin, Jiarui
Ren, Xiangyuan
Wang, Xiaodan
Xu, Ruiwen
author_facet Zhang, Yubin
Huang, Yanhua
Xu, Haiming
Qi, Mingliang
Wang, Chang
Jin, Jiarui
Ren, Xiangyuan
Wang, Xiaodan
Xu, Ruiwen
contents Multimodal recommendation has emerged as a critical technique in modern recommender systems, leveraging content representations from advanced multimodal large language models (MLLMs). To ensure these representations are well-adapted, alignment with the recommender system is essential. However, evaluating the alignment of MLLMs for recommendation presents significant challenges due to three key issues: (1) static benchmarks are inaccurate because of the dynamism in real-world applications, (2) evaluations with online system, while accurate, are prohibitively expensive at scale, and (3) conventional metrics fail to provide actionable insights when learned representations underperform. To address these challenges, we propose the Leakage Impact Score (LIS), a novel metric for multimodal recommendation. Rather than directly assessing MLLMs, LIS efficiently measures the upper bound of preference data. We also share practical insights on deploying MLLMs with LIS in real-world scenarios. Online A/B tests on both Content Feed and Display Ads of Xiaohongshu's Explore Feed production demonstrate the effectiveness of our proposed method, showing significant improvements in user spent time and advertiser value.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Metric for MLLM Alignment in Large-scale Recommendation
Zhang, Yubin
Huang, Yanhua
Xu, Haiming
Qi, Mingliang
Wang, Chang
Jin, Jiarui
Ren, Xiangyuan
Wang, Xiaodan
Xu, Ruiwen
Information Retrieval
Machine Learning
Multimodal recommendation has emerged as a critical technique in modern recommender systems, leveraging content representations from advanced multimodal large language models (MLLMs). To ensure these representations are well-adapted, alignment with the recommender system is essential. However, evaluating the alignment of MLLMs for recommendation presents significant challenges due to three key issues: (1) static benchmarks are inaccurate because of the dynamism in real-world applications, (2) evaluations with online system, while accurate, are prohibitively expensive at scale, and (3) conventional metrics fail to provide actionable insights when learned representations underperform. To address these challenges, we propose the Leakage Impact Score (LIS), a novel metric for multimodal recommendation. Rather than directly assessing MLLMs, LIS efficiently measures the upper bound of preference data. We also share practical insights on deploying MLLMs with LIS in real-world scenarios. Online A/B tests on both Content Feed and Display Ads of Xiaohongshu's Explore Feed production demonstrate the effectiveness of our proposed method, showing significant improvements in user spent time and advertiser value.
title A Metric for MLLM Alignment in Large-scale Recommendation
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2508.04963