Serendipitous Recommendation with Multimodal LLM

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Haoting, Wang, Jianling, Li, Hao, Yi, Fangjun, Fu, Mengyu, Zhang, Youwei, Liu, Yifan, Liu, Liang, Chen, Minmin, Chi, Ed H., Hong, Lichan, Lu, Haokai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911167330385920
author Wang, Haoting
Wang, Jianling
Li, Hao
Yi, Fangjun
Fu, Mengyu
Zhang, Youwei
Liu, Yifan
Liu, Liang
Chen, Minmin
Chi, Ed H.
Hong, Lichan
Lu, Haokai
author_facet Wang, Haoting
Wang, Jianling
Li, Hao
Yi, Fangjun
Fu, Mengyu
Zhang, Youwei
Liu, Yifan
Liu, Liang
Chen, Minmin
Chi, Ed H.
Hong, Lichan
Lu, Haokai
contents Conventional recommendation systems succeed in identifying relevant content but often fail to provide users with surprising or novel items. Multimodal Large Language Models (MLLMs) possess the world knowledge and multimodal understanding needed for serendipity, but their integration into billion-item-scale platforms presents significant challenges. In this paper, we propose a novel hierarchical framework where fine-tuned MLLMs provide high-level guidance to conventional recommendation models, steering them towards more serendipitous suggestions. This approach leverages MLLM strengths in understanding multimodal content and user interests while retaining the efficiency of traditional models for item-level recommendation. This mitigates the complexity of applying MLLMs directly to vast action spaces. We also demonstrate a chain-of-thought strategy enabling MLLMs to discover novel user interests by first understanding video content and then identifying relevant yet unexplored interest clusters. Through live experiments within a commercial short-form video platform serving billions of users, we show that our MLLM-powered approach significantly improves both recommendation serendipity and user satisfaction.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Serendipitous Recommendation with Multimodal LLM
Wang, Haoting
Wang, Jianling
Li, Hao
Yi, Fangjun
Fu, Mengyu
Zhang, Youwei
Liu, Yifan
Liu, Liang
Chen, Minmin
Chi, Ed H.
Hong, Lichan
Lu, Haokai
Information Retrieval
Conventional recommendation systems succeed in identifying relevant content but often fail to provide users with surprising or novel items. Multimodal Large Language Models (MLLMs) possess the world knowledge and multimodal understanding needed for serendipity, but their integration into billion-item-scale platforms presents significant challenges. In this paper, we propose a novel hierarchical framework where fine-tuned MLLMs provide high-level guidance to conventional recommendation models, steering them towards more serendipitous suggestions. This approach leverages MLLM strengths in understanding multimodal content and user interests while retaining the efficiency of traditional models for item-level recommendation. This mitigates the complexity of applying MLLMs directly to vast action spaces. We also demonstrate a chain-of-thought strategy enabling MLLMs to discover novel user interests by first understanding video content and then identifying relevant yet unexplored interest clusters. Through live experiments within a commercial short-form video platform serving billions of users, we show that our MLLM-powered approach significantly improves both recommendation serendipity and user satisfaction.
title Serendipitous Recommendation with Multimodal LLM
topic Information Retrieval
url https://arxiv.org/abs/2506.08283