VisualLens: Personalization through Task-Agnostic Visual History

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Wang Bill, Fu, Deqing, Sun, Kai, Lu, Yi, Lin, Zhaojiang, Moon, Seungwhan, Narang, Kanika, Canim, Mustafa, Liu, Yue, Kumar, Anuj, Dong, Xin Luna
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917023042240512
author Zhu, Wang Bill
Fu, Deqing
Sun, Kai
Lu, Yi
Lin, Zhaojiang
Moon, Seungwhan
Narang, Kanika
Canim, Mustafa
Liu, Yue
Kumar, Anuj
Dong, Xin Luna
author_facet Zhu, Wang Bill
Fu, Deqing
Sun, Kai
Lu, Yi
Lin, Zhaojiang
Moon, Seungwhan
Narang, Kanika
Canim, Mustafa
Liu, Yue
Kumar, Anuj
Dong, Xin Luna
contents Existing recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals. However, item-based histories are not always accessible, and are not generalizable for multimodal recommendation. We hypothesize that a user's visual history -- comprising images from daily life -- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization. To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history. VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation. We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10% on Hit@3, and outperforms GPT-4o by 2-5%. Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16034
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VisualLens: Personalization through Task-Agnostic Visual History
Zhu, Wang Bill
Fu, Deqing
Sun, Kai
Lu, Yi
Lin, Zhaojiang
Moon, Seungwhan
Narang, Kanika
Canim, Mustafa
Liu, Yue
Kumar, Anuj
Dong, Xin Luna
Computer Vision and Pattern Recognition
Existing recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals. However, item-based histories are not always accessible, and are not generalizable for multimodal recommendation. We hypothesize that a user's visual history -- comprising images from daily life -- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization. To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history. VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation. We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10% on Hit@3, and outperforms GPT-4o by 2-5%. Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories.
title VisualLens: Personalization through Task-Agnostic Visual History
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16034