NoteLLM-2: Multimodal Large Representation Models for Recommendation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Chao, Zhang, Haoxin, Wu, Shiwei, Wu, Di, Xu, Tong, Zhao, Xiangyu, Gao, Yan, Hu, Yao, Chen, Enhong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912195420356608
author Zhang, Chao
Zhang, Haoxin
Wu, Shiwei
Wu, Di
Xu, Tong
Zhao, Xiangyu
Gao, Yan
Hu, Yao
Chen, Enhong
author_facet Zhang, Chao
Zhang, Haoxin
Wu, Shiwei
Wu, Di
Xu, Tong
Zhao, Xiangyu
Gao, Yan
Hu, Yao
Chen, Enhong
contents Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularly for item-to-item (I2I) recommendations, remains underexplored. While leveraging existing Multimodal Large Language Models (MLLMs) for such tasks is promising, challenges arise due to their delayed release compared to corresponding LLMs and the inefficiency in representation tasks. To address these issues, we propose an end-to-end fine-tuning method that customizes the integration of any existing LLMs and vision encoders for efficient multimodal representation. Preliminary experiments revealed that fine-tuned LLMs often neglect image content. To counteract this, we propose NoteLLM-2, a novel framework that enhances visual information. Specifically, we propose two approaches: first, a prompt-based method that segregates visual and textual content, employing a multimodal In-Context Learning strategy to balance focus across modalities; second, a late fusion technique that directly integrates visual information into the final representations. Extensive experiments, both online and offline, demonstrate the effectiveness of our approach. Code is available at https://github.com/Applied-Machine-Learning-Lab/NoteLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16789
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NoteLLM-2: Multimodal Large Representation Models for Recommendation
Zhang, Chao
Zhang, Haoxin
Wu, Shiwei
Wu, Di
Xu, Tong
Zhao, Xiangyu
Gao, Yan
Hu, Yao
Chen, Enhong
Information Retrieval
Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularly for item-to-item (I2I) recommendations, remains underexplored. While leveraging existing Multimodal Large Language Models (MLLMs) for such tasks is promising, challenges arise due to their delayed release compared to corresponding LLMs and the inefficiency in representation tasks. To address these issues, we propose an end-to-end fine-tuning method that customizes the integration of any existing LLMs and vision encoders for efficient multimodal representation. Preliminary experiments revealed that fine-tuned LLMs often neglect image content. To counteract this, we propose NoteLLM-2, a novel framework that enhances visual information. Specifically, we propose two approaches: first, a prompt-based method that segregates visual and textual content, employing a multimodal In-Context Learning strategy to balance focus across modalities; second, a late fusion technique that directly integrates visual information into the final representations. Extensive experiments, both online and offline, demonstrate the effectiveness of our approach. Code is available at https://github.com/Applied-Machine-Learning-Lab/NoteLLM.
title NoteLLM-2: Multimodal Large Representation Models for Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2405.16789