QARM: Quantitative Alignment Multi-Modal Recommendation at Kuaishou

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Xinchen, Cao, Jiangxia, Sun, Tianyu, Yu, Jinkai, Huang, Rui, Yuan, Wei, Lin, Hezheng, Zheng, Yichen, Wang, Shiyao, Hu, Qigen, Qiu, Changqing, Zhang, Jiaqi, Zhang, Xu, Yan, Zhiheng, Zhang, Jingming, Zhang, Simin, Wen, Mingxing, Liu, Zhaojie, Gai, Kun, Zhou, Guorui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912123436662784
author Luo, Xinchen
Cao, Jiangxia
Sun, Tianyu
Yu, Jinkai
Huang, Rui
Yuan, Wei
Lin, Hezheng
Zheng, Yichen
Wang, Shiyao
Hu, Qigen
Qiu, Changqing
Zhang, Jiaqi
Zhang, Xu
Yan, Zhiheng
Zhang, Jingming
Zhang, Simin
Wen, Mingxing
Liu, Zhaojie
Gai, Kun
Zhou, Guorui
author_facet Luo, Xinchen
Cao, Jiangxia
Sun, Tianyu
Yu, Jinkai
Huang, Rui
Yuan, Wei
Lin, Hezheng
Zheng, Yichen
Wang, Shiyao
Hu, Qigen
Qiu, Changqing
Zhang, Jiaqi
Zhang, Xu
Yan, Zhiheng
Zhang, Jingming
Zhang, Simin
Wen, Mingxing
Liu, Zhaojie
Gai, Kun
Zhou, Guorui
contents In recent years, with the significant evolution of multi-modal large models, many recommender researchers realized the potential of multi-modal information for user interest modeling. In industry, a wide-used modeling architecture is a cascading paradigm: (1) first pre-training a multi-modal model to provide omnipotent representations for downstream services; (2) The downstream recommendation model takes the multi-modal representation as additional input to fit real user-item behaviours. Although such paradigm achieves remarkable improvements, however, there still exist two problems that limit model performance: (1) Representation Unmatching: The pre-trained multi-modal model is always supervised by the classic NLP/CV tasks, while the recommendation models are supervised by real user-item interaction. As a result, the two fundamentally different tasks' goals were relatively separate, and there was a lack of consistent objective on their representations; (2) Representation Unlearning: The generated multi-modal representations are always stored in cache store and serve as extra fixed input of recommendation model, thus could not be updated by recommendation model gradient, further unfriendly for downstream training. Inspired by the two difficulties challenges in downstream tasks usage, we introduce a quantitative multi-modal framework to customize the specialized and trainable multi-modal information for different downstream models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QARM: Quantitative Alignment Multi-Modal Recommendation at Kuaishou
Luo, Xinchen
Cao, Jiangxia
Sun, Tianyu
Yu, Jinkai
Huang, Rui
Yuan, Wei
Lin, Hezheng
Zheng, Yichen
Wang, Shiyao
Hu, Qigen
Qiu, Changqing
Zhang, Jiaqi
Zhang, Xu
Yan, Zhiheng
Zhang, Jingming
Zhang, Simin
Wen, Mingxing
Liu, Zhaojie
Gai, Kun
Zhou, Guorui
Information Retrieval
Artificial Intelligence
N/A
In recent years, with the significant evolution of multi-modal large models, many recommender researchers realized the potential of multi-modal information for user interest modeling. In industry, a wide-used modeling architecture is a cascading paradigm: (1) first pre-training a multi-modal model to provide omnipotent representations for downstream services; (2) The downstream recommendation model takes the multi-modal representation as additional input to fit real user-item behaviours. Although such paradigm achieves remarkable improvements, however, there still exist two problems that limit model performance: (1) Representation Unmatching: The pre-trained multi-modal model is always supervised by the classic NLP/CV tasks, while the recommendation models are supervised by real user-item interaction. As a result, the two fundamentally different tasks' goals were relatively separate, and there was a lack of consistent objective on their representations; (2) Representation Unlearning: The generated multi-modal representations are always stored in cache store and serve as extra fixed input of recommendation model, thus could not be updated by recommendation model gradient, further unfriendly for downstream training. Inspired by the two difficulties challenges in downstream tasks usage, we introduce a quantitative multi-modal framework to customize the specialized and trainable multi-modal information for different downstream models.
title QARM: Quantitative Alignment Multi-Modal Recommendation at Kuaishou
topic Information Retrieval
Artificial Intelligence
N/A
url https://arxiv.org/abs/2411.11739