MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Baixuan, Wang, Weiqi, Shi, Haochen, Ding, Wenxuan, Jing, Huihao, Fang, Tianqing, Bai, Jiaxin, Liu, Xin, Yu, Changlong, Li, Zheng, Luo, Chen, Yin, Qingyu, Yin, Bing, Chen, Long, Song, Yangqiu
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913542563692544
author Xu, Baixuan
Wang, Weiqi
Shi, Haochen
Ding, Wenxuan
Jing, Huihao
Fang, Tianqing
Bai, Jiaxin
Liu, Xin
Yu, Changlong
Li, Zheng
Luo, Chen
Yin, Qingyu
Yin, Bing
Chen, Long
Song, Yangqiu
author_facet Xu, Baixuan
Wang, Weiqi
Shi, Haochen
Ding, Wenxuan
Jing, Huihao
Fang, Tianqing
Bai, Jiaxin
Liu, Xin
Yu, Changlong
Li, Zheng
Luo, Chen
Yin, Qingyu
Yin, Bing
Chen, Long
Song, Yangqiu
contents Improving user experience and providing personalized search results in E-commerce platforms heavily rely on understanding purchase intention. However, existing methods for acquiring large-scale intentions bank on distilling large language models with human annotation for verification. Such an approach tends to generate product-centric intentions, overlook valuable visual information from product images, and incurs high costs for scalability. To address these issues, we introduce MIND, a multimodal framework that allows Large Vision-Language Models (LVLMs) to infer purchase intentions from multimodal product metadata and prioritize human-centric ones. Using Amazon Review data, we apply MIND and create a multimodal intention knowledge base, which contains 1,264,441 million intentions derived from 126,142 co-buy shopping records across 107,215 products. Extensive human evaluations demonstrate the high plausibility and typicality of our obtained intentions and validate the effectiveness of our distillation framework and filtering mechanism. Additional experiments reveal that our obtained intentions significantly enhance large language models in two intention comprehension tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10701
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding
Xu, Baixuan
Wang, Weiqi
Shi, Haochen
Ding, Wenxuan
Jing, Huihao
Fang, Tianqing
Bai, Jiaxin
Liu, Xin
Yu, Changlong
Li, Zheng
Luo, Chen
Yin, Qingyu
Yin, Bing
Chen, Long
Song, Yangqiu
Computation and Language
Improving user experience and providing personalized search results in E-commerce platforms heavily rely on understanding purchase intention. However, existing methods for acquiring large-scale intentions bank on distilling large language models with human annotation for verification. Such an approach tends to generate product-centric intentions, overlook valuable visual information from product images, and incurs high costs for scalability. To address these issues, we introduce MIND, a multimodal framework that allows Large Vision-Language Models (LVLMs) to infer purchase intentions from multimodal product metadata and prioritize human-centric ones. Using Amazon Review data, we apply MIND and create a multimodal intention knowledge base, which contains 1,264,441 million intentions derived from 126,142 co-buy shopping records across 107,215 products. Extensive human evaluations demonstrate the high plausibility and typicality of our obtained intentions and validate the effectiveness of our distillation framework and filtering mechanism. Additional experiments reveal that our obtained intentions significantly enhance large language models in two intention comprehension tasks.
title MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding
topic Computation and Language
url https://arxiv.org/abs/2406.10701