ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Ruixiang, Jia, Jian, Li, Yan, Bai, Xuehan, Chen, Quan, Li, Han, Jiang, Peng, Li, Xirong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909657728024576
author Zhao, Ruixiang
Jia, Jian
Li, Yan
Bai, Xuehan
Chen, Quan
Li, Han
Jiang, Peng
Li, Xirong
author_facet Zhao, Ruixiang
Jia, Jian
Li, Yan
Bai, Xuehan
Chen, Quan
Li, Han
Jiang, Peng
Li, Xirong
contents E-commerce is increasingly multimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We propose ASR-enhanced Multimodal Product Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text, AMPere uses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness of AMPere in obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02978
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
Zhao, Ruixiang
Jia, Jian
Li, Yan
Bai, Xuehan
Chen, Quan
Li, Han
Jiang, Peng
Li, Xirong
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
E-commerce is increasingly multimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We propose ASR-enhanced Multimodal Product Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text, AMPere uses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness of AMPere in obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval.
title ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.02978