ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909657728024576 |
|---|---|
| author | Zhao, Ruixiang Jia, Jian Li, Yan Bai, Xuehan Chen, Quan Li, Han Jiang, Peng Li, Xirong |
| author_facet | Zhao, Ruixiang Jia, Jian Li, Yan Bai, Xuehan Chen, Quan Li, Han Jiang, Peng Li, Xirong |
| contents | E-commerce is increasingly multimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We propose ASR-enhanced Multimodal Product Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text, AMPere uses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness of AMPere in obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_02978 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval Zhao, Ruixiang Jia, Jian Li, Yan Bai, Xuehan Chen, Quan Li, Han Jiang, Peng Li, Xirong Multimedia Artificial Intelligence Computer Vision and Pattern Recognition E-commerce is increasingly multimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We propose ASR-enhanced Multimodal Product Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text, AMPere uses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness of AMPere in obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval. |
| title | ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval |
| topic | Multimedia Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2408.02978 |