Leveraging Large Language Models for Multimodal Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barbany, Oriol, Huang, Michael, Zhu, Xinliang, Dhua, Arnab
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909180276768768
author Barbany, Oriol
Huang, Michael
Zhu, Xinliang
Dhua, Arnab
author_facet Barbany, Oriol
Huang, Michael
Zhu, Xinliang
Dhua, Arnab
contents Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of the desired products, while text allows for easily incorporating search modifications. However, some existing multimodal search systems are unreliable and fail to address simple queries. The problem becomes harder with the large variability of natural language text queries, which may contain ambiguous, implicit, and irrelevant in-formation. Addressing these issues may require systems with enhanced matching capabilities, reasoning abilities, and context-aware query parsing and rewriting. This paper introduces a novel multimodal search model that achieves a new performance milestone on the Fashion200K dataset. Additionally, we propose a novel search interface integrating Large Language Models (LLMs) to facilitate natural language interaction. This interface routes queries to search systems while conversationally engaging with users and considering previous searches. When coupled with our multimodal search model, it heralds a new era of shopping assistants capable of offering human-like interaction and enhancing the overall search experience.
format Preprint
id arxiv_https___arxiv_org_abs_2404_15790
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Large Language Models for Multimodal Search
Barbany, Oriol
Huang, Michael
Zhu, Xinliang
Dhua, Arnab
Computer Vision and Pattern Recognition
Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of the desired products, while text allows for easily incorporating search modifications. However, some existing multimodal search systems are unreliable and fail to address simple queries. The problem becomes harder with the large variability of natural language text queries, which may contain ambiguous, implicit, and irrelevant in-formation. Addressing these issues may require systems with enhanced matching capabilities, reasoning abilities, and context-aware query parsing and rewriting. This paper introduces a novel multimodal search model that achieves a new performance milestone on the Fashion200K dataset. Additionally, we propose a novel search interface integrating Large Language Models (LLMs) to facilitate natural language interaction. This interface routes queries to search systems while conversationally engaging with users and considering previous searches. When coupled with our multimodal search model, it heralds a new era of shopping assistants capable of offering human-like interaction and enhancing the overall search experience.
title Leveraging Large Language Models for Multimodal Search
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.15790