Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhixin, Zhang, Yiyuan, Ding, Xiaohan, Yue, Xiangyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913565896605696
author Zhang, Zhixin
Zhang, Yiyuan
Ding, Xiaohan
Yue, Xiangyu
author_facet Zhang, Zhixin
Zhang, Yiyuan
Ding, Xiaohan
Yue, Xiangyu
contents Search engines enable the retrieval of unknown information with texts. However, traditional methods fall short when it comes to understanding unfamiliar visual content, such as identifying an object that the model has never seen before. This challenge is particularly pronounced for large vision-language models (VLMs): if the model has not been exposed to the object depicted in an image, it struggles to generate reliable answers to the user's question regarding that image. Moreover, as new objects and events continuously emerge, frequently updating VLMs is impractical due to heavy computational burdens. To address this limitation, we propose Vision Search Assistant, a novel framework that facilitates collaboration between VLMs and web agents. This approach leverages VLMs' visual understanding capabilities and web agents' real-time information access to perform open-world Retrieval-Augmented Generation via the web. By integrating visual and textual representations through this collaboration, the model can provide informed responses even when the image is novel to the system. Extensive experiments conducted on both open-set and closed-set QA benchmarks demonstrate that the Vision Search Assistant significantly outperforms the other models and can be widely applied to existing VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21220
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
Zhang, Zhixin
Zhang, Yiyuan
Ding, Xiaohan
Yue, Xiangyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Machine Learning
Search engines enable the retrieval of unknown information with texts. However, traditional methods fall short when it comes to understanding unfamiliar visual content, such as identifying an object that the model has never seen before. This challenge is particularly pronounced for large vision-language models (VLMs): if the model has not been exposed to the object depicted in an image, it struggles to generate reliable answers to the user's question regarding that image. Moreover, as new objects and events continuously emerge, frequently updating VLMs is impractical due to heavy computational burdens. To address this limitation, we propose Vision Search Assistant, a novel framework that facilitates collaboration between VLMs and web agents. This approach leverages VLMs' visual understanding capabilities and web agents' real-time information access to perform open-world Retrieval-Augmented Generation via the web. By integrating visual and textual representations through this collaboration, the model can provide informed responses even when the image is novel to the system. Extensive experiments conducted on both open-set and closed-set QA benchmarks demonstrate that the Vision Search Assistant significantly outperforms the other models and can be widely applied to existing VLMs.
title Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2410.21220