ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Pengfei, Zhou, Jingbo, Xu, Tong, Xia, Yuan, Xu, Linli, Chen, Enhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916651699535872
author Luo, Pengfei
Zhou, Jingbo
Xu, Tong
Xia, Yuan
Xu, Linli
Chen, Enhong
author_facet Luo, Pengfei
Zhou, Jingbo
Xu, Tong
Xia, Yuan
Xu, Linli
Chen, Enhong
contents With the proliferation of images in online content, language-guided image retrieval (LGIR) has emerged as a research hotspot over the past decade, encompassing a variety of subtasks with diverse input forms. While the development of large multimodal models (LMMs) has significantly facilitated these tasks, existing approaches often address them in isolation, requiring the construction of separate systems for each task. This not only increases system complexity and maintenance costs, but also exacerbates challenges stemming from language ambiguity and complex image content, making it difficult for retrieval systems to provide accurate and reliable results. To this end, we propose ImageScope, a training-free, three-stage framework that leverages collective reasoning to unify LGIR tasks. The key insight behind the unification lies in the compositional nature of language, which transforms diverse LGIR tasks into a generalized text-to-image retrieval process, along with the reasoning of LMMs serving as a universal verification to refine the results. To be specific, in the first stage, we improve the robustness of the framework by synthesizing search intents across varying levels of semantic granularity using chain-of-thought (CoT) reasoning. In the second and third stages, we then reflect on retrieval results by verifying predicate propositions locally, and performing pairwise evaluations globally. Experiments conducted on six LGIR datasets demonstrate that ImageScope outperforms competitive baselines. Comprehensive evaluations and ablation studies further confirm the effectiveness of our design.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10166
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning
Luo, Pengfei
Zhou, Jingbo
Xu, Tong
Xia, Yuan
Xu, Linli
Chen, Enhong
Information Retrieval
Artificial Intelligence
Multimedia
With the proliferation of images in online content, language-guided image retrieval (LGIR) has emerged as a research hotspot over the past decade, encompassing a variety of subtasks with diverse input forms. While the development of large multimodal models (LMMs) has significantly facilitated these tasks, existing approaches often address them in isolation, requiring the construction of separate systems for each task. This not only increases system complexity and maintenance costs, but also exacerbates challenges stemming from language ambiguity and complex image content, making it difficult for retrieval systems to provide accurate and reliable results. To this end, we propose ImageScope, a training-free, three-stage framework that leverages collective reasoning to unify LGIR tasks. The key insight behind the unification lies in the compositional nature of language, which transforms diverse LGIR tasks into a generalized text-to-image retrieval process, along with the reasoning of LMMs serving as a universal verification to refine the results. To be specific, in the first stage, we improve the robustness of the framework by synthesizing search intents across varying levels of semantic granularity using chain-of-thought (CoT) reasoning. In the second and third stages, we then reflect on retrieval results by verifying predicate propositions locally, and performing pairwise evaluations globally. Experiments conducted on six LGIR datasets demonstrate that ImageScope outperforms competitive baselines. Comprehensive evaluations and ablation studies further confirm the effectiveness of our design.
title ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning
topic Information Retrieval
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2503.10166