Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Attimonelli, Matteo, De Bellis, Alessandro, Gema, Aryo Pradipta, Saxena, Rohit, Sekoyan, Monica, Kwan, Wai-Chung, Pomo, Claudio, Suglia, Alessandro, Jannach, Dietmar, Di Noia, Tommaso, Minervini, Pasquale
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914571016470528
author Attimonelli, Matteo
De Bellis, Alessandro
Gema, Aryo Pradipta
Saxena, Rohit
Sekoyan, Monica
Kwan, Wai-Chung
Pomo, Claudio
Suglia, Alessandro
Jannach, Dietmar
Di Noia, Tommaso
Minervini, Pasquale
author_facet Attimonelli, Matteo
De Bellis, Alessandro
Gema, Aryo Pradipta
Saxena, Rohit
Sekoyan, Monica
Kwan, Wai-Chung
Pomo, Claudio
Suglia, Alessandro
Jannach, Dietmar
Di Noia, Tommaso
Minervini, Pasquale
contents Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR benchmarks is assumed to require multimodal composition, i.e., combining complementary information from reference image and textual modification. In this work, we show that this assumption does not always hold. Across four widely used CIR benchmarks and eleven Generalist Multimodal Embedding models, a large fraction of queries can be solved using a single modality (from 32.2% to 83.6%), revealing pervasive unimodal shortcuts. Thus, high CIR performance can arise from unimodal signals rather than true multimodal composition. To better understand this issue, we perform a two-stage audit. First, we identify shortcut-solvable queries through cross-model analysis. Second, we conduct human validation on 4,741 shortcut-free queries, of which only 1,689 are well-formed, with common issues including ambiguous edits and mismatched targets. Re-evaluating models on this validated subset reveals qualitatively different behaviour: queries can no longer be solved with a single modality, and successful retrieval requires combining both inputs. While accuracy decreases, reliance on multimodal information increases. Overall, current CIR benchmarks conflate shortcut-solvable, noisy, and genuinely compositional queries, leading to an overestimation of model capability in multimodal composition.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14787
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
Attimonelli, Matteo
De Bellis, Alessandro
Gema, Aryo Pradipta
Saxena, Rohit
Sekoyan, Monica
Kwan, Wai-Chung
Pomo, Claudio
Suglia, Alessandro
Jannach, Dietmar
Di Noia, Tommaso
Minervini, Pasquale
Computer Vision and Pattern Recognition
Computation and Language
Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR benchmarks is assumed to require multimodal composition, i.e., combining complementary information from reference image and textual modification. In this work, we show that this assumption does not always hold. Across four widely used CIR benchmarks and eleven Generalist Multimodal Embedding models, a large fraction of queries can be solved using a single modality (from 32.2% to 83.6%), revealing pervasive unimodal shortcuts. Thus, high CIR performance can arise from unimodal signals rather than true multimodal composition. To better understand this issue, we perform a two-stage audit. First, we identify shortcut-solvable queries through cross-model analysis. Second, we conduct human validation on 4,741 shortcut-free queries, of which only 1,689 are well-formed, with common issues including ambiguous edits and mismatched targets. Re-evaluating models on this validated subset reveals qualitatively different behaviour: queries can no longer be solved with a single modality, and successful retrieval requires combining both inputs. While accuracy decreases, reliance on multimodal information increases. Overall, current CIR benchmarks conflate shortcut-solvable, noisy, and genuinely compositional queries, leading to an overestimation of model capability in multimodal composition.
title Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2605.14787