Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Tingyu, Zhang, Yanzhao, Li, Mingxin, Guo, Zhuoning, Long, Dingkun, Xie, Pengjun, Zhang, Siyue, Zhao, Yilun, Wu, Shu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915747840655360
author Song, Tingyu
Zhang, Yanzhao
Li, Mingxin
Guo, Zhuoning
Long, Dingkun
Xie, Pengjun
Zhang, Siyue
Zhao, Yilun
Wu, Shu
author_facet Song, Tingyu
Zhang, Yanzhao
Li, Mingxin
Guo, Zhuoning
Long, Dingkun
Xie, Pengjun
Zhang, Siyue
Zhao, Yilun
Wu, Shu
contents Composed Image Retrieval (CIR) is a pivotal and complex task in multimodal understanding. Current CIR benchmarks typically feature limited query categories and fail to capture the diverse requirements of real-world scenarios. To bridge this evaluation gap, we leverage image editing to achieve precise control over modification types and content, enabling a pipeline for synthesizing queries across a broad spectrum of categories. Using this pipeline, we construct EDIR, a novel fine-grained CIR benchmark. EDIR encompasses 5,000 high-quality queries structured across five main categories and fifteen subcategories. Our comprehensive evaluation of 13 multimodal embedding models reveals a significant capability gap; even state-of-the-art models (e.g., RzenEmbed and GME) struggle to perform consistently across all subcategories, highlighting the rigorous nature of our benchmark. Through comparative analysis, we further uncover inherent limitations in existing benchmarks, such as modality biases and insufficient categorical coverage. Furthermore, an in-domain training experiment demonstrates the feasibility of our benchmark. This experiment clarifies the task challenges by distinguishing between categories that are solvable with targeted data and those that expose intrinsic limitations of current model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16125
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
Song, Tingyu
Zhang, Yanzhao
Li, Mingxin
Guo, Zhuoning
Long, Dingkun
Xie, Pengjun
Zhang, Siyue
Zhao, Yilun
Wu, Shu
Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Composed Image Retrieval (CIR) is a pivotal and complex task in multimodal understanding. Current CIR benchmarks typically feature limited query categories and fail to capture the diverse requirements of real-world scenarios. To bridge this evaluation gap, we leverage image editing to achieve precise control over modification types and content, enabling a pipeline for synthesizing queries across a broad spectrum of categories. Using this pipeline, we construct EDIR, a novel fine-grained CIR benchmark. EDIR encompasses 5,000 high-quality queries structured across five main categories and fifteen subcategories. Our comprehensive evaluation of 13 multimodal embedding models reveals a significant capability gap; even state-of-the-art models (e.g., RzenEmbed and GME) struggle to perform consistently across all subcategories, highlighting the rigorous nature of our benchmark. Through comparative analysis, we further uncover inherent limitations in existing benchmarks, such as modality biases and insufficient categorical coverage. Furthermore, an in-domain training experiment demonstrates the feasibility of our benchmark. This experiment clarifies the task challenges by distinguishing between categories that are solvable with targeted data and those that expose intrinsic limitations of current model architectures.
title Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
topic Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2601.16125