CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Takeda, Tomohisa, Lin, Yu-Chieh, Nozawa, Yuji, Ng, Youyang, Torii, Osamu, Matsui, Yusuke
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911718616072192
author Takeda, Tomohisa
Lin, Yu-Chieh
Nozawa, Yuji
Ng, Youyang
Torii, Osamu
Matsui, Yusuke
author_facet Takeda, Tomohisa
Lin, Yu-Chieh
Nozawa, Yuji
Ng, Youyang
Torii, Osamu
Matsui, Yusuke
contents Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query at each turn progressively approaches the target image. Data are generated via a CIReVL-based retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multi-turn sessions across nine subsets, substantially exceeding Multi-turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, high-quality benchmark to facilitate future research on multi-turn CIR. The dataset and code are publicly available at https://huggingface.co/datasets/tk1441/CIRCLED and https://github.com/mti-lab/circled.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26734
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
Takeda, Tomohisa
Lin, Yu-Chieh
Nozawa, Yuji
Ng, Youyang
Torii, Osamu
Matsui, Yusuke
Computer Vision and Pattern Recognition
Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query at each turn progressively approaches the target image. Data are generated via a CIReVL-based retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multi-turn sessions across nine subsets, substantially exceeding Multi-turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, high-quality benchmark to facilitate future research on multi-turn CIR. The dataset and code are publicly available at https://huggingface.co/datasets/tk1441/CIRCLED and https://github.com/mti-lab/circled.
title CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26734