CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911718616072192 |
|---|---|
| author | Takeda, Tomohisa Lin, Yu-Chieh Nozawa, Yuji Ng, Youyang Torii, Osamu Matsui, Yusuke |
| author_facet | Takeda, Tomohisa Lin, Yu-Chieh Nozawa, Yuji Ng, Youyang Torii, Osamu Matsui, Yusuke |
| contents | Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query at each turn progressively approaches the target image. Data are generated via a CIReVL-based retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multi-turn sessions across nine subsets, substantially exceeding Multi-turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, high-quality benchmark to facilitate future research on multi-turn CIR. The dataset and code are publicly available at https://huggingface.co/datasets/tk1441/CIRCLED and https://github.com/mti-lab/circled. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_26734 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains Takeda, Tomohisa Lin, Yu-Chieh Nozawa, Yuji Ng, Youyang Torii, Osamu Matsui, Yusuke Computer Vision and Pattern Recognition Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In CIRCLED, the query at each turn progressively approaches the target image. Data are generated via a CIReVL-based retrieval pipeline and curated with multiple filters on retrieval success, turn length, consistency, and information redundancy to ensure quality. In total, we collect 22,608 multi-turn sessions across nine subsets, substantially exceeding Multi-turn FashionIQ (11,505 sessions) in both scale and generality. We further apply multiple baseline methods and quantitatively assess retrieval accuracy on CIRCLED. Our work provides a practical, high-quality benchmark to facilitate future research on multi-turn CIR. The dataset and code are publicly available at https://huggingface.co/datasets/tk1441/CIRCLED and https://github.com/mti-lab/circled. |
| title | CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.26734 |