IntRec: Intent-based Retrieval with Contrastive Refinement

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shamsolmoali, Pourya, Zareapoor, Masoumeh, Granger, Eric, Lu, Yue
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914339193094144
author Shamsolmoali, Pourya
Zareapoor, Masoumeh
Granger, Eric
Lu, Yue
author_facet Shamsolmoali, Pourya
Zareapoor, Masoumeh
Granger, Eric
Lu, Yue
contents Retrieving user-specified objects from complex scenes remains a challenging task, especially when queries are ambiguous or involve multiple similar objects. Existing open-vocabulary detectors operate in a one-shot manner, lacking the ability to refine predictions based on user feedback. To address this, we propose IntRec, an interactive object retrieval framework that refines predictions based on user feedback. At its core is an Intent State (IS) that maintains dual memory sets for positive anchors (confirmed cues) and negative constraints (rejected hypotheses). A contrastive alignment function ranks candidate objects by maximizing similarity to positive cues while penalizing rejected ones, enabling fine-grained disambiguation in cluttered scenes. Our interactive framework provides substantial improvements in retrieval accuracy without additional supervision. On LVIS, IntRec achieves 35.4 AP, outperforming OVMR, CoDet, and CAKE by +2.3, +3.7, and +0.5, respectively. On the challenging LVIS-Ambiguous benchmark, it improves performance by +7.9 AP over its one-shot baseline after a single corrective feedback, with less than 30 ms of added latency per interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17639
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IntRec: Intent-based Retrieval with Contrastive Refinement
Shamsolmoali, Pourya
Zareapoor, Masoumeh
Granger, Eric
Lu, Yue
Computer Vision and Pattern Recognition
Retrieving user-specified objects from complex scenes remains a challenging task, especially when queries are ambiguous or involve multiple similar objects. Existing open-vocabulary detectors operate in a one-shot manner, lacking the ability to refine predictions based on user feedback. To address this, we propose IntRec, an interactive object retrieval framework that refines predictions based on user feedback. At its core is an Intent State (IS) that maintains dual memory sets for positive anchors (confirmed cues) and negative constraints (rejected hypotheses). A contrastive alignment function ranks candidate objects by maximizing similarity to positive cues while penalizing rejected ones, enabling fine-grained disambiguation in cluttered scenes. Our interactive framework provides substantial improvements in retrieval accuracy without additional supervision. On LVIS, IntRec achieves 35.4 AP, outperforming OVMR, CoDet, and CAKE by +2.3, +3.7, and +0.5, respectively. On the challenging LVIS-Ambiguous benchmark, it improves performance by +7.9 AP over its one-shot baseline after a single corrective feedback, with less than 30 ms of added latency per interaction.
title IntRec: Intent-based Retrieval with Contrastive Refinement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.17639