DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Siyan, Li, Deshpande, Darshan, Kannappan, Anand, Qian, Rebecca
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917236834304000
author Siyan, Li
Deshpande, Darshan
Kannappan, Anand
Qian, Rebecca
author_facet Siyan, Li
Deshpande, Darshan
Kannappan, Anand
Qian, Rebecca
contents When recalling information in conversation, people often arrive at the recollection after multiple turns. However, existing benchmarks for evaluating agent capabilities in such tip-of-the-tongue search processes are restricted to single-turn settings. To more realistically simulate tip-of-the-tongue search, we introduce Dual-agent based Evaluation Through Obscure Under-specified Retrieval (DETOUR), a dual-agent evaluation benchmark containing 1,011 prompts. The benchmark design involves a Primary Agent, which is the subject of evaluation, tasked with identifying the recollected entity through querying a Memory Agent that is held consistent across evaluations. Our results indicate that current state-of-the-art models still struggle with our benchmark, only achieving 36% accuracy when evaluated on all modalities (text, image, audio, and video), highlighting the importance of enhancing capabilities in underspecified scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00352
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
Siyan, Li
Deshpande, Darshan
Kannappan, Anand
Qian, Rebecca
Computation and Language
When recalling information in conversation, people often arrive at the recollection after multiple turns. However, existing benchmarks for evaluating agent capabilities in such tip-of-the-tongue search processes are restricted to single-turn settings. To more realistically simulate tip-of-the-tongue search, we introduce Dual-agent based Evaluation Through Obscure Under-specified Retrieval (DETOUR), a dual-agent evaluation benchmark containing 1,011 prompts. The benchmark design involves a Primary Agent, which is the subject of evaluation, tasked with identifying the recollected entity through querying a Memory Agent that is held consistent across evaluations. Our results indicate that current state-of-the-art models still struggle with our benchmark, only achieving 36% accuracy when evaluated on all modalities (text, image, audio, and video), highlighting the importance of enhancing capabilities in underspecified scenarios.
title DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
topic Computation and Language
url https://arxiv.org/abs/2602.00352