DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Siyan, Li, Deshpande, Darshan, Kannappan, Anand, Qian, Rebecca |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
TRAIL: Trace Reasoning and Agentic Issue Localization
by: Deshpande, Darshan, et al.
Published: (2025)
by: Deshpande, Darshan, et al.
Published: (2025)
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
by: Deshpande, Darshan, et al.
Published: (2025)
by: Deshpande, Darshan, et al.
Published: (2025)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
by: Deshpande, Darshan, et al.
Published: (2026)
by: Deshpande, Darshan, et al.
Published: (2026)
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
by: Deshpande, Darshan, et al.
Published: (2024)
by: Deshpande, Darshan, et al.
Published: (2024)
Lynx: An Open Source Hallucination Evaluation Model
by: Ravi, Selvan Sunitha, et al.
Published: (2024)
by: Ravi, Selvan Sunitha, et al.
Published: (2024)
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
by: Fujinuma, Yoshinari, et al.
Published: (2026)
by: Fujinuma, Yoshinari, et al.
Published: (2026)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023)
by: Vidgen, Bertie, et al.
Published: (2023)
Contextualizing Argument Quality Assessment with Relevant Knowledge
by: Deshpande, Darshan, et al.
Published: (2023)
by: Deshpande, Darshan, et al.
Published: (2023)
Robust Text Classification: Analyzing Prototype-Based Networks
by: Sourati, Zhivar, et al.
Published: (2023)
by: Sourati, Zhivar, et al.
Published: (2023)
RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents
by: Li, Mingchen, et al.
Published: (2026)
by: Li, Mingchen, et al.
Published: (2026)
GNOME: Generating Negotiations through Open-Domain Mapping of Exchanges
by: Deshpande, Darshan, et al.
Published: (2024)
by: Deshpande, Darshan, et al.
Published: (2024)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent
by: Li, Bingxuan, et al.
Published: (2026)
by: Li, Bingxuan, et al.
Published: (2026)
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
by: Li, Bobo, et al.
Published: (2025)
by: Li, Bobo, et al.
Published: (2025)
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
by: Wang, Siyan, et al.
Published: (2024)
by: Wang, Siyan, et al.
Published: (2024)
Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
by: Fofadiya, Darshan
Published: (2025)
by: Fofadiya, Darshan
Published: (2025)
SRR-Judge: Step-Level Rating and Refinement for Enhancing Search-Integrated Reasoning in Search Agents
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
by: Zhao, Siyan, et al.
Published: (2025)
by: Zhao, Siyan, et al.
Published: (2025)
InteractComp: Evaluating Search Agents With Ambiguous Queries
by: Deng, Mingyi, et al.
Published: (2025)
by: Deng, Mingyi, et al.
Published: (2025)
LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge
by: Zhou, Heng, et al.
Published: (2025)
by: Zhou, Heng, et al.
Published: (2025)
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
by: Mou, Xinyi, et al.
Published: (2024)
by: Mou, Xinyi, et al.
Published: (2024)
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions
by: Ye, Jingheng, et al.
Published: (2024)
by: Ye, Jingheng, et al.
Published: (2024)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
by: Gupta, Aarav, et al.
Published: (2026)
by: Gupta, Aarav, et al.
Published: (2026)
Grounded Language Agent for Product Search via Intelligent Web Interactions
by: Fereidouni, Moghis, et al.
Published: (2024)
by: Fereidouni, Moghis, et al.
Published: (2024)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
by: Qian, Hongjin, et al.
Published: (2025)
by: Qian, Hongjin, et al.
Published: (2025)
BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models
by: Chen, Jiangxi, et al.
Published: (2026)
by: Chen, Jiangxi, et al.
Published: (2026)
CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
by: Agarwal, Aradhye, et al.
Published: (2026)
by: Agarwal, Aradhye, et al.
Published: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
by: Li, Chance Jiajie, et al.
Published: (2025)
by: Li, Chance Jiajie, et al.
Published: (2025)
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
by: Kong, Quyu, et al.
Published: (2025)
by: Kong, Quyu, et al.
Published: (2025)
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
by: Kartha, Aaryaman, et al.
Published: (2025)
by: Kartha, Aaryaman, et al.
Published: (2025)
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
by: Patel, Maya, et al.
Published: (2024)
by: Patel, Maya, et al.
Published: (2024)
EvolveSearch: An Iterative Self-Evolving Search Agent
by: Zhang, Dingchu, et al.
Published: (2025)
by: Zhang, Dingchu, et al.
Published: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
by: Jing, Huihao, et al.
Published: (2025)
by: Jing, Huihao, et al.
Published: (2025)
A Study on Leveraging Search and Self-Feedback for Agent Reasoning
by: K, Karthikeyan, et al.
Published: (2025)
by: K, Karthikeyan, et al.
Published: (2025)
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
by: Anand, Ashwani, et al.
Published: (2026)
by: Anand, Ashwani, et al.
Published: (2026)
EDEN: Empathetic Dialogues for English learning
by: Siyan, Li, et al.
Published: (2024)
by: Siyan, Li, et al.
Published: (2024)
Similar Items
-
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
by: CH-Wang, Sky, et al.
Published: (2025) -
TRAIL: Trace Reasoning and Agentic Issue Localization
by: Deshpande, Darshan, et al.
Published: (2025) -
MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
by: Deshpande, Darshan, et al.
Published: (2025) -
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
by: Deshpande, Darshan, et al.
Published: (2026) -
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
by: Deshpande, Darshan, et al.
Published: (2024)