SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Tiansheng, Zhao, Yilun, Zhang, Canyu, Cohan, Arman, Zhao, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
LimRank: Less is More for Reasoning-Intensive Information Reranking
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
by: Zhang, Siyue, et al.
Published: (2025)
by: Zhang, Siyue, et al.
Published: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
by: Du, Mingxuan, et al.
Published: (2025)
by: Du, Mingxuan, et al.
Published: (2025)
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
by: Zhao, Xinran, et al.
Published: (2024)
by: Zhao, Xinran, et al.
Published: (2024)
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
by: Chen, Zijian, et al.
Published: (2025)
by: Chen, Zijian, et al.
Published: (2025)
Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement Learning
by: Chen, Yiqun, et al.
Published: (2025)
by: Chen, Yiqun, et al.
Published: (2025)
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
by: Lyu, Xinxi, et al.
Published: (2025)
by: Lyu, Xinxi, et al.
Published: (2025)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
by: Weller, Orion, et al.
Published: (2024)
by: Weller, Orion, et al.
Published: (2024)
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
by: Jin, Jiajie, et al.
Published: (2024)
by: Jin, Jiajie, et al.
Published: (2024)
ReFIT: Relevance Feedback from a Reranker during Inference
by: Reddy, Revanth Gangi, et al.
Published: (2023)
by: Reddy, Revanth Gangi, et al.
Published: (2023)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
by: Song, Tingyu, et al.
Published: (2026)
by: Song, Tingyu, et al.
Published: (2026)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
by: Cheng, Yiruo, et al.
Published: (2024)
by: Cheng, Yiruo, et al.
Published: (2024)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
by: Hu, Tiansheng, et al.
Published: (2025)
by: Hu, Tiansheng, et al.
Published: (2025)
CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment
by: Zhao, Zhengyun, et al.
Published: (2025)
by: Zhao, Zhengyun, et al.
Published: (2025)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
by: Ding, Hang, et al.
Published: (2025)
by: Ding, Hang, et al.
Published: (2025)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
by: Li, Haitao, et al.
Published: (2025)
by: Li, Haitao, et al.
Published: (2025)
Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA
by: Ou, Justice, et al.
Published: (2025)
by: Ou, Justice, et al.
Published: (2025)
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
by: Kim, Eric Y., et al.
Published: (2026)
by: Kim, Eric Y., et al.
Published: (2026)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
by: Chen, Jianlyu, et al.
Published: (2024)
by: Chen, Jianlyu, et al.
Published: (2024)
RAR-b: Reasoning as Retrieval Benchmark
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
by: Chen, Yiqun, et al.
Published: (2025)
by: Chen, Yiqun, et al.
Published: (2025)
Revela: Dense Retriever Learning via Language Modeling
by: Cai, Fengyu, et al.
Published: (2025)
by: Cai, Fengyu, et al.
Published: (2025)
PJB: A Reasoning-Aware Benchmark for Person-Job Retrieval
by: Wang, Guangzhi, et al.
Published: (2026)
by: Wang, Guangzhi, et al.
Published: (2026)
PosIR: Position-Aware Heterogeneous Information Retrieval Benchmark
by: Zeng, Ziyang, et al.
Published: (2026)
by: Zeng, Ziyang, et al.
Published: (2026)
$\texttt{MixGR}$: Enhancing Retriever Generalization for Scientific Domain through Complementary Granularity
by: Cai, Fengyu, et al.
Published: (2024)
by: Cai, Fengyu, et al.
Published: (2024)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
by: Jin, Rihui, et al.
Published: (2026)
by: Jin, Rihui, et al.
Published: (2026)
MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation
by: Chang, Chia-Yuan, et al.
Published: (2024)
by: Chang, Chia-Yuan, et al.
Published: (2024)
Towards Personalized Deep Research: Benchmarks and Evaluations
by: Liang, Yuan, et al.
Published: (2025)
by: Liang, Yuan, et al.
Published: (2025)
Generalizing Conversational Dense Retrieval via LLM-Cognition Data Augmentation
by: Chen, Haonan, et al.
Published: (2024)
by: Chen, Haonan, et al.
Published: (2024)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
by: Li, Xiangyang, et al.
Published: (2024)
by: Li, Xiangyang, et al.
Published: (2024)
Building Russian Benchmark for Evaluation of Information Retrieval Models
by: Kovalev, Grigory, et al.
Published: (2025)
by: Kovalev, Grigory, et al.
Published: (2025)
Similar Items
-
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
by: Zhao, Yilun, et al.
Published: (2026) -
LimRank: Less is More for Reasoning-Intensive Information Reranking
by: Song, Tingyu, et al.
Published: (2025) -
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval
by: Song, Tingyu, et al.
Published: (2025) -
Investigating Retrieval-Augmented Generation Systems on Unanswerable, Uncheatable, Realistic, Multi-hop Queries
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025) -
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
by: Zhang, Siyue, et al.
Published: (2025)