InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xi, Yunjia, Lin, Jianghao, Zhu, Menghui, Xiao, Yongzhao, Ou, Zhuoying, Liu, Jiaqi, Wan, Tong, Chen, Bo, Liu, Weiwen, Wang, Yasheng, Tang, Ruiming, Zhang, Weinan, Yu, Yong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916754097176576
author Xi, Yunjia
Lin, Jianghao
Zhu, Menghui
Xiao, Yongzhao
Ou, Zhuoying
Liu, Jiaqi
Wan, Tong
Chen, Bo
Liu, Weiwen
Wang, Yasheng
Tang, Ruiming
Zhang, Weinan
Yu, Yong
author_facet Xi, Yunjia
Lin, Jianghao
Zhu, Menghui
Xiao, Yongzhao
Ou, Zhuoying
Liu, Jiaqi
Wan, Tong
Chen, Bo
Liu, Weiwen
Wang, Yasheng
Tang, Ruiming
Zhang, Weinan
Yu, Yong
contents Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by grounding responses with retrieved information. As an emerging paradigm, Agentic RAG further enhances this process by introducing autonomous LLM agents into the information seeking process. However, existing benchmarks fall short in evaluating such systems, as they are confined to a static retrieval environment with a fixed, limited corpus} and simple queries that fail to elicit agentic behavior. Moreover, their evaluation protocols assess information seeking effectiveness by pre-defined gold sets of documents, making them unsuitable for the open-ended and dynamic nature of real-world web environments. To bridge this gap, we present InfoDeepSeek, a new benchmark with challenging questions designed for assessing agentic information seeking in real-world, dynamic web environments. We propose a systematic methodology for constructing challenging queries satisfying the criteria of determinacy, difficulty, and diversity. Based on this, we develop the first evaluation framework tailored to dynamic agentic information seeking, including fine-grained metrics about the accuracy, utility, and compactness of information seeking outcomes. Through extensive experiments across LLMs, search engines, and question types, InfoDeepSeek reveals nuanced agent behaviors and offers actionable insights for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15872
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
Xi, Yunjia
Lin, Jianghao
Zhu, Menghui
Xiao, Yongzhao
Ou, Zhuoying
Liu, Jiaqi
Wan, Tong
Chen, Bo
Liu, Weiwen
Wang, Yasheng
Tang, Ruiming
Zhang, Weinan
Yu, Yong
Information Retrieval
Computation and Language
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by grounding responses with retrieved information. As an emerging paradigm, Agentic RAG further enhances this process by introducing autonomous LLM agents into the information seeking process. However, existing benchmarks fall short in evaluating such systems, as they are confined to a static retrieval environment with a fixed, limited corpus} and simple queries that fail to elicit agentic behavior. Moreover, their evaluation protocols assess information seeking effectiveness by pre-defined gold sets of documents, making them unsuitable for the open-ended and dynamic nature of real-world web environments. To bridge this gap, we present InfoDeepSeek, a new benchmark with challenging questions designed for assessing agentic information seeking in real-world, dynamic web environments. We propose a systematic methodology for constructing challenging queries satisfying the criteria of determinacy, difficulty, and diversity. Based on this, we develop the first evaluation framework tailored to dynamic agentic information seeking, including fine-grained metrics about the accuracy, utility, and compactness of information seeking outcomes. Through extensive experiments across LLMs, search engines, and question types, InfoDeepSeek reveals nuanced agent behaviors and offers actionable insights for future research.
title InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2505.15872