DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Du, Mingxuan, Xu, Benfeng, Zhu, Chiwei, Wang, Xiaorui, Mao, Zhendong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916793392562176
author Du, Mingxuan
Xu, Benfeng
Zhu, Chiwei
Wang, Xiaorui
Mao, Zhendong
author_facet Du, Mingxuan
Xu, Benfeng
Zhu, Chiwei
Wang, Xiaorui
Mao, Zhendong
contents Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they transform vast amounts of online information into analyst-grade, citation-rich reports--compressing hours of manual desk research into minutes. However, a comprehensive benchmark for systematically evaluating the capabilities of these agents remains absent. To bridge this gap, we present DeepResearch Bench, a benchmark consisting of 100 PhD-level research tasks, each meticulously crafted by domain experts across 22 distinct fields. Evaluating DRAs is inherently complex and labor-intensive. We therefore propose two novel methodologies that achieve strong alignment with human judgment. The first is a reference-based method with adaptive criteria to assess the quality of generated research reports. The other framework is introduced to evaluate DRA's information retrieval and collection capabilities by assessing its effective citation count and overall citation accuracy. We have open-sourced DeepResearch Bench and key components of these frameworks at https://github.com/Ayanami0730/deep_research_bench to accelerate the development of practical LLM-based agents.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Du, Mingxuan
Xu, Benfeng
Zhu, Chiwei
Wang, Xiaorui
Mao, Zhendong
Computation and Language
Information Retrieval
Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they transform vast amounts of online information into analyst-grade, citation-rich reports--compressing hours of manual desk research into minutes. However, a comprehensive benchmark for systematically evaluating the capabilities of these agents remains absent. To bridge this gap, we present DeepResearch Bench, a benchmark consisting of 100 PhD-level research tasks, each meticulously crafted by domain experts across 22 distinct fields. Evaluating DRAs is inherently complex and labor-intensive. We therefore propose two novel methodologies that achieve strong alignment with human judgment. The first is a reference-based method with adaptive criteria to assess the quality of generated research reports. The other framework is introduced to evaluate DRA's information retrieval and collection capabilities by assessing its effective citation count and overall citation accuracy. We have open-sourced DeepResearch Bench and key components of these frameworks at https://github.com/Ayanami0730/deep_research_bench to accelerate the development of practical LLM-based agents.
title DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2506.11763