LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiayu, Ming, Yifei, Dulepet, Riya, Chen, Qinglin, Xu, Austin, Ke, Zixuan, Sala, Frederic, Albarghouthi, Aws, Xiong, Caiming, Joty, Shafiq
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914484305526784
author Wang, Jiayu
Ming, Yifei
Dulepet, Riya
Chen, Qinglin
Xu, Austin
Ke, Zixuan
Sala, Frederic
Albarghouthi, Aws
Xiong, Caiming
Joty, Shafiq
author_facet Wang, Jiayu
Ming, Yifei
Dulepet, Riya
Chen, Qinglin
Xu, Austin
Ke, Zixuan
Sala, Frederic
Albarghouthi, Aws
Xiong, Caiming
Joty, Shafiq
contents Deep research -- producing comprehensive, citation-grounded reports by searching and synthesizing information from hundreds of live web sources -- marks an important frontier for agentic systems. To rigorously evaluate this ability, four principles are essential: tasks should be (1) user-centric, reflecting realistic information needs, (2) dynamic, requiring up-to-date information beyond parametric knowledge, (3) unambiguous, ensuring consistent interpretation across users, and (4) multi-faceted and search-intensive, requiring search over numerous web sources and in-depth analysis. Existing benchmarks fall short of these principles, often focusing on narrow domains or posing ambiguous questions that hinder fair comparison. Guided by these principles, we introduce LiveResearchBench, a benchmark of 100 expert-curated tasks spanning daily life, enterprise, and academia, each requiring extensive, dynamic, real-time web search and synthesis. Built with over 1,500 hours of human labor, LiveResearchBench provides a rigorous basis for systematic evaluation. To evaluate citation-grounded long-form reports, we introduce DeepEval, a comprehensive suite covering both content- and report-level quality, including coverage, presentation, citation accuracy and association, consistency and depth of analysis. DeepEval integrates four complementary evaluation protocols, each designed to ensure stable assessment and high agreement with human judgments. Using LiveResearchBench and DeepEval, we conduct a comprehensive evaluation of 17 frontier deep research systems, including single-agent web search, single-agent deep research, and multi-agent systems. Our analysis reveals current strengths, recurring failure modes, and key system components needed to advance reliable, insightful deep research. Our code is available at: https://github.com/SalesforceAIResearch/LiveResearchBench.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
Wang, Jiayu
Ming, Yifei
Dulepet, Riya
Chen, Qinglin
Xu, Austin
Ke, Zixuan
Sala, Frederic
Albarghouthi, Aws
Xiong, Caiming
Joty, Shafiq
Artificial Intelligence
Deep research -- producing comprehensive, citation-grounded reports by searching and synthesizing information from hundreds of live web sources -- marks an important frontier for agentic systems. To rigorously evaluate this ability, four principles are essential: tasks should be (1) user-centric, reflecting realistic information needs, (2) dynamic, requiring up-to-date information beyond parametric knowledge, (3) unambiguous, ensuring consistent interpretation across users, and (4) multi-faceted and search-intensive, requiring search over numerous web sources and in-depth analysis. Existing benchmarks fall short of these principles, often focusing on narrow domains or posing ambiguous questions that hinder fair comparison. Guided by these principles, we introduce LiveResearchBench, a benchmark of 100 expert-curated tasks spanning daily life, enterprise, and academia, each requiring extensive, dynamic, real-time web search and synthesis. Built with over 1,500 hours of human labor, LiveResearchBench provides a rigorous basis for systematic evaluation. To evaluate citation-grounded long-form reports, we introduce DeepEval, a comprehensive suite covering both content- and report-level quality, including coverage, presentation, citation accuracy and association, consistency and depth of analysis. DeepEval integrates four complementary evaluation protocols, each designed to ensure stable assessment and high agreement with human judgments. Using LiveResearchBench and DeepEval, we conduct a comprehensive evaluation of 17 frontier deep research systems, including single-agent web search, single-agent deep research, and multi-agent systems. Our analysis reveals current strengths, recurring failure modes, and key system components needed to advance reliable, insightful deep research. Our code is available at: https://github.com/SalesforceAIResearch/LiveResearchBench.
title LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
topic Artificial Intelligence
url https://arxiv.org/abs/2510.14240