Understanding DeepResearch via Reports

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fan, Tianyu, Niu, Xinyao, Zheng, Yuxiang, Zhang, Fengji, Huang, Chengen, Chen, Bei, Lin, Junyang, Huang, Chao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916998314721280
author Fan, Tianyu
Niu, Xinyao
Zheng, Yuxiang
Zhang, Fengji
Huang, Chengen
Chen, Bei
Lin, Junyang
Huang, Chao
author_facet Fan, Tianyu
Niu, Xinyao
Zheng, Yuxiang
Zhang, Fengji
Huang, Chengen
Chen, Bei
Lin, Junyang
Huang, Chao
contents DeepResearch agents represent a transformative AI paradigm, conducting expert-level research through sophisticated reasoning and multi-tool integration. However, evaluating these systems remains critically challenging due to open-ended research scenarios and existing benchmarks that focus on isolated capabilities rather than holistic performance. Unlike traditional LLM tasks, DeepResearch systems must synthesize diverse sources, generate insights, and present coherent findings, which are capabilities that resist simple verification. To address this gap, we introduce DeepResearch-ReportEval, a comprehensive framework designed to assess DeepResearch systems through their most representative outputs: research reports. Our approach systematically measures three dimensions: quality, redundancy, and factuality, using an innovative LLM-as-a-Judge methodology achieving strong expert concordance. We contribute a standardized benchmark of 100 curated queries spanning 12 real-world categories, enabling systematic capability comparison. Our evaluation of four leading commercial systems reveals distinct design philosophies and performance trade-offs, establishing foundational insights as DeepResearch evolves from information assistants toward intelligent research partners. Source code and data are available at: https://github.com/HKUDS/DeepResearch-Eval.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding DeepResearch via Reports
Fan, Tianyu
Niu, Xinyao
Zheng, Yuxiang
Zhang, Fengji
Huang, Chengen
Chen, Bei
Lin, Junyang
Huang, Chao
Artificial Intelligence
DeepResearch agents represent a transformative AI paradigm, conducting expert-level research through sophisticated reasoning and multi-tool integration. However, evaluating these systems remains critically challenging due to open-ended research scenarios and existing benchmarks that focus on isolated capabilities rather than holistic performance. Unlike traditional LLM tasks, DeepResearch systems must synthesize diverse sources, generate insights, and present coherent findings, which are capabilities that resist simple verification. To address this gap, we introduce DeepResearch-ReportEval, a comprehensive framework designed to assess DeepResearch systems through their most representative outputs: research reports. Our approach systematically measures three dimensions: quality, redundancy, and factuality, using an innovative LLM-as-a-Judge methodology achieving strong expert concordance. We contribute a standardized benchmark of 100 curated queries spanning 12 real-world categories, enabling systematic capability comparison. Our evaluation of four leading commercial systems reveals distinct design philosophies and performance trade-offs, establishing foundational insights as DeepResearch evolves from information assistants toward intelligent research partners. Source code and data are available at: https://github.com/HKUDS/DeepResearch-Eval.
title Understanding DeepResearch via Reports
topic Artificial Intelligence
url https://arxiv.org/abs/2510.07861