Guardado en:
Detalles Bibliográficos
Autores principales: Wan, Haiyuan, Yang, Chen, Yu, Junchi, Tu, Meiqi, Lu, Jiaxuan, Yu, Di, Cao, Jianbao, Gao, Ben, Xie, Jiaqing, Wang, Aoran, Zhang, Wenlong, Torr, Philip, Zhou, Dongzhan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2509.01396
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915605599223808
author Wan, Haiyuan
Yang, Chen
Yu, Junchi
Tu, Meiqi
Lu, Jiaxuan
Yu, Di
Cao, Jianbao
Gao, Ben
Xie, Jiaqing
Wang, Aoran
Zhang, Wenlong
Torr, Philip
Zhou, Dongzhan
author_facet Wan, Haiyuan
Yang, Chen
Yu, Junchi
Tu, Meiqi
Lu, Jiaxuan
Yu, Di
Cao, Jianbao
Gao, Ben
Xie, Jiaqing
Wang, Aoran
Zhang, Wenlong
Torr, Philip
Zhou, Dongzhan
contents Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating their research capability faithfully is rather challenging due to the difficulty of collecting frontier research questions that genuinely capture researchers' attention and intellectual curiosity. To address this gap, we introduce DeepResearch Arena, a benchmark grounded in academic seminars that capture rich expert discourse and interaction, better reflecting real-world research environments and reducing the risk of data leakage. To automatically construct DeepResearch Arena, we propose a Multi-Agent Hierarchical Task Generation (MAHTG) system that extracts research-worthy inspirations from seminar transcripts. The MAHTG system further translates research-worthy inspirations into high-quality research tasks, ensuring the traceability of research task formulation while filtering noise. With the MAHTG system, we curate DeepResearch Arena with over 10,000 high-quality research tasks from over 200 academic seminars, spanning 12 disciplines, such as literature, history, and science. Our extensive evaluation shows that DeepResearch Arena presents substantial challenges for current state-of-the-art agents, with clear performance gaps observed across different models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01396
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
Wan, Haiyuan
Yang, Chen
Yu, Junchi
Tu, Meiqi
Lu, Jiaxuan
Yu, Di
Cao, Jianbao
Gao, Ben
Xie, Jiaqing
Wang, Aoran
Zhang, Wenlong
Torr, Philip
Zhou, Dongzhan
Artificial Intelligence
Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating their research capability faithfully is rather challenging due to the difficulty of collecting frontier research questions that genuinely capture researchers' attention and intellectual curiosity. To address this gap, we introduce DeepResearch Arena, a benchmark grounded in academic seminars that capture rich expert discourse and interaction, better reflecting real-world research environments and reducing the risk of data leakage. To automatically construct DeepResearch Arena, we propose a Multi-Agent Hierarchical Task Generation (MAHTG) system that extracts research-worthy inspirations from seminar transcripts. The MAHTG system further translates research-worthy inspirations into high-quality research tasks, ensuring the traceability of research task formulation while filtering noise. With the MAHTG system, we curate DeepResearch Arena with over 10,000 high-quality research tasks from over 200 academic seminars, spanning 12 disciplines, such as literature, history, and science. Our extensive evaluation shows that DeepResearch Arena presents substantial challenges for current state-of-the-art agents, with clear performance gaps observed across different models.
title DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
topic Artificial Intelligence
url https://arxiv.org/abs/2509.01396