DRBench: A Realistic Benchmark for Enterprise Deep Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abaskohi, Amirhossein, Chen, Tianyi, Muñoz-Mármol, Miguel, Fox, Curtis, Ramesh, Amrutha Varshini, Marcotte, Étienne, Lù, Xing Han, Chapados, Nicolas, Gella, Spandana, West, Peter, Carenini, Giuseppe, Pal, Christopher, Drouin, Alexandre, Laradji, Issam H.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910046835703808
author Abaskohi, Amirhossein
Chen, Tianyi
Muñoz-Mármol, Miguel
Fox, Curtis
Ramesh, Amrutha Varshini
Marcotte, Étienne
Lù, Xing Han
Chapados, Nicolas
Gella, Spandana
West, Peter
Carenini, Giuseppe
Pal, Christopher
Drouin, Alexandre
Laradji, Issam H.
author_facet Abaskohi, Amirhossein
Chen, Tianyi
Muñoz-Mármol, Miguel
Fox, Curtis
Ramesh, Amrutha Varshini
Marcotte, Étienne
Lù, Xing Han
Chapados, Nicolas
Gella, Spandana
West, Peter
Carenini, Giuseppe
Pal, Christopher
Drouin, Alexandre
Laradji, Issam H.
contents We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (for example, "What changes should we make to our product roadmap to ensure compliance with this standard?") that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web. Tasks are generated through a carefully designed synthesis pipeline with human-in-the-loop verification, and agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports. We release 100 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance. We demonstrate the effectiveness of DRBench by evaluating diverse DR agents across open- and closed-source models (such as GPT, Llama, and Qwen) and DR strategies, highlighting their strengths, weaknesses, and the critical path for advancing enterprise deep research. Code and data are available at https://github.com/ServiceNow/drbench.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DRBench: A Realistic Benchmark for Enterprise Deep Research
Abaskohi, Amirhossein
Chen, Tianyi
Muñoz-Mármol, Miguel
Fox, Curtis
Ramesh, Amrutha Varshini
Marcotte, Étienne
Lù, Xing Han
Chapados, Nicolas
Gella, Spandana
West, Peter
Carenini, Giuseppe
Pal, Christopher
Drouin, Alexandre
Laradji, Issam H.
Computation and Language
We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (for example, "What changes should we make to our product roadmap to ensure compliance with this standard?") that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web. Tasks are generated through a carefully designed synthesis pipeline with human-in-the-loop verification, and agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports. We release 100 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance. We demonstrate the effectiveness of DRBench by evaluating diverse DR agents across open- and closed-source models (such as GPT, Llama, and Qwen) and DR strategies, highlighting their strengths, weaknesses, and the critical path for advancing enterprise deep research. Code and data are available at https://github.com/ServiceNow/drbench.
title DRBench: A Realistic Benchmark for Enterprise Deep Research
topic Computation and Language
url https://arxiv.org/abs/2510.00172