Benchmarking Deep Search over Heterogeneous Enterprise Data

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choubey, Prafulla Kumar, Peng, Xiangyu, Bhagavath, Shilpa, Huang, Kung-Hsiang, Xiong, Caiming, Wu, Chien-Sheng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909665424572416
author Choubey, Prafulla Kumar
Peng, Xiangyu
Bhagavath, Shilpa
Huang, Kung-Hsiang
Xiong, Caiming
Wu, Chien-Sheng
author_facet Choubey, Prafulla Kumar
Peng, Xiangyu
Bhagavath, Shilpa
Huang, Kung-Hsiang
Xiong, Caiming
Wu, Chien-Sheng
contents We present a new benchmark for evaluating Deep Search--a realistic and complex form of retrieval-augmented generation (RAG) that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources. These include documents, meeting transcripts, Slack messages, GitHub, and URLs, which vary in structure and often contain human-to-human interactions. We build it using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers. We release our benchmark with both answerable and unanswerable queries, and retrieval pool of 39,190 enterprise artifacts, enabling fine-grained evaluation of long-context LLM and RAG systems. Our experiments reveal that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on our benchmark. With further analysis, we highlight retrieval as the main bottleneck: existing methods struggle to conduct deep searches and retrieve all necessary evidence. Consequently, they often reason over partial context, leading to significant performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23139
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Deep Search over Heterogeneous Enterprise Data
Choubey, Prafulla Kumar
Peng, Xiangyu
Bhagavath, Shilpa
Huang, Kung-Hsiang
Xiong, Caiming
Wu, Chien-Sheng
Computation and Language
Artificial Intelligence
We present a new benchmark for evaluating Deep Search--a realistic and complex form of retrieval-augmented generation (RAG) that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources. These include documents, meeting transcripts, Slack messages, GitHub, and URLs, which vary in structure and often contain human-to-human interactions. We build it using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers. We release our benchmark with both answerable and unanswerable queries, and retrieval pool of 39,190 enterprise artifacts, enabling fine-grained evaluation of long-context LLM and RAG systems. Our experiments reveal that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on our benchmark. With further analysis, we highlight retrieval as the main bottleneck: existing methods struggle to conduct deep searches and retrieve all necessary evidence. Consequently, they often reason over partial context, leading to significant performance degradation.
title Benchmarking Deep Search over Heterogeneous Enterprise Data
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.23139