Benchmarking Deep Search over Heterogeneous Enterprise Data
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909665424572416 |
|---|---|
| author | Choubey, Prafulla Kumar Peng, Xiangyu Bhagavath, Shilpa Huang, Kung-Hsiang Xiong, Caiming Wu, Chien-Sheng |
| author_facet | Choubey, Prafulla Kumar Peng, Xiangyu Bhagavath, Shilpa Huang, Kung-Hsiang Xiong, Caiming Wu, Chien-Sheng |
| contents | We present a new benchmark for evaluating Deep Search--a realistic and complex form of retrieval-augmented generation (RAG) that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources. These include documents, meeting transcripts, Slack messages, GitHub, and URLs, which vary in structure and often contain human-to-human interactions. We build it using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers. We release our benchmark with both answerable and unanswerable queries, and retrieval pool of 39,190 enterprise artifacts, enabling fine-grained evaluation of long-context LLM and RAG systems. Our experiments reveal that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on our benchmark. With further analysis, we highlight retrieval as the main bottleneck: existing methods struggle to conduct deep searches and retrieve all necessary evidence. Consequently, they often reason over partial context, leading to significant performance degradation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_23139 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Benchmarking Deep Search over Heterogeneous Enterprise Data Choubey, Prafulla Kumar Peng, Xiangyu Bhagavath, Shilpa Huang, Kung-Hsiang Xiong, Caiming Wu, Chien-Sheng Computation and Language Artificial Intelligence We present a new benchmark for evaluating Deep Search--a realistic and complex form of retrieval-augmented generation (RAG) that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources. These include documents, meeting transcripts, Slack messages, GitHub, and URLs, which vary in structure and often contain human-to-human interactions. We build it using a synthetic data pipeline that simulates business workflows across product planning, development, and support stages, generating interconnected content with realistic noise and multi-hop questions with guaranteed ground-truth answers. We release our benchmark with both answerable and unanswerable queries, and retrieval pool of 39,190 enterprise artifacts, enabling fine-grained evaluation of long-context LLM and RAG systems. Our experiments reveal that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on our benchmark. With further analysis, we highlight retrieval as the main bottleneck: existing methods struggle to conduct deep searches and retrieve all necessary evidence. Consequently, they often reason over partial context, leading to significant performance degradation. |
| title | Benchmarking Deep Search over Heterogeneous Enterprise Data |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2506.23139 |