Benchmarking Deep Search over Heterogeneous Enterprise Data
Fuente:
arXiv
Saved in:
| Main Authors: | Choubey, Prafulla Kumar, Peng, Xiangyu, Bhagavath, Shilpa, Huang, Kung-Hsiang, Xiong, Caiming, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
by: Zhang, Nan, et al.
Published: (2024)
by: Zhang, Nan, et al.
Published: (2024)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)
by: Huang, Kung-Hsiang, et al.
Published: (2023)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
by: Xie, Kaige, et al.
Published: (2024)
by: Xie, Kaige, et al.
Published: (2024)
Agentic Confidence Calibration
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
by: Huang, Kung-Hsiang, et al.
Published: (2024)
by: Huang, Kung-Hsiang, et al.
Published: (2024)
DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
by: Prabhakar, Akshara, et al.
Published: (2025)
by: Prabhakar, Akshara, et al.
Published: (2025)
Nudging the Boundaries of LLM Reasoning
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
by: Chen, Justin Chih-Yao, et al.
Published: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
by: Hoang, Thai, et al.
Published: (2025)
by: Hoang, Thai, et al.
Published: (2025)
Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems
by: Wang, Ling, et al.
Published: (2026)
by: Wang, Ling, et al.
Published: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
SafeWorld: Geo-Diverse Safety Alignment
by: Yin, Da, et al.
Published: (2024)
by: Yin, Da, et al.
Published: (2024)
The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
HPE:Answering Complex Questions over Text by Hybrid Question Parsing and Execution
by: Liu, Ye, et al.
Published: (2023)
by: Liu, Ye, et al.
Published: (2023)
NewsEdits 2.0: Learning the Intentions Behind Updating News
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
by: Tan, Jiejun, et al.
Published: (2025)
by: Tan, Jiejun, et al.
Published: (2025)
BEAVER: An Enterprise Benchmark for Text-to-SQL
by: Chen, Peter Baile, et al.
Published: (2024)
by: Chen, Peter Baile, et al.
Published: (2024)
Parameter-Efficient Detoxification with Contrastive Decoding
by: Niu, Tong, et al.
Published: (2024)
by: Niu, Tong, et al.
Published: (2024)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
by: Peng, Xiangyu, et al.
Published: (2025)
by: Peng, Xiangyu, et al.
Published: (2025)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
by: Nguyen, Xuan-Phi, et al.
Published: (2025)
by: Nguyen, Xuan-Phi, et al.
Published: (2025)
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
EnterpriseEM: Fine-tuned Embeddings for Enterprise Semantic Search
by: Rathinasamy, Kamalkumar, et al.
Published: (2024)
by: Rathinasamy, Kamalkumar, et al.
Published: (2024)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
Dingtalk DeepResearch: A Unified Multi Agent Framework for Adaptive Intelligence in Enterprise Environments
by: Chen, Mengyuan, et al.
Published: (2025)
by: Chen, Mengyuan, et al.
Published: (2025)
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
by: Lei, Fangyu, et al.
Published: (2024)
by: Lei, Fangyu, et al.
Published: (2024)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
by: Luo, Wenzhen, et al.
Published: (2025)
by: Luo, Wenzhen, et al.
Published: (2025)
DeepJSONEval: Benchmarking Complex Nested JSON Data Mining for Large Language Models
by: Zhou, Zhicheng, et al.
Published: (2025)
by: Zhou, Zhicheng, et al.
Published: (2025)
Similar Items
-
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026) -
Turning Conversations into Workflows: A Framework to Extract and Evaluate Dialog Workflows for Service AI Agents
by: Choubey, Prafulla Kumar, et al.
Published: (2025) -
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025) -
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026) -
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024)