EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Malay, Shiva Krishna Reddy, Nayak, Shravan, Nair, Jishnu Sethumadhavan, Davasam, Sagar, Tiwari, Aman, Madhusudhan, Sathwik Tejaswi, Nemala, Sridhar Krishna, Sunkara, Srinivas, Rajeswar, Sai
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911514847346688
author Malay, Shiva Krishna Reddy
Nayak, Shravan
Nair, Jishnu Sethumadhavan
Davasam, Sagar
Tiwari, Aman
Madhusudhan, Sathwik Tejaswi
Nemala, Sridhar Krishna
Sunkara, Srinivas
Rajeswar, Sai
author_facet Malay, Shiva Krishna Reddy
Nayak, Shravan
Nair, Jishnu Sethumadhavan
Davasam, Sagar
Tiwari, Aman
Madhusudhan, Sathwik Tejaswi
Nemala, Sridhar Krishna
Sunkara, Srinivas
Rajeswar, Sai
contents Large language models are shifting from passive information providers to active agents intended for complex workflows. However, their deployment as reliable AI workers in enterprise is stalled by benchmarks that fail to capture the intricacies of professional environments, specifically, the need for long-horizon planning amidst persistent state changes and strict access protocols. In this work, we introduce EnterpriseOps-Gym, a benchmark designed to evaluate agentic planning in realistic enterprise settings. Specifically, EnterpriseOps-Gym features a containerized sandbox with 164 database tables and 512 functional tools to mimic real-world search friction. Within this environment, agents are evaluated on 1,150 expert-curated tasks across eight mission-critical verticals (including Customer Service, HR, and IT). Our evaluation of 14 frontier models reveals critical limitations in state-of-the-art models: the top-performing Claude Opus 4.5 achieves only 37.4% success. Further analysis shows that providing oracle human plans improves performance by 14-35 percentage points, pinpointing strategic reasoning as the primary bottleneck. Additionally, agents frequently fail to refuse infeasible tasks (best model achieves 53.9%), leading to unintended and potentially harmful side effects. Our findings underscore that current agents are not yet ready for autonomous enterprise deployment. More broadly, EnterpriseOps-Gym provides a concrete testbed to advance the robustness of agentic planning in professional workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13594
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Malay, Shiva Krishna Reddy
Nayak, Shravan
Nair, Jishnu Sethumadhavan
Davasam, Sagar
Tiwari, Aman
Madhusudhan, Sathwik Tejaswi
Nemala, Sridhar Krishna
Sunkara, Srinivas
Rajeswar, Sai
Artificial Intelligence
Machine Learning
Large language models are shifting from passive information providers to active agents intended for complex workflows. However, their deployment as reliable AI workers in enterprise is stalled by benchmarks that fail to capture the intricacies of professional environments, specifically, the need for long-horizon planning amidst persistent state changes and strict access protocols. In this work, we introduce EnterpriseOps-Gym, a benchmark designed to evaluate agentic planning in realistic enterprise settings. Specifically, EnterpriseOps-Gym features a containerized sandbox with 164 database tables and 512 functional tools to mimic real-world search friction. Within this environment, agents are evaluated on 1,150 expert-curated tasks across eight mission-critical verticals (including Customer Service, HR, and IT). Our evaluation of 14 frontier models reveals critical limitations in state-of-the-art models: the top-performing Claude Opus 4.5 achieves only 37.4% success. Further analysis shows that providing oracle human plans improves performance by 14-35 percentage points, pinpointing strategic reasoning as the primary bottleneck. Additionally, agents frequently fail to refuse infeasible tasks (best model achieves 53.9%), leading to unintended and potentially harmful side effects. Our findings underscore that current agents are not yet ready for autonomous enterprise deployment. More broadly, EnterpriseOps-Gym provides a concrete testbed to advance the robustness of agentic planning in professional workflows.
title EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.13594