Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Tao, Wang, Hao, Li, Changyu, Chai, Shenghua, Zhang, Minghui, Luo, Zhongtian, Zhou, Yuxuan, Jin, Haopeng, Kang, Zhaolu, Yang, Jiabing, Zhang, YiFan, Wang, Xinming, Yi, Hongzhu, He, Zheqi, Zheng, Jing-Shu, Yang, Xi, Huang, Yan, Wang, Liang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915997437394944
author Yu, Tao
Wang, Hao
Li, Changyu
Chai, Shenghua
Zhang, Minghui
Luo, Zhongtian
Zhou, Yuxuan
Jin, Haopeng
Kang, Zhaolu
Yang, Jiabing
Zhang, YiFan
Wang, Xinming
Yi, Hongzhu
He, Zheqi
Zheng, Jing-Shu
Yang, Xi
Huang, Yan
Wang, Liang
author_facet Yu, Tao
Wang, Hao
Li, Changyu
Chai, Shenghua
Zhang, Minghui
Luo, Zhongtian
Zhou, Yuxuan
Jin, Haopeng
Kang, Zhaolu
Yang, Jiabing
Zhang, YiFan
Wang, Xinming
Yi, Hongzhu
He, Zheqi
Zheng, Jing-Shu
Yang, Xi
Huang, Yan
Wang, Liang
contents Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08761
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
Yu, Tao
Wang, Hao
Li, Changyu
Chai, Shenghua
Zhang, Minghui
Luo, Zhongtian
Zhou, Yuxuan
Jin, Haopeng
Kang, Zhaolu
Yang, Jiabing
Zhang, YiFan
Wang, Xinming
Yi, Hongzhu
He, Zheqi
Zheng, Jing-Shu
Yang, Xi
Huang, Yan
Wang, Liang
Multiagent Systems
Machine Learning
Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments.
title Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
topic Multiagent Systems
Machine Learning
url https://arxiv.org/abs/2605.08761