Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915997437394944 |
|---|---|
| author | Yu, Tao Wang, Hao Li, Changyu Chai, Shenghua Zhang, Minghui Luo, Zhongtian Zhou, Yuxuan Jin, Haopeng Kang, Zhaolu Yang, Jiabing Zhang, YiFan Wang, Xinming Yi, Hongzhu He, Zheqi Zheng, Jing-Shu Yang, Xi Huang, Yan Wang, Liang |
| author_facet | Yu, Tao Wang, Hao Li, Changyu Chai, Shenghua Zhang, Minghui Luo, Zhongtian Zhou, Yuxuan Jin, Haopeng Kang, Zhaolu Yang, Jiabing Zhang, YiFan Wang, Xinming Yi, Hongzhu He, Zheqi Zheng, Jing-Shu Yang, Xi Huang, Yan Wang, Liang |
| contents | Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_08761 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows Yu, Tao Wang, Hao Li, Changyu Chai, Shenghua Zhang, Minghui Luo, Zhongtian Zhou, Yuxuan Jin, Haopeng Kang, Zhaolu Yang, Jiabing Zhang, YiFan Wang, Xinming Yi, Hongzhu He, Zheqi Zheng, Jing-Shu Yang, Xi Huang, Yan Wang, Liang Multiagent Systems Machine Learning Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. However, existing enterprise benchmarks largely evaluate single agents with broad tool access, while existing multi-agent benchmarks rarely capture realistic enterprise constraints such as role specialization, access control, stateful business systems, and policy-based approvals. We introduce \textsc{EntCollabBench}, a benchmark for evaluating enterprise multi-agent collaboration. \textsc{EntCollabBench} simulates a permission-isolated organization with 11 role-specialized agents across six departments and contains two evaluation subsets: a Workflow subset, where agents collaboratively modify enterprise system states, and an Approval subset, where agents make policy-grounded decisions. Evaluation is based on execution traces, database state verification, and deterministic policy adjudication rather than natural-language response judging. Experiments with representative LLM agents show that current models still struggle with end-to-end enterprise collaboration, especially in delegation, context transfer, parameter grounding, workflow closure, and decision commitment. \textsc{EntCollabBench} provides a reproducible testbed for measuring and improving agent systems intended for realistic organizational environments. |
| title | Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows |
| topic | Multiagent Systems Machine Learning |
| url | https://arxiv.org/abs/2605.08761 |