AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bogavelli, Tara, Sharma, Roshnee, Subramani, Hari
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909982483546112
author Bogavelli, Tara
Sharma, Roshnee
Subramani, Hari
author_facet Bogavelli, Tara
Sharma, Roshnee
Subramani, Hari
contents While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact within complex multi-agent systems. This study aims to address these gaps by providing a comprehensive enterprise-specific benchmark evaluating 18 distinct agentic configurations across state-of-the-art large language models. We examine four critical agentic system dimensions: orchestration strategy, agent prompt implementation (ReAct versus function calling), memory architecture, and thinking tool integration. Our benchmark reveals significant model-specific architectural preferences that challenge the prevalent one-size-fits-all paradigm in agentic AI systems. It also reveals significant weaknesses in overall agentic performance on enterprise tasks with the highest scoring models achieving a maximum of only 35.3\% success on the more complex task and 70.8\% on the simpler task. We hope these findings inform the design of future agentic systems by enabling more empirically backed decisions regarding architectural components and model selection.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10769
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
Bogavelli, Tara
Sharma, Roshnee
Subramani, Hari
Artificial Intelligence
Computation and Language
Multiagent Systems
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact within complex multi-agent systems. This study aims to address these gaps by providing a comprehensive enterprise-specific benchmark evaluating 18 distinct agentic configurations across state-of-the-art large language models. We examine four critical agentic system dimensions: orchestration strategy, agent prompt implementation (ReAct versus function calling), memory architecture, and thinking tool integration. Our benchmark reveals significant model-specific architectural preferences that challenge the prevalent one-size-fits-all paradigm in agentic AI systems. It also reveals significant weaknesses in overall agentic performance on enterprise tasks with the highest scoring models achieving a maximum of only 35.3\% success on the more complex task and 70.8\% on the simpler task. We hope these findings inform the design of future agentic systems by enabling more empirically backed decisions regarding architectural components and model selection.
title AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
topic Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2509.10769