AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bogavelli, Tara, Sharma, Roshnee, Subramani, Hari |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
von: Sun, Haochen, et al.
Veröffentlicht: (2025)
von: Sun, Haochen, et al.
Veröffentlicht: (2025)
Governed Memory: A Production Architecture for Multi-Agent Workflows
von: Taheri, Hamed
Veröffentlicht: (2026)
von: Taheri, Hamed
Veröffentlicht: (2026)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
von: Ye, Rui, et al.
Veröffentlicht: (2025)
von: Ye, Rui, et al.
Veröffentlicht: (2025)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
von: Styles, Olly, et al.
Veröffentlicht: (2024)
von: Styles, Olly, et al.
Veröffentlicht: (2024)
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
von: Yang, Yiqian, et al.
Veröffentlicht: (2025)
von: Yang, Yiqian, et al.
Veröffentlicht: (2025)
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
Layered Chain-of-Thought Prompting for Multi-Agent LLM Systems: A Comprehensive Approach to Explainable Large Language Models
von: Sanwal, Manish
Veröffentlicht: (2025)
von: Sanwal, Manish
Veröffentlicht: (2025)
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
von: Fang, Jinyuan, et al.
Veröffentlicht: (2025)
von: Fang, Jinyuan, et al.
Veröffentlicht: (2025)
CACA Agent: Capability Collaboration based AI Agent
von: Xu, Peng, et al.
Veröffentlicht: (2024)
von: Xu, Peng, et al.
Veröffentlicht: (2024)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
von: Ke, Zixuan, et al.
Veröffentlicht: (2026)
von: Ke, Zixuan, et al.
Veröffentlicht: (2026)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
von: Siegel, Zachary S., et al.
Veröffentlicht: (2024)
von: Siegel, Zachary S., et al.
Veröffentlicht: (2024)
SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning
von: Li, Xuchen, et al.
Veröffentlicht: (2025)
von: Li, Xuchen, et al.
Veröffentlicht: (2025)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
von: Liu, Zijun, et al.
Veröffentlicht: (2023)
von: Liu, Zijun, et al.
Veröffentlicht: (2023)
Efficient Agents: Building Effective Agents While Reducing Cost
von: Wang, Ningning, et al.
Veröffentlicht: (2025)
von: Wang, Ningning, et al.
Veröffentlicht: (2025)
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
von: Sun, Yiqun, et al.
Veröffentlicht: (2026)
von: Sun, Yiqun, et al.
Veröffentlicht: (2026)
ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis
von: Zhao, Huiya, et al.
Veröffentlicht: (2025)
von: Zhao, Huiya, et al.
Veröffentlicht: (2025)
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
von: Zhu, Jianing, et al.
Veröffentlicht: (2026)
von: Zhu, Jianing, et al.
Veröffentlicht: (2026)
Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
von: Jin, Haibo, et al.
Veröffentlicht: (2026)
von: Jin, Haibo, et al.
Veröffentlicht: (2026)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
von: Zhang, Boxuan, et al.
Veröffentlicht: (2026)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2026)
BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
von: Ahmed, Fahim, et al.
Veröffentlicht: (2025)
von: Ahmed, Fahim, et al.
Veröffentlicht: (2025)
Enhancing Online Learning Efficiency Through Heterogeneous Resource Integration with a Multi-Agent RAG System
von: Srivastav, Devansh, et al.
Veröffentlicht: (2025)
von: Srivastav, Devansh, et al.
Veröffentlicht: (2025)
The Orchestration of Multi-Agent Systems: Architectures, Protocols, and Enterprise Adoption
von: Adimulam, Apoorva, et al.
Veröffentlicht: (2026)
von: Adimulam, Apoorva, et al.
Veröffentlicht: (2026)
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents
von: Rosati, Riccardo, et al.
Veröffentlicht: (2026)
von: Rosati, Riccardo, et al.
Veröffentlicht: (2026)
LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
von: Zhang, Danqing, et al.
Veröffentlicht: (2025)
von: Zhang, Danqing, et al.
Veröffentlicht: (2025)
AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2026)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2026)
Investigating the Potential of Large Language Model-Based Router Multi-Agent Architectures for Foundation Design Automation: A Task Classification and Expert Selection Study
von: Youwai, Sompote, et al.
Veröffentlicht: (2025)
von: Youwai, Sompote, et al.
Veröffentlicht: (2025)
From Competition to Coordination: Market Making as a Scalable Framework for Safe and Aligned Multi-Agent LLM Systems
von: Gho, Brendan, et al.
Veröffentlicht: (2025)
von: Gho, Brendan, et al.
Veröffentlicht: (2025)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
von: Zhu, Yinghao, et al.
Veröffentlicht: (2025)
von: Zhu, Yinghao, et al.
Veröffentlicht: (2025)
Grammar Search for Multi-Agent Systems
von: Singh, Mayank, et al.
Veröffentlicht: (2025)
von: Singh, Mayank, et al.
Veröffentlicht: (2025)
Cognitive Duality for Adaptive Web Agents
von: Liu, Jiarun, et al.
Veröffentlicht: (2025)
von: Liu, Jiarun, et al.
Veröffentlicht: (2025)
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
von: Sinha, Aarush, et al.
Veröffentlicht: (2026)
von: Sinha, Aarush, et al.
Veröffentlicht: (2026)
Multimodal Safety Evaluation in Generative Agent Social Simulations
von: Vera, Alhim, et al.
Veröffentlicht: (2025)
von: Vera, Alhim, et al.
Veröffentlicht: (2025)
Multi-Agent Collaboration via Evolving Orchestration
von: Dang, Yufan, et al.
Veröffentlicht: (2025)
von: Dang, Yufan, et al.
Veröffentlicht: (2025)
Adaptive Memory Admission Control for LLM Agents
von: Zhang, Guilin, et al.
Veröffentlicht: (2026)
von: Zhang, Guilin, et al.
Veröffentlicht: (2026)
Scaling Small Agents Through Strategy Auctions
von: Alazraki, Lisa, et al.
Veröffentlicht: (2026)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2026)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025)
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025)
Achieving Unanimous Consensus Through Multi-Agent Deliberation
von: Pokharel, Apurba, et al.
Veröffentlicht: (2025)
von: Pokharel, Apurba, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
von: Sun, Haochen, et al.
Veröffentlicht: (2025) -
Governed Memory: A Production Architecture for Multi-Agent Workflows
von: Taheri, Hamed
Veröffentlicht: (2026) -
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
von: Zhang, Yifei, et al.
Veröffentlicht: (2026) -
MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
von: Ye, Rui, et al.
Veröffentlicht: (2025) -
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
von: Styles, Olly, et al.
Veröffentlicht: (2024)