AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Qinghua, Zhao, Dehai, Liu, Yue, Zhang, Hao, Zhu, Liming, Xu, Xiwei, Shi, Angela, Tan, Tristan, Kazman, Rick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agent Design Pattern Catalogue: A Collection of Architectural Patterns for Foundation Model based Agents
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents
by: Shamsujjoha, Md, et al.
Published: (2024)
by: Shamsujjoha, Md, et al.
Published: (2024)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model
by: Zhou, Jingwen, et al.
Published: (2024)
by: Zhou, Jingwen, et al.
Published: (2024)
SeeAction: Towards Reverse Engineering How-What-Where of HCI Actions from Screencasts for UI Automation
by: Zhao, Dehai, et al.
Published: (2025)
by: Zhao, Dehai, et al.
Published: (2025)
A Reference Architecture for Designing Foundation Model based Systems
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
A Taxonomy of Foundation Model based Systems through the Lens of Software Architecture
by: Lu, Qinghua, et al.
Published: (2023)
by: Lu, Qinghua, et al.
Published: (2023)
Decentralised Governance-Driven Architecture for Designing Foundation Model based Systems: Exploring the Role of Blockchain in Responsible AI
by: Liu, Yue, et al.
Published: (2023)
by: Liu, Yue, et al.
Published: (2023)
SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
by: Zhou, Jingwen, et al.
Published: (2025)
by: Zhou, Jingwen, et al.
Published: (2025)
A Systematic Mapping Study on Architectural Approaches to Software Performance Analysis
by: Zhao, Yutong, et al.
Published: (2024)
by: Zhao, Yutong, et al.
Published: (2024)
An Extended Pattern Collection for Blockchain-based Applications
by: Xu, Xiwei, et al.
Published: (2025)
by: Xu, Xiwei, et al.
Published: (2025)
AgentOps: Enabling Observability of LLM Agents
by: Dong, Liming, et al.
Published: (2024)
by: Dong, Liming, et al.
Published: (2024)
RAGOps: Operating and Managing Retrieval-Augmented Generation Pipelines
by: Xu, Xiwei, et al.
Published: (2025)
by: Xu, Xiwei, et al.
Published: (2025)
Practitioner Views on Mobile App Accessibility: Practices and Challenges
by: Indika, Amila, et al.
Published: (2026)
by: Indika, Amila, et al.
Published: (2026)
Automated Soap Opera Testing Directed by LLMs and Scenario Knowledge: Feasibility, Challenges, and Road Ahead
by: Su, Yanqi, et al.
Published: (2024)
by: Su, Yanqi, et al.
Published: (2024)
Oops!... I did it again. Conclusion (In-)Stability in Quantitative Empirical Software Engineering: A Large-Scale Analysis
by: Hoess, Nicole, et al.
Published: (2025)
by: Hoess, Nicole, et al.
Published: (2025)
Leveraging Sustainable Systematic Literature Reviews
by: Santos, Vinicius dos, et al.
Published: (2025)
by: Santos, Vinicius dos, et al.
Published: (2025)
Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories
by: Hoess, Nicole, et al.
Published: (2025)
by: Hoess, Nicole, et al.
Published: (2025)
To Be Forgotten or To Be Fair: Unveiling Fairness Implications of Machine Unlearning Methods
by: Zhang, Dawen, et al.
Published: (2023)
by: Zhang, Dawen, et al.
Published: (2023)
Towards a Maturity Model for Systematic Literature Review Process
by: Santos, Vinicius dos, et al.
Published: (2022)
by: Santos, Vinicius dos, et al.
Published: (2022)
Trust in Software Supply Chains: Blockchain-Enabled SBOM and the AIBOM Future
by: Xia, Boming, et al.
Published: (2023)
by: Xia, Boming, et al.
Published: (2023)
Explaining the Contributing Factors for Vulnerability Detection in Machine Learning
by: Mouine, Esma, et al.
Published: (2024)
by: Mouine, Esma, et al.
Published: (2024)
Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language Models
by: Zhang, Zejun, et al.
Published: (2024)
by: Zhang, Zejun, et al.
Published: (2024)
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
by: Zhang, Zejun, et al.
Published: (2026)
by: Zhang, Zejun, et al.
Published: (2026)
Towards a Responsible AI Metrics Catalogue: A Collection of Metrics for AI Accountability
by: Xia, Boming, et al.
Published: (2023)
by: Xia, Boming, et al.
Published: (2023)
VulAgent: Hypothesis-Validation based Multi-Agent Vulnerability Detection
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
Exploring Accessibility Trends and Challenges in Mobile App Development: A Study of Stack Overflow Questions
by: Indika, Amila, et al.
Published: (2024)
by: Indika, Amila, et al.
Published: (2024)
ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
by: Li, Jialin, et al.
Published: (2026)
by: Li, Jialin, et al.
Published: (2026)
A Structured Approach to Safety Case Construction for AI Systems
by: Lee, Sung Une, et al.
Published: (2026)
by: Lee, Sung Une, et al.
Published: (2026)
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Architectural Patterns for Designing Quantum Artificial Intelligence Systems
by: Klymenko, Mykhailo, et al.
Published: (2024)
by: Klymenko, Mykhailo, et al.
Published: (2024)
An LLM-assisted approach to designing software architectures using ADD
by: Cervantes, Humberto, et al.
Published: (2025)
by: Cervantes, Humberto, et al.
Published: (2025)
Moderately Mighty: To What Extent Can Internal Software Metrics Predict App Popularity at Launch?
by: Opu, Md Nahidul Islam, et al.
Published: (2025)
by: Opu, Md Nahidul Islam, et al.
Published: (2025)
Detection of Technical Debt in Java Source Code
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
Everything is Context: Agentic File System Abstraction for Context Engineering
by: Xu, Xiwei, et al.
Published: (2025)
by: Xu, Xiwei, et al.
Published: (2025)
ESG Reporting Lifecycle Management with Large Language Models and AI Agents
by: Hoang, Thong, et al.
Published: (2026)
by: Hoang, Thong, et al.
Published: (2026)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
by: Hong, Sirui, et al.
Published: (2026)
by: Hong, Sirui, et al.
Published: (2026)
EvalSVA: Multi-Agent Evaluators for Next-Gen Software Vulnerability Assessment
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
by: Widanapathiranage, Dilani, et al.
Published: (2025)
by: Widanapathiranage, Dilani, et al.
Published: (2025)
Similar Items
-
Agent Design Pattern Catalogue: A Collection of Architectural Patterns for Foundation Model based Agents
by: Liu, Yue, et al.
Published: (2024) -
Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents
by: Shamsujjoha, Md, et al.
Published: (2024) -
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
by: Xia, Boming, et al.
Published: (2024) -
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
by: Lu, Qinghua, et al.
Published: (2023) -
A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model
by: Zhou, Jingwen, et al.
Published: (2024)