Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xia, Boming, Lu, Qinghua, Zhu, Liming, Xing, Zhenchang, Zhao, Dehai, Zhang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents
von: Shamsujjoha, Md, et al.
Veröffentlicht: (2024)
von: Shamsujjoha, Md, et al.
Veröffentlicht: (2024)
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
von: Xia, Boming, et al.
Veröffentlicht: (2024)
von: Xia, Boming, et al.
Veröffentlicht: (2024)
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
A Reference Architecture for Designing Foundation Model based Systems
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
von: Lu, Qinghua, et al.
Veröffentlicht: (2025)
von: Lu, Qinghua, et al.
Veröffentlicht: (2025)
A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model
von: Zhou, Jingwen, et al.
Veröffentlicht: (2024)
von: Zhou, Jingwen, et al.
Veröffentlicht: (2024)
Agent Design Pattern Catalogue: A Collection of Architectural Patterns for Foundation Model based Agents
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
Uncertainty Propagation in LLM-Based Systems
von: Xia, Boming, et al.
Veröffentlicht: (2026)
von: Xia, Boming, et al.
Veröffentlicht: (2026)
A Taxonomy of Foundation Model based Systems through the Lens of Software Architecture
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
SeeAction: Towards Reverse Engineering How-What-Where of HCI Actions from Screencasts for UI Automation
von: Zhao, Dehai, et al.
Veröffentlicht: (2025)
von: Zhao, Dehai, et al.
Veröffentlicht: (2025)
Still Manual? Automated Linter Configuration via DSL-Based LLM Compilation of Coding Standards
von: Zhang, Zejun, et al.
Veröffentlicht: (2026)
von: Zhang, Zejun, et al.
Veröffentlicht: (2026)
Decentralised Governance-Driven Architecture for Designing Foundation Model based Systems: Exploring the Role of Blockchain in Responsible AI
von: Liu, Yue, et al.
Veröffentlicht: (2023)
von: Liu, Yue, et al.
Veröffentlicht: (2023)
Trust in Software Supply Chains: Blockchain-Enabled SBOM and the AIBOM Future
von: Xia, Boming, et al.
Veröffentlicht: (2023)
von: Xia, Boming, et al.
Veröffentlicht: (2023)
Privacy and Copyright Protection in Generative AI: A Lifecycle Perspective
von: Zhang, Dawen, et al.
Veröffentlicht: (2023)
von: Zhang, Dawen, et al.
Veröffentlicht: (2023)
Towards a Responsible AI Metrics Catalogue: A Collection of Metrics for AI Accountability
von: Xia, Boming, et al.
Veröffentlicht: (2023)
von: Xia, Boming, et al.
Veröffentlicht: (2023)
SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
von: Zhou, Jingwen, et al.
Veröffentlicht: (2025)
von: Zhou, Jingwen, et al.
Veröffentlicht: (2025)
To Be Forgotten or To Be Fair: Unveiling Fairness Implications of Machine Unlearning Methods
von: Zhang, Dawen, et al.
Veröffentlicht: (2023)
von: Zhang, Dawen, et al.
Veröffentlicht: (2023)
When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
von: Xing, Zhenchang, et al.
Veröffentlicht: (2025)
von: Xing, Zhenchang, et al.
Veröffentlicht: (2025)
AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development
von: Zhu, Yuecai, et al.
Veröffentlicht: (2026)
von: Zhu, Yuecai, et al.
Veröffentlicht: (2026)
AgentOps: Enabling Observability of LLM Agents
von: Dong, Liming, et al.
Veröffentlicht: (2024)
von: Dong, Liming, et al.
Veröffentlicht: (2024)
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
von: Jin, Haolin, et al.
Veröffentlicht: (2025)
von: Jin, Haolin, et al.
Veröffentlicht: (2025)
A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems
von: Zhang, Dawen, et al.
Veröffentlicht: (2024)
von: Zhang, Dawen, et al.
Veröffentlicht: (2024)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2025)
Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language Models
von: Liao, Dianshu, et al.
Veröffentlicht: (2025)
von: Liao, Dianshu, et al.
Veröffentlicht: (2025)
Grid-Orch: An LLM-Powered Orchestrator for Distribution Grid Simulation and Analytics
von: Liu, Boming, et al.
Veröffentlicht: (2026)
von: Liu, Boming, et al.
Veröffentlicht: (2026)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language Models
von: Zhang, Zejun, et al.
Veröffentlicht: (2024)
von: Zhang, Zejun, et al.
Veröffentlicht: (2024)
EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents
von: Liu, Junwei, et al.
Veröffentlicht: (2025)
von: Liu, Junwei, et al.
Veröffentlicht: (2025)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
von: He, Jiawei, et al.
Veröffentlicht: (2026)
von: He, Jiawei, et al.
Veröffentlicht: (2026)
Explore-Construct-Filter: An Automated Framework for Rich and Reliable API Knowledge Graph Construction
von: Sun, Yanbang, et al.
Veröffentlicht: (2025)
von: Sun, Yanbang, et al.
Veröffentlicht: (2025)
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
von: Ma, Wei, et al.
Veröffentlicht: (2025)
von: Ma, Wei, et al.
Veröffentlicht: (2025)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
von: Li, Dawei, et al.
Veröffentlicht: (2026)
von: Li, Dawei, et al.
Veröffentlicht: (2026)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
von: Yan, Shuo, et al.
Veröffentlicht: (2025)
von: Yan, Shuo, et al.
Veröffentlicht: (2025)
A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
von: Srinivasan, Vasundra
Veröffentlicht: (2026)
von: Srinivasan, Vasundra
Veröffentlicht: (2026)
MAAD: Automate Software Architecture Design through Knowledge-Driven Multi-Agent Collaboration
von: Li, Ruiyin, et al.
Veröffentlicht: (2025)
von: Li, Ruiyin, et al.
Veröffentlicht: (2025)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
von: Trae Research Team, et al.
Veröffentlicht: (2025)
von: Trae Research Team, et al.
Veröffentlicht: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
From Exploration to Revelation: Detecting Dark Patterns in Mobile Apps
von: Chen, Jieshan, et al.
Veröffentlicht: (2024)
von: Chen, Jieshan, et al.
Veröffentlicht: (2024)
Automated Soap Opera Testing Directed by LLMs and Scenario Knowledge: Feasibility, Challenges, and Road Ahead
von: Su, Yanqi, et al.
Veröffentlicht: (2024)
von: Su, Yanqi, et al.
Veröffentlicht: (2024)
From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents
von: Ma, Murong, et al.
Veröffentlicht: (2026)
von: Ma, Murong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Swiss Cheese Model for AI Safety: A Taxonomy and Reference Architecture for Multi-Layered Guardrails of Foundation Model Based Agents
von: Shamsujjoha, Md, et al.
Veröffentlicht: (2024) -
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
von: Xia, Boming, et al.
Veröffentlicht: (2024) -
Towards Responsible Generative AI: A Reference Architecture for Designing Foundation Model based Agents
von: Lu, Qinghua, et al.
Veröffentlicht: (2023) -
A Reference Architecture for Designing Foundation Model based Systems
von: Lu, Qinghua, et al.
Veröffentlicht: (2023) -
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
von: Lu, Qinghua, et al.
Veröffentlicht: (2025)