Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Akshathala, Sreemaee, Adnan, Bassam, Ramesh, Mahisha, Vaidhyanathan, Karthik, Muhammed, Basil, Parthasarathy, Kannan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
di: Adnan, Bassam, et al.
Pubblicazione: (2026)
di: Adnan, Bassam, et al.
Pubblicazione: (2026)
Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI
di: Koch, Christopher, et al.
Pubblicazione: (2026)
di: Koch, Christopher, et al.
Pubblicazione: (2026)
Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps
di: Parthasarathy, Kannan, et al.
Pubblicazione: (2025)
di: Parthasarathy, Kannan, et al.
Pubblicazione: (2025)
Sherlock: Reliable and Efficient Agentic Workflow Execution
di: Ro, Yeonju, et al.
Pubblicazione: (2025)
di: Ro, Yeonju, et al.
Pubblicazione: (2025)
AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations
di: Mitra, Shaswata, et al.
Pubblicazione: (2026)
di: Mitra, Shaswata, et al.
Pubblicazione: (2026)
A Generic Modelling Framework for Last-Mile Delivery Systems
di: Gürcan, Önder, et al.
Pubblicazione: (2025)
di: Gürcan, Önder, et al.
Pubblicazione: (2025)
Runtime Composition in Dynamic System of Systems: A Systematic Review of Challenges, Solutions, Tools, and Evaluation Methods
di: Ashfaq, Muhammad, et al.
Pubblicazione: (2025)
di: Ashfaq, Muhammad, et al.
Pubblicazione: (2025)
Toward Agentic Software Engineering Beyond Code: Framing Vision, Values, and Vocabulary
di: Hoda, Rashina
Pubblicazione: (2025)
di: Hoda, Rashina
Pubblicazione: (2025)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
di: Feng, Yukang, et al.
Pubblicazione: (2026)
di: Feng, Yukang, et al.
Pubblicazione: (2026)
AutoFSM: A Multi-agent Framework for FSM Code Generation with IR and SystemC-Based Testing
di: Luo, Qiuming, et al.
Pubblicazione: (2025)
di: Luo, Qiuming, et al.
Pubblicazione: (2025)
Can AI Agents Generate Microservices? How Far are We?
di: Adnan, Bassam, et al.
Pubblicazione: (2026)
di: Adnan, Bassam, et al.
Pubblicazione: (2026)
ABMax: A JAX-based Agent-based Modeling Framework
di: Chaturvedi, Siddharth, et al.
Pubblicazione: (2025)
di: Chaturvedi, Siddharth, et al.
Pubblicazione: (2025)
Accelerating Drug Discovery Through Agentic AI: A Multi-Agent Approach to Laboratory Automation in the DMTA Cycle
di: Fehlis, Yao, et al.
Pubblicazione: (2025)
di: Fehlis, Yao, et al.
Pubblicazione: (2025)
The Lifecycle Workbench -- A Configurable Framework for Digitized Product Maintenance Services
di: Briechle, Dominique, et al.
Pubblicazione: (2025)
di: Briechle, Dominique, et al.
Pubblicazione: (2025)
The Hidden Bloat in Machine Learning Systems
di: Zhang, Huaifeng, et al.
Pubblicazione: (2025)
di: Zhang, Huaifeng, et al.
Pubblicazione: (2025)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
di: Saxena, Siddhant, et al.
Pubblicazione: (2026)
di: Saxena, Siddhant, et al.
Pubblicazione: (2026)
Nexus: A Lightweight and Scalable Multi-Agent Framework for Complex Tasks Automation
di: Sami, Humza, et al.
Pubblicazione: (2025)
di: Sami, Humza, et al.
Pubblicazione: (2025)
Extending the OWASP Multi-Agentic System Threat Modeling Guide: Insights from Multi-Agent Security Research
di: Krawiecka, Klaudia, et al.
Pubblicazione: (2025)
di: Krawiecka, Klaudia, et al.
Pubblicazione: (2025)
A Research Agenda on Agents and Software Engineering: Outcomes from the Rio A2SE Seminar
di: Taibi, Davide, et al.
Pubblicazione: (2026)
di: Taibi, Davide, et al.
Pubblicazione: (2026)
Bridging the Prototype-Production Gap: A Multi-Agent System for Notebooks Transformation
di: Elhashemy, Hanya, et al.
Pubblicazione: (2025)
di: Elhashemy, Hanya, et al.
Pubblicazione: (2025)
Fairness in Multi-Agent Systems for Software Engineering: An SDLC-Oriented Rapid Review
di: Yang-Smith, Corey, et al.
Pubblicazione: (2026)
di: Yang-Smith, Corey, et al.
Pubblicazione: (2026)
Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development
di: Bowers, Brian, et al.
Pubblicazione: (2025)
di: Bowers, Brian, et al.
Pubblicazione: (2025)
TDFlow: Agentic Workflows for Test Driven Development
di: Han, Kevin, et al.
Pubblicazione: (2025)
di: Han, Kevin, et al.
Pubblicazione: (2025)
Spec Kit Agents: Context-Grounded Agentic Workflows
di: Taghavi, Pardis, et al.
Pubblicazione: (2026)
di: Taghavi, Pardis, et al.
Pubblicazione: (2026)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
di: Salim, Mohamad, et al.
Pubblicazione: (2026)
di: Salim, Mohamad, et al.
Pubblicazione: (2026)
Enhancing Holonic Architecture with Natural Language Processing for System of Systems
di: Ashfaq, Muhammad, et al.
Pubblicazione: (2024)
di: Ashfaq, Muhammad, et al.
Pubblicazione: (2024)
Qualixar OS: A Universal Operating System for AI Agent Orchestration
di: Bhardwaj, Varun Pratap
Pubblicazione: (2026)
di: Bhardwaj, Varun Pratap
Pubblicazione: (2026)
Agentic Agile-V: From Vibe Coding to Verified Engineering in Software and Hardware Development
di: Koch, Christopher
Pubblicazione: (2026)
di: Koch, Christopher
Pubblicazione: (2026)
Agile V: A Compliance-Ready Framework for AI-Augmented Engineering -- From Concept to Audit-Ready Delivery
di: Koch, Christopher, et al.
Pubblicazione: (2026)
di: Koch, Christopher, et al.
Pubblicazione: (2026)
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
di: Li, Yang, et al.
Pubblicazione: (2025)
di: Li, Yang, et al.
Pubblicazione: (2025)
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
di: Zhou, Chenyu, et al.
Pubblicazione: (2026)
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
di: Joos, Pascal, et al.
Pubblicazione: (2025)
di: Joos, Pascal, et al.
Pubblicazione: (2025)
Cognitive Agents Powered by Large Language Models for Agile Software Project Management
di: Cinkusz, Konrad, et al.
Pubblicazione: (2025)
di: Cinkusz, Konrad, et al.
Pubblicazione: (2025)
REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
di: Pabba, Anvith, et al.
Pubblicazione: (2025)
di: Pabba, Anvith, et al.
Pubblicazione: (2025)
Deterministic vs. LLM-Controlled Orchestration for COBOL-to-Python Modernization
di: Lwin, Naing Oo, et al.
Pubblicazione: (2026)
di: Lwin, Naing Oo, et al.
Pubblicazione: (2026)
Real-Time BDI Agents: a model and its implementation
di: Traldi, Andrea, et al.
Pubblicazione: (2022)
di: Traldi, Andrea, et al.
Pubblicazione: (2022)
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls
di: Zhang, Zeyu, et al.
Pubblicazione: (2026)
di: Zhang, Zeyu, et al.
Pubblicazione: (2026)
A Step Towards a Universal Method for Modeling and Implementing Cross-Organizational Business Processes
di: Zeisler, Gerhard, et al.
Pubblicazione: (2024)
di: Zeisler, Gerhard, et al.
Pubblicazione: (2024)
Securing Smart Contract Languages with a Unified Agentic Framework for Vulnerability Repair in Solidity and Move
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
di: Adnan, Bassam, et al.
Pubblicazione: (2026) -
Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI
di: Koch, Christopher, et al.
Pubblicazione: (2026) -
Engineering LLM Powered Multi-agent Framework for Autonomous CloudOps
di: Parthasarathy, Kannan, et al.
Pubblicazione: (2025) -
Sherlock: Reliable and Efficient Agentic Workflow Execution
di: Ro, Yeonju, et al.
Pubblicazione: (2025) -
AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations
di: Mitra, Shaswata, et al.
Pubblicazione: (2026)