Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
Fuente:
arXiv
Guardado en:
| Autor principal: | Pandey, Mukund |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluation Framework for AI Systems in "the Wild"
por: Jabbour, Sarah, et al.
Publicado: (2025)
por: Jabbour, Sarah, et al.
Publicado: (2025)
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
por: Bhonsle, Roshita, et al.
Publicado: (2025)
por: Bhonsle, Roshita, et al.
Publicado: (2025)
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
por: Janet, et al.
Publicado: (2025)
por: Janet, et al.
Publicado: (2025)
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
por: Zhang, Hong, et al.
Publicado: (2026)
por: Zhang, Hong, et al.
Publicado: (2026)
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
por: Mehta, Sushant
Publicado: (2025)
por: Mehta, Sushant
Publicado: (2025)
Detecting Silent Failures in Multi-Agentic AI Trajectories
por: Pathak, Divya, et al.
Publicado: (2025)
por: Pathak, Divya, et al.
Publicado: (2025)
A Unified Framework for the Evaluation of LLM Agentic Capabilities
por: Zhu, Pengyu, et al.
Publicado: (2026)
por: Zhu, Pengyu, et al.
Publicado: (2026)
Holistic Evaluation and Failure Diagnosis of AI Agents
por: Madvil, Netta, et al.
Publicado: (2026)
por: Madvil, Netta, et al.
Publicado: (2026)
Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
por: Akshathala, Sreemaee, et al.
Publicado: (2025)
por: Akshathala, Sreemaee, et al.
Publicado: (2025)
Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
por: Chhabra, Anshuman, et al.
Publicado: (2025)
por: Chhabra, Anshuman, et al.
Publicado: (2025)
LightAgent: Production-level Open-source Agentic AI Framework
por: Cai, Weige, et al.
Publicado: (2025)
por: Cai, Weige, et al.
Publicado: (2025)
Creative Adversarial Testing (CAT): A Novel Framework for Evaluating Goal-Oriented Agentic AI Systems
por: Dhrif, Hassen
Publicado: (2025)
por: Dhrif, Hassen
Publicado: (2025)
RAIL in the Wild: Operationalizing Responsible AI Evaluation Using Anthropic's Value Dataset
por: Verma, Sumit, et al.
Publicado: (2025)
por: Verma, Sumit, et al.
Publicado: (2025)
The Auton Agentic AI Framework
por: Cao, Sheng, et al.
Publicado: (2026)
por: Cao, Sheng, et al.
Publicado: (2026)
DAO-AI: Evaluating Collective Decision-Making through Agentic AI in Decentralized Governance
por: Capponi, Agostino, et al.
Publicado: (2025)
por: Capponi, Agostino, et al.
Publicado: (2025)
Agentic Design Patterns: A System-Theoretic Framework
por: Dao, Minh-Dung, et al.
Publicado: (2026)
por: Dao, Minh-Dung, et al.
Publicado: (2026)
Zero-Direction Probing: A Linear-Algebraic Framework for Deep Analysis of Large-Language-Model Drift
por: Pandey, Amit
Publicado: (2025)
por: Pandey, Amit
Publicado: (2025)
AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems
por: Lee, YenTing, et al.
Publicado: (2026)
por: Lee, YenTing, et al.
Publicado: (2026)
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
por: Kartik, NVJK, et al.
Publicado: (2025)
por: Kartik, NVJK, et al.
Publicado: (2025)
WildSpoof Challenge Evaluation Plan
por: Wu, Yihan, et al.
Publicado: (2025)
por: Wu, Yihan, et al.
Publicado: (2025)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
por: Li, Jingyue, et al.
Publicado: (2026)
por: Li, Jingyue, et al.
Publicado: (2026)
Proper Scoring Rules for Agentic Uncertainty Quantification
por: Raghu, Suresh, et al.
Publicado: (2026)
por: Raghu, Suresh, et al.
Publicado: (2026)
Performant LLM Agentic Framework for Conversational AI
por: Casella, Alex, et al.
Publicado: (2025)
por: Casella, Alex, et al.
Publicado: (2025)
Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier
por: Henry, Jazmia
Publicado: (2026)
por: Henry, Jazmia
Publicado: (2026)
A Conceptual Framework for AI Capability Evaluations
por: Carro, María Victoria, et al.
Publicado: (2025)
por: Carro, María Victoria, et al.
Publicado: (2025)
Control Plane as a Tool: A Scalable Design Pattern for Agentic AI Systems
por: Kandasamy, Sivasathivel
Publicado: (2025)
por: Kandasamy, Sivasathivel
Publicado: (2025)
Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
por: Vinay, Vaishali
Publicado: (2025)
por: Vinay, Vaishali
Publicado: (2025)
GuidelineGuard: An Agentic Framework for Medical Note Evaluation with Guideline Adherence
por: Shahriyear, MD Ragib
Publicado: (2024)
por: Shahriyear, MD Ragib
Publicado: (2024)
Ethical AI: Towards Defining a Collective Evaluation Framework
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
DREAM: Deep Research Evaluation with Agentic Metrics
por: Avraham, Elad Ben, et al.
Publicado: (2026)
por: Avraham, Elad Ben, et al.
Publicado: (2026)
Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems
por: Shukla, Manish
Publicado: (2025)
por: Shukla, Manish
Publicado: (2025)
Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals
por: Menon, Achyutha, et al.
Publicado: (2026)
por: Menon, Achyutha, et al.
Publicado: (2026)
Agentic AI Frameworks: Architectures, Protocols, and Design Challenges
por: Derouiche, Hana, et al.
Publicado: (2025)
por: Derouiche, Hana, et al.
Publicado: (2025)
Digital Twin and Agentic AI for Wild Fire Disaster Management: Intelligent Virtual Situation Room
por: Morsali, Mohammad, et al.
Publicado: (2026)
por: Morsali, Mohammad, et al.
Publicado: (2026)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
por: Wang, Yixu, et al.
Publicado: (2025)
por: Wang, Yixu, et al.
Publicado: (2025)
Agentic Architect: An Agentic AI Framework for Architecture Design Exploration and Optimization
por: Blasberg, Alexander, et al.
Publicado: (2026)
por: Blasberg, Alexander, et al.
Publicado: (2026)
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
por: van der Maden, Willem, et al.
Publicado: (2026)
por: van der Maden, Willem, et al.
Publicado: (2026)
Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
por: Bluethgen, Christian, et al.
Publicado: (2025)
por: Bluethgen, Christian, et al.
Publicado: (2025)
Stochasticity in Agentic Evaluations: Quantifying Inconsistency with Intraclass Correlation
por: Mustahsan, Zairah, et al.
Publicado: (2025)
por: Mustahsan, Zairah, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluation Framework for AI Systems in "the Wild"
por: Jabbour, Sarah, et al.
Publicado: (2025) -
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
por: Bhonsle, Roshita, et al.
Publicado: (2025) -
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
por: Janet, et al.
Publicado: (2025) -
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
por: Zhang, Hong, et al.
Publicado: (2026) -
Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
por: Mehta, Sushant
Publicado: (2025)