Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
Fuente:
arXiv
Guardado en:
| Autor principal: | Vinay, Vaishali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
por: Janet, et al.
Publicado: (2025)
por: Janet, et al.
Publicado: (2025)
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
por: Qi, Zhengyang, et al.
Publicado: (2026)
por: Qi, Zhengyang, et al.
Publicado: (2026)
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
por: Garg, Ishir, et al.
Publicado: (2026)
por: Garg, Ishir, et al.
Publicado: (2026)
Entropy Collapse: A Universal Failure Mode of Intelligent Systems
por: Khanh, Truong Xuan, et al.
Publicado: (2025)
por: Khanh, Truong Xuan, et al.
Publicado: (2025)
The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines
por: Vinay, Vaishali
Publicado: (2025)
por: Vinay, Vaishali
Publicado: (2025)
LLM-Powered AI Agent Systems and Their Applications in Industry
por: Liang, Guannan, et al.
Publicado: (2025)
por: Liang, Guannan, et al.
Publicado: (2025)
Failures in Perspective-taking of Multimodal AI Systems
por: Leonard, Bridget, et al.
Publicado: (2024)
por: Leonard, Bridget, et al.
Publicado: (2024)
A Taxonomy of Prompt Defects in LLM Systems
por: Tian, Haoye, et al.
Publicado: (2025)
por: Tian, Haoye, et al.
Publicado: (2025)
AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
por: Sapkota, Ranjan, et al.
Publicado: (2025)
por: Sapkota, Ranjan, et al.
Publicado: (2025)
On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
por: Mangold, Aline, et al.
Publicado: (2025)
por: Mangold, Aline, et al.
Publicado: (2025)
LLM Agents in Law: Taxonomy, Applications, and Challenges
por: Liu, Shuang, et al.
Publicado: (2026)
por: Liu, Shuang, et al.
Publicado: (2026)
Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025
por: Ansari, Samar
Publicado: (2026)
por: Ansari, Samar
Publicado: (2026)
Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness
por: Nguyen, Phat, et al.
Publicado: (2026)
por: Nguyen, Phat, et al.
Publicado: (2026)
Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications
por: Sun, Esther, et al.
Publicado: (2025)
por: Sun, Esther, et al.
Publicado: (2025)
Time, Causality, and Observability Failures in Distributed AI Inference Systems
por: Sharma, Ankur, et al.
Publicado: (2026)
por: Sharma, Ankur, et al.
Publicado: (2026)
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
por: Dalrymple, David "davidad", et al.
Publicado: (2024)
por: Dalrymple, David "davidad", et al.
Publicado: (2024)
LLM-enabled Applications Require System-Level Threat Monitoring
por: Zhang, Yedi, et al.
Publicado: (2026)
por: Zhang, Yedi, et al.
Publicado: (2026)
Authenticated Workflows: A Systems Approach to Protecting Agentic AI
por: Rajagopalan, Mohan, et al.
Publicado: (2026)
por: Rajagopalan, Mohan, et al.
Publicado: (2026)
Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs
por: Dhulshette, Nilesh, et al.
Publicado: (2025)
por: Dhulshette, Nilesh, et al.
Publicado: (2025)
Automated Analysis of Global AI Safety Initiatives: A Taxonomy-Driven LLM Approach
por: Semitsu, Takayuki, et al.
Publicado: (2026)
por: Semitsu, Takayuki, et al.
Publicado: (2026)
SymptomWise: A Deterministic Reasoning Layer for Reliable and Efficient AI Systems
por: Henry, Isaac, et al.
Publicado: (2026)
por: Henry, Isaac, et al.
Publicado: (2026)
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
por: Feuer, Benjamin, et al.
Publicado: (2024)
por: Feuer, Benjamin, et al.
Publicado: (2024)
LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
por: Agarwal, Shradha, et al.
Publicado: (2026)
por: Agarwal, Shradha, et al.
Publicado: (2026)
A Taxonomy of Hierarchical Multi-Agent Systems: Design Patterns, Coordination Mechanisms, and Industrial Applications
por: Moore, David J.
Publicado: (2025)
por: Moore, David J.
Publicado: (2025)
LiveFC: A System for Live Fact-Checking of Audio Streams
por: V, Venktesh, et al.
Publicado: (2024)
por: V, Venktesh, et al.
Publicado: (2024)
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
por: Pandey, Mukund
Publicado: (2026)
por: Pandey, Mukund
Publicado: (2026)
Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety
por: Zhang, Wenxiao, et al.
Publicado: (2025)
por: Zhang, Wenxiao, et al.
Publicado: (2025)
AI Agent Systems: Architectures, Applications, and Evaluation
por: Xu, Bin
Publicado: (2026)
por: Xu, Bin
Publicado: (2026)
Perspectives on a Reliability Monitoring Framework for Agentic AI Systems
por: Flehmig, Niclas, et al.
Publicado: (2025)
por: Flehmig, Niclas, et al.
Publicado: (2025)
Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems
por: Mishra, Saurabh, et al.
Publicado: (2024)
por: Mishra, Saurabh, et al.
Publicado: (2024)
Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability
por: Li, Xianyou, et al.
Publicado: (2026)
por: Li, Xianyou, et al.
Publicado: (2026)
AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
por: Mulian, Hadar, et al.
Publicado: (2026)
por: Mulian, Hadar, et al.
Publicado: (2026)
The Cognitive Circuit Breaker: A Systems Engineering Framework for Intrinsic AI Reliability
por: Pan, Jonathan
Publicado: (2026)
por: Pan, Jonathan
Publicado: (2026)
Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems
por: Or, Barak
Publicado: (2026)
por: Or, Barak
Publicado: (2026)
Quantifying Automation Risk in High-Automation AI Systems: A Bayesian Framework for Failure Propagation and Optimal Oversight
por: Srivastava, Vishal, et al.
Publicado: (2026)
por: Srivastava, Vishal, et al.
Publicado: (2026)
Revolutionizing System Reliability: The Role of AI in Predictive Maintenance Strategies
por: Bidollahkhani, Michael, et al.
Publicado: (2024)
por: Bidollahkhani, Michael, et al.
Publicado: (2024)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
por: Huang, Donghao, et al.
Publicado: (2026)
por: Huang, Donghao, et al.
Publicado: (2026)
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
por: Xia, Boming, et al.
Publicado: (2024)
por: Xia, Boming, et al.
Publicado: (2024)
A Survey on Failure Analysis and Fault Injection in AI Systems
por: Yu, Guangba, et al.
Publicado: (2024)
por: Yu, Guangba, et al.
Publicado: (2024)
Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems
por: Qi, Shihao, et al.
Publicado: (2026)
por: Qi, Shihao, et al.
Publicado: (2026)
Ejemplares similares
-
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
por: Janet, et al.
Publicado: (2025) -
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
por: Qi, Zhengyang, et al.
Publicado: (2026) -
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
por: Garg, Ishir, et al.
Publicado: (2026) -
Entropy Collapse: A Universal Failure Mode of Intelligent Systems
por: Khanh, Truong Xuan, et al.
Publicado: (2025) -
The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines
por: Vinay, Vaishali
Publicado: (2025)