Detecting Silent Failures in Multi-Agentic AI Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | Pathak, Divya, Kumar, Harshit, Roy, Anuska, George, Felix, Verma, Mudit, Moogi, Pratibha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Cycle Detection in Agentic Applications
by: George, Felix, et al.
Published: (2025)
by: George, Felix, et al.
Published: (2025)
Metric Criticality Identification for Cloud Microservices
by: Singal, Akanksha, et al.
Published: (2025)
by: Singal, Akanksha, et al.
Published: (2025)
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models
by: Verma, Mudit, et al.
Published: (2024)
by: Verma, Mudit, et al.
Published: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
by: Bisht, Harshit, et al.
Published: (2026)
by: Bisht, Harshit, et al.
Published: (2026)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
by: Desikan, Prasanna, et al.
Published: (2026)
by: Desikan, Prasanna, et al.
Published: (2026)
Hindsight PRIORs for Reward Learning from Human Preferences
by: Verma, Mudit, et al.
Published: (2024)
by: Verma, Mudit, et al.
Published: (2024)
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
by: Yin, Bo, et al.
Published: (2026)
by: Yin, Bo, et al.
Published: (2026)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
by: Roy, Joyjit, et al.
Published: (2026)
by: Roy, Joyjit, et al.
Published: (2026)
AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
by: Roy, Shovan
Published: (2025)
by: Roy, Shovan
Published: (2025)
Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems
by: Or, Barak
Published: (2026)
by: Or, Barak
Published: (2026)
Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
by: Verma, Mudit, et al.
Published: (2024)
by: Verma, Mudit, et al.
Published: (2024)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
by: Bhambri, Siddhant, et al.
Published: (2023)
by: Bhambri, Siddhant, et al.
Published: (2023)
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
by: Zunjare, Pratibha, et al.
Published: (2026)
by: Zunjare, Pratibha, et al.
Published: (2026)
Taming Silent Failures: A Framework for Verifiable AI Reliability
by: Yang, Guan-Yan, et al.
Published: (2025)
by: Yang, Guan-Yan, et al.
Published: (2025)
Willful Disobedience: Automatically Detecting Failures in Agentic Traces
by: Sharma, Reshabh K, et al.
Published: (2026)
by: Sharma, Reshabh K, et al.
Published: (2026)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
by: Gundawar, Atharva, et al.
Published: (2024)
by: Gundawar, Atharva, et al.
Published: (2024)
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
by: Barke, Shraddha, et al.
Published: (2026)
by: Barke, Shraddha, et al.
Published: (2026)
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
by: Roy, Joyjit, et al.
Published: (2026)
by: Roy, Joyjit, et al.
Published: (2026)
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
by: Jha, Saurabh, et al.
Published: (2025)
by: Jha, Saurabh, et al.
Published: (2025)
SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
by: Tripathy, Arihant, et al.
Published: (2025)
by: Tripathy, Arihant, et al.
Published: (2025)
TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning
by: Tayebati, Sina, et al.
Published: (2026)
by: Tayebati, Sina, et al.
Published: (2026)
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
by: Janet, et al.
Published: (2025)
by: Janet, et al.
Published: (2025)
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
by: Advani, Laksh
Published: (2026)
by: Advani, Laksh
Published: (2026)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
by: Pathak, Gangesh, et al.
Published: (2025)
by: Pathak, Gangesh, et al.
Published: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
by: Khoo, Shaun, et al.
Published: (2025)
by: Khoo, Shaun, et al.
Published: (2025)
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
by: Pandey, Mukund
Published: (2026)
by: Pandey, Mukund
Published: (2026)
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
by: Gundawar, Atharva, et al.
Published: (2024)
by: Gundawar, Atharva, et al.
Published: (2024)
Just-in-Time Detection of Silent Security Patches
by: Tang, Xunzhu, et al.
Published: (2023)
by: Tang, Xunzhu, et al.
Published: (2023)
Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue
by: Bhagtani, Kratika, et al.
Published: (2026)
by: Bhagtani, Kratika, et al.
Published: (2026)
Accuracy-Constrained CNN Pruning for Efficient and Reliable EEG-Based Seizure Detection
by: K, Mounvik, et al.
Published: (2025)
by: K, Mounvik, et al.
Published: (2025)
AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
by: Mulian, Hadar, et al.
Published: (2026)
by: Mulian, Hadar, et al.
Published: (2026)
Detecting Object Tracking Failure via Sequential Hypothesis Testing
by: Muñoz, Alejandro Monroy, et al.
Published: (2026)
by: Muñoz, Alejandro Monroy, et al.
Published: (2026)
Agentic Risk-Aware Set-Based Engineering Design
by: Kumar, Varun, et al.
Published: (2026)
by: Kumar, Varun, et al.
Published: (2026)
Agentic AI-based Coverage Closure for Formal Verification
by: Pothireddypalli, Sivaram, et al.
Published: (2026)
by: Pothireddypalli, Sivaram, et al.
Published: (2026)
De-Fake: Style based Anomaly Deepfake Detection
by: Padhi, Sudev Kumar, et al.
Published: (2025)
by: Padhi, Sudev Kumar, et al.
Published: (2025)
Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems
by: Mavridou, Anastasia, et al.
Published: (2025)
by: Mavridou, Anastasia, et al.
Published: (2025)
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
ADR: An Agentic Detection System for Enterprise Agentic AI Security
by: Li, Chenning, et al.
Published: (2026)
by: Li, Chenning, et al.
Published: (2026)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)
by: Harshit
Published: (2025)
Similar Items
-
Unsupervised Cycle Detection in Agentic Applications
by: George, Felix, et al.
Published: (2025) -
Metric Criticality Identification for Cloud Microservices
by: Singal, Akanksha, et al.
Published: (2025) -
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models
by: Verma, Mudit, et al.
Published: (2024) -
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
by: Sahoo, Subramanyam, et al.
Published: (2026) -
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
by: Bisht, Harshit, et al.
Published: (2026)