Detecting Silent Failures in Multi-Agentic AI Trajectories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pathak, Divya, Kumar, Harshit, Roy, Anuska, George, Felix, Verma, Mudit, Moogi, Pratibha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unsupervised Cycle Detection in Agentic Applications
von: George, Felix, et al.
Veröffentlicht: (2025)
von: George, Felix, et al.
Veröffentlicht: (2025)
Metric Criticality Identification for Cloud Microservices
von: Singal, Akanksha, et al.
Veröffentlicht: (2025)
von: Singal, Akanksha, et al.
Veröffentlicht: (2025)
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
von: Bisht, Harshit, et al.
Veröffentlicht: (2026)
von: Bisht, Harshit, et al.
Veröffentlicht: (2026)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
von: Desikan, Prasanna, et al.
Veröffentlicht: (2026)
von: Desikan, Prasanna, et al.
Veröffentlicht: (2026)
Hindsight PRIORs for Reward Learning from Human Preferences
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
von: Yin, Bo, et al.
Veröffentlicht: (2026)
von: Yin, Bo, et al.
Veröffentlicht: (2026)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
von: Roy, Joyjit, et al.
Veröffentlicht: (2026)
von: Roy, Joyjit, et al.
Veröffentlicht: (2026)
AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
von: Roy, Shovan
Veröffentlicht: (2025)
von: Roy, Shovan
Veröffentlicht: (2025)
Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems
von: Or, Barak
Veröffentlicht: (2026)
von: Or, Barak
Veröffentlicht: (2026)
Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2023)
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2023)
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2026)
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2026)
Taming Silent Failures: A Framework for Verifiable AI Reliability
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2025)
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2025)
Willful Disobedience: Automatically Detecting Failures in Agentic Traces
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2026)
von: Sharma, Reshabh K, et al.
Veröffentlicht: (2026)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
von: Barke, Shraddha, et al.
Veröffentlicht: (2026)
von: Barke, Shraddha, et al.
Veröffentlicht: (2026)
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
von: Roy, Joyjit, et al.
Veröffentlicht: (2026)
von: Roy, Joyjit, et al.
Veröffentlicht: (2026)
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
von: Jha, Saurabh, et al.
Veröffentlicht: (2025)
von: Jha, Saurabh, et al.
Veröffentlicht: (2025)
SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
von: Tripathy, Arihant, et al.
Veröffentlicht: (2025)
von: Tripathy, Arihant, et al.
Veröffentlicht: (2025)
TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning
von: Tayebati, Sina, et al.
Veröffentlicht: (2026)
von: Tayebati, Sina, et al.
Veröffentlicht: (2026)
From Failure Modes to Reliability Awareness in Generative and Agentic AI System
von: Janet, et al.
Veröffentlicht: (2025)
von: Janet, et al.
Veröffentlicht: (2025)
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
von: Advani, Laksh
Veröffentlicht: (2026)
von: Advani, Laksh
Veröffentlicht: (2026)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
von: Pathak, Gangesh, et al.
Veröffentlicht: (2025)
von: Pathak, Gangesh, et al.
Veröffentlicht: (2025)
With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
von: Khoo, Shaun, et al.
Veröffentlicht: (2025)
von: Khoo, Shaun, et al.
Veröffentlicht: (2025)
Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework
von: Pandey, Mukund
Veröffentlicht: (2026)
von: Pandey, Mukund
Veröffentlicht: (2026)
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
Just-in-Time Detection of Silent Security Patches
von: Tang, Xunzhu, et al.
Veröffentlicht: (2023)
von: Tang, Xunzhu, et al.
Veröffentlicht: (2023)
Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue
von: Bhagtani, Kratika, et al.
Veröffentlicht: (2026)
von: Bhagtani, Kratika, et al.
Veröffentlicht: (2026)
Accuracy-Constrained CNN Pruning for Efficient and Reliable EEG-Based Seizure Detection
von: K, Mounvik, et al.
Veröffentlicht: (2025)
von: K, Mounvik, et al.
Veröffentlicht: (2025)
AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
von: Mulian, Hadar, et al.
Veröffentlicht: (2026)
von: Mulian, Hadar, et al.
Veröffentlicht: (2026)
Detecting Object Tracking Failure via Sequential Hypothesis Testing
von: Muñoz, Alejandro Monroy, et al.
Veröffentlicht: (2026)
von: Muñoz, Alejandro Monroy, et al.
Veröffentlicht: (2026)
Agentic Risk-Aware Set-Based Engineering Design
von: Kumar, Varun, et al.
Veröffentlicht: (2026)
von: Kumar, Varun, et al.
Veröffentlicht: (2026)
Agentic AI-based Coverage Closure for Formal Verification
von: Pothireddypalli, Sivaram, et al.
Veröffentlicht: (2026)
von: Pothireddypalli, Sivaram, et al.
Veröffentlicht: (2026)
De-Fake: Style based Anomaly Deepfake Detection
von: Padhi, Sudev Kumar, et al.
Veröffentlicht: (2025)
von: Padhi, Sudev Kumar, et al.
Veröffentlicht: (2025)
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
von: Feuer, Benjamin, et al.
Veröffentlicht: (2025)
von: Feuer, Benjamin, et al.
Veröffentlicht: (2025)
Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems
von: Mavridou, Anastasia, et al.
Veröffentlicht: (2025)
von: Mavridou, Anastasia, et al.
Veröffentlicht: (2025)
ADR: An Agentic Detection System for Enterprise Agentic AI Security
von: Li, Chenning, et al.
Veröffentlicht: (2026)
von: Li, Chenning, et al.
Veröffentlicht: (2026)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
von: Harshit
Veröffentlicht: (2025)
von: Harshit
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unsupervised Cycle Detection in Agentic Applications
von: George, Felix, et al.
Veröffentlicht: (2025) -
Metric Criticality Identification for Cloud Microservices
von: Singal, Akanksha, et al.
Veröffentlicht: (2025) -
On the Brittle Foundations of ReAct Prompting for Agentic Large Language Models
von: Verma, Mudit, et al.
Veröffentlicht: (2024) -
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
von: Bisht, Harshit, et al.
Veröffentlicht: (2026)