Willful Disobedience: Automatically Detecting Failures in Agentic Traces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharma, Reshabh K, Barke, Shraddha, Zorn, Benjamin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918490389086208
author Sharma, Reshabh K
Barke, Shraddha
Zorn, Benjamin
author_facet Sharma, Reshabh K
Barke, Shraddha
Zorn, Benjamin
contents AI agents are increasingly embedded in real software systems, where they execute multi-step workflows through multi-turn dialogue, tool invocations, and intermediate decisions. These long execution histories, called agentic traces, make validation difficult. Outcome-only benchmarks can miss critical procedural failures, such as incorrect workflow routing, unsafe tool usage, or violations of prompt-specified rules. This paper presents AgentPex, an AI-powered tool designed to systematically evaluate agentic traces. AgentPex extracts behavioral rules from agent prompts and system instructions, then uses these specifications to automatically evaluate traces for compliance. We evaluate AgentPex on 424 traces from $τ^2$-bench across models in telecom, retail, and airline customer service. Our results show that AgentPex distinguishes agent behavior across models and surfaces specification violations that are not captured by outcome-only scoring. It also provides fine-grained analysis by domain and metric, enabling developers to understand agent strengths and weaknesses at scale. The source code of AgentPex is available at https://github.com/microsoft/agentpex.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Willful Disobedience: Automatically Detecting Failures in Agentic Traces
Sharma, Reshabh K
Barke, Shraddha
Zorn, Benjamin
Software Engineering
Artificial Intelligence
AI agents are increasingly embedded in real software systems, where they execute multi-step workflows through multi-turn dialogue, tool invocations, and intermediate decisions. These long execution histories, called agentic traces, make validation difficult. Outcome-only benchmarks can miss critical procedural failures, such as incorrect workflow routing, unsafe tool usage, or violations of prompt-specified rules. This paper presents AgentPex, an AI-powered tool designed to systematically evaluate agentic traces. AgentPex extracts behavioral rules from agent prompts and system instructions, then uses these specifications to automatically evaluate traces for compliance. We evaluate AgentPex on 424 traces from $τ^2$-bench across models in telecom, retail, and airline customer service. Our results show that AgentPex distinguishes agent behavior across models and surfaces specification violations that are not captured by outcome-only scoring. It also provides fine-grained analysis by domain and metric, enabling developers to understand agent strengths and weaknesses at scale. The source code of AgentPex is available at https://github.com/microsoft/agentpex.
title Willful Disobedience: Automatically Detecting Failures in Agentic Traces
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2603.23806