TRAIL: Trace Reasoning and Agentic Issue Localization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Deshpande, Darshan, Gangal, Varun, Mehta, Hersh, Krishnan, Jitin, Kannappan, Anand, Qian, Rebecca
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916808550776832
author Deshpande, Darshan
Gangal, Varun
Mehta, Hersh
Krishnan, Jitin
Kannappan, Anand
Qian, Rebecca
author_facet Deshpande, Darshan
Gangal, Varun
Mehta, Hersh
Krishnan, Jitin
Kannappan, Anand
Qian, Rebecca
contents The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate. Current evaluation methods depend on manual, domain-specific human analysis of lengthy workflow traces - an approach that does not scale with the growing complexity and volume of agentic outputs. Error analysis in these settings is further complicated by the interplay of external tool outputs and language model reasoning, making it more challenging than traditional software debugging. In this work, we (1) articulate the need for robust and dynamic evaluation methods for agentic workflow traces, (2) introduce a formal taxonomy of error types encountered in agentic systems, and (3) present a set of 148 large human-annotated traces (TRAIL) constructed using this taxonomy and grounded in established agentic benchmarks. To ensure ecological validity, we curate traces from both single and multi-agent systems, focusing on real-world applications such as software engineering and open-world information retrieval. Our evaluations reveal that modern long context LLMs perform poorly at trace debugging, with the best Gemini-2.5-pro model scoring a mere 11% on TRAIL. Our dataset and code are made publicly available to support and accelerate future research in scalable evaluation for agentic workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TRAIL: Trace Reasoning and Agentic Issue Localization
Deshpande, Darshan
Gangal, Varun
Mehta, Hersh
Krishnan, Jitin
Kannappan, Anand
Qian, Rebecca
Artificial Intelligence
Computation and Language
The increasing adoption of agentic workflows across diverse domains brings a critical need to scalably and systematically evaluate the complex traces these systems generate. Current evaluation methods depend on manual, domain-specific human analysis of lengthy workflow traces - an approach that does not scale with the growing complexity and volume of agentic outputs. Error analysis in these settings is further complicated by the interplay of external tool outputs and language model reasoning, making it more challenging than traditional software debugging. In this work, we (1) articulate the need for robust and dynamic evaluation methods for agentic workflow traces, (2) introduce a formal taxonomy of error types encountered in agentic systems, and (3) present a set of 148 large human-annotated traces (TRAIL) constructed using this taxonomy and grounded in established agentic benchmarks. To ensure ecological validity, we curate traces from both single and multi-agent systems, focusing on real-world applications such as software engineering and open-world information retrieval. Our evaluations reveal that modern long context LLMs perform poorly at trace debugging, with the best Gemini-2.5-pro model scoring a mere 11% on TRAIL. Our dataset and code are made publicly available to support and accelerate future research in scalable evaluation for agentic workflows.
title TRAIL: Trace Reasoning and Agentic Issue Localization
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.08638