AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kartik, NVJK, Sapra, Garvit, Hada, Rishav, Pareek, Nikhil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912591491629056
author Kartik, NVJK
Sapra, Garvit
Hada, Rishav
Pareek, Nikhil
author_facet Kartik, NVJK
Sapra, Garvit
Hada, Rishav
Pareek, Nikhil
contents With the growing adoption of Large Language Models (LLMs) in automating complex, multi-agent workflows, organizations face mounting risks from errors, emergent behaviors, and systemic failures that current evaluation methods fail to capture. We present AgentCompass, the first evaluation framework designed specifically for post-deployment monitoring and debugging of agentic workflows. AgentCompass models the reasoning process of expert debuggers through a structured, multi-stage analytical pipeline: error identification and categorization, thematic clustering, quantitative scoring, and strategic summarization. The framework is further enhanced with a dual memory system-episodic and semantic-that enables continual learning across executions. Through collaborations with design partners, we demonstrate the framework's practical utility on real-world deployments, before establishing its efficacy against the publicly available TRAIL benchmark. AgentCompass achieves state-of-the-art results on key metrics, while uncovering critical issues missed in human annotations, underscoring its role as a robust, developer-centric tool for reliable monitoring and improvement of agentic systems in production.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14647
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
Kartik, NVJK
Sapra, Garvit
Hada, Rishav
Pareek, Nikhil
Artificial Intelligence
Computation and Language
With the growing adoption of Large Language Models (LLMs) in automating complex, multi-agent workflows, organizations face mounting risks from errors, emergent behaviors, and systemic failures that current evaluation methods fail to capture. We present AgentCompass, the first evaluation framework designed specifically for post-deployment monitoring and debugging of agentic workflows. AgentCompass models the reasoning process of expert debuggers through a structured, multi-stage analytical pipeline: error identification and categorization, thematic clustering, quantitative scoring, and strategic summarization. The framework is further enhanced with a dual memory system-episodic and semantic-that enables continual learning across executions. Through collaborations with design partners, we demonstrate the framework's practical utility on real-world deployments, before establishing its efficacy against the publicly available TRAIL benchmark. AgentCompass achieves state-of-the-art results on key metrics, while uncovering critical issues missed in human annotations, underscoring its role as a robust, developer-centric tool for reliable monitoring and improvement of agentic systems in production.
title AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.14647