Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manglik, Akshay, Shanker, Apaar, Deshpande, Kaustubh, Qin, Jason, Maurya, Yash, Chatrath, Veronica, Kalmath, Vijay S., Lentz, Levi, Yuan, Xue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916035055058944
author Manglik, Akshay
Shanker, Apaar
Deshpande, Kaustubh
Qin, Jason
Maurya, Yash
Chatrath, Veronica
Kalmath, Vijay S.
Lentz, Levi
Yuan
Xue
author_facet Manglik, Akshay
Shanker, Apaar
Deshpande, Kaustubh
Qin, Jason
Maurya, Yash
Chatrath, Veronica
Kalmath, Vijay S.
Lentz, Levi
Yuan
Xue
contents Diagnosing failures in LLM agents remains largely manual. Practitioners inspect a small subset of execution traces, form ad-hoc hypotheses, and iterate. This process misses patterns that only emerge across trace populations and does not scale to production corpora where individual traces span tens of thousands of tokens. We formalize the problem of corpus-level trace diagnostics. Given a corpus of execution traces, the goal is to produce grounded natural-language insights that characterize systematic behavioral patterns across trace groups, each linked to supporting evidence. We present the Insights Generator (IG), a multi-agent system that answers diagnostic questions by proposing and testing hypotheses across the trace corpus to produce an evidence-backed insights report. We evaluate IG across qualitative and objective dimensions, spanning rubric-based report assessment and downstream performance improvements achieved by implementing IG insights. Human experts using IG reports improve scaffold performance by 30.4pp over the unmodified baseline scaffold, and coding agents leveraging IG-derived insights show consistent and stable gains. Across benchmarks, IG's scout-investigator architecture produces findings comparable in detection coverage to competing approaches, while domain experts rated IG reports as leading depth and evidence quality.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21347
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
Manglik, Akshay
Shanker, Apaar
Deshpande, Kaustubh
Qin, Jason
Maurya, Yash
Chatrath, Veronica
Kalmath, Vijay S.
Lentz, Levi
Yuan
Xue
Artificial Intelligence
Machine Learning
Software Engineering
Diagnosing failures in LLM agents remains largely manual. Practitioners inspect a small subset of execution traces, form ad-hoc hypotheses, and iterate. This process misses patterns that only emerge across trace populations and does not scale to production corpora where individual traces span tens of thousands of tokens. We formalize the problem of corpus-level trace diagnostics. Given a corpus of execution traces, the goal is to produce grounded natural-language insights that characterize systematic behavioral patterns across trace groups, each linked to supporting evidence. We present the Insights Generator (IG), a multi-agent system that answers diagnostic questions by proposing and testing hypotheses across the trace corpus to produce an evidence-backed insights report. We evaluate IG across qualitative and objective dimensions, spanning rubric-based report assessment and downstream performance improvements achieved by implementing IG insights. Human experts using IG reports improve scaffold performance by 30.4pp over the unmodified baseline scaffold, and coding agents leveraging IG-derived insights show consistent and stable gains. Across benchmarks, IG's scout-investigator architecture produces findings comparable in detection coverage to competing approaches, while domain experts rated IG reports as leading depth and evidence quality.
title Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
topic Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2605.21347