AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mulian, Hadar, Zeltyn, Sergey, Levy, Ido, Galanti, Liane, Yaeli, Avi, Shlomov, Segev
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912992455557120
author Mulian, Hadar
Zeltyn, Sergey
Levy, Ido
Galanti, Liane
Yaeli, Avi
Shlomov, Segev
author_facet Mulian, Hadar
Zeltyn, Sergey
Levy, Ido
Galanti, Liane
Yaeli, Avi
Shlomov, Segev
contents We introduce a comprehensive validation framework for LLM-based agentic systems that provides systematic diagnosis and improvement of reliability failures. The framework includes fifteen failure-detection tools and two root-cause analysis modules that jointly uncover weaknesses across input handling, prompt design, and output generation. It integrates lightweight rule-based checks with LLM-as-a-judge assessments to support structured incident detection, classification, and repair. We applied the framework to IBM CUGA, evaluating its performance on the AppWorld and WebArena benchmarks. The analysis revealed recurrent planner misalignments, schema violations, brittle prompt dependencies, and more. Based on these insights, we refined both prompting and coding strategies, maintaining CUGA's benchmark results while enabling mid-sized models such as Llama 4 and Mistral Medium to achieve notable accuracy gains, substantially narrowing the gap with frontier models. Beyond quantitative validation, we conducted an exploratory study that fed the framework's diagnostic outputs and agent description into an LLM for self-reflection and prioritization. This interactive analysis produced actionable insights on recurring failure patterns and focus areas for improvement, demonstrating how validation itself can evolve into an agentic, dialogue-driven process. These results show a path toward scalable, quality assurance, and adaptive validation in production agentic systems, offering a foundation for more robust, interpretable, and self-improving agentic architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29848
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
Mulian, Hadar
Zeltyn, Sergey
Levy, Ido
Galanti, Liane
Yaeli, Avi
Shlomov, Segev
Artificial Intelligence
Multiagent Systems
We introduce a comprehensive validation framework for LLM-based agentic systems that provides systematic diagnosis and improvement of reliability failures. The framework includes fifteen failure-detection tools and two root-cause analysis modules that jointly uncover weaknesses across input handling, prompt design, and output generation. It integrates lightweight rule-based checks with LLM-as-a-judge assessments to support structured incident detection, classification, and repair. We applied the framework to IBM CUGA, evaluating its performance on the AppWorld and WebArena benchmarks. The analysis revealed recurrent planner misalignments, schema violations, brittle prompt dependencies, and more. Based on these insights, we refined both prompting and coding strategies, maintaining CUGA's benchmark results while enabling mid-sized models such as Llama 4 and Mistral Medium to achieve notable accuracy gains, substantially narrowing the gap with frontier models. Beyond quantitative validation, we conducted an exploratory study that fed the framework's diagnostic outputs and agent description into an LLM for self-reflection and prioritization. This interactive analysis produced actionable insights on recurring failure patterns and focus areas for improvement, demonstrating how validation itself can evolve into an agentic, dialogue-driven process. These results show a path toward scalable, quality assurance, and adaptive validation in production agentic systems, offering a foundation for more robust, interpretable, and self-improving agentic architectures.
title AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2603.29848