Maestro: Joint Graph & Config Optimization for Reliable AI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wenxiao, Kattakinda, Priyatham, Feizi, Soheil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909771499569152
author Wang, Wenxiao
Kattakinda, Priyatham
Feizi, Soheil
author_facet Wang, Wenxiao
Kattakinda, Priyatham
Feizi, Soheil
contents Building reliable LLM agents requires decisions at two levels: the graph (which modules exist and how information flows) and the configuration of each node (models, prompts, tools, control knobs). Most existing optimizers tune configurations while holding the graph fixed, leaving structural failure modes unaddressed. We introduce Maestro, a framework-agnostic holistic optimizer for LLM agents that jointly searches over graphs and configurations to maximize agent quality, subject to explicit rollout/token budgets. Beyond numeric metrics, Maestro leverages reflective textual feedback from traces to prioritize edits, improving sample efficiency and targeting specific failure modes. On the IFBench and HotpotQA benchmarks, Maestro consistently surpasses leading prompt optimizers--MIPROv2, GEPA, and GEPA+Merge--by an average of 12%, 4.9%, and 4.86%, respectively; even when restricted to prompt-only optimization, it still leads by 9.65%, 2.37%, and 2.41%. Maestro achieves these results with far fewer rollouts than GEPA. We further show large gains on two applications (interviewer & RAG agents), highlighting that joint graph & configuration search addresses structural failure modes that prompt tuning alone cannot fix.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04642
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Maestro: Joint Graph & Config Optimization for Reliable AI Agents
Wang, Wenxiao
Kattakinda, Priyatham
Feizi, Soheil
Artificial Intelligence
Computation and Language
Machine Learning
Software Engineering
Building reliable LLM agents requires decisions at two levels: the graph (which modules exist and how information flows) and the configuration of each node (models, prompts, tools, control knobs). Most existing optimizers tune configurations while holding the graph fixed, leaving structural failure modes unaddressed. We introduce Maestro, a framework-agnostic holistic optimizer for LLM agents that jointly searches over graphs and configurations to maximize agent quality, subject to explicit rollout/token budgets. Beyond numeric metrics, Maestro leverages reflective textual feedback from traces to prioritize edits, improving sample efficiency and targeting specific failure modes. On the IFBench and HotpotQA benchmarks, Maestro consistently surpasses leading prompt optimizers--MIPROv2, GEPA, and GEPA+Merge--by an average of 12%, 4.9%, and 4.86%, respectively; even when restricted to prompt-only optimization, it still leads by 9.65%, 2.37%, and 2.41%. Maestro achieves these results with far fewer rollouts than GEPA. We further show large gains on two applications (interviewer & RAG agents), highlighting that joint graph & configuration search addresses structural failure modes that prompt tuning alone cannot fix.
title Maestro: Joint Graph & Config Optimization for Reliable AI Agents
topic Artificial Intelligence
Computation and Language
Machine Learning
Software Engineering
url https://arxiv.org/abs/2509.04642