Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenwu, Song, Yuran, Zhao, Mingze, Jin, Bo, Li, Wenhao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917544685731840
author Li, Wenwu
Song, Yuran
Zhao, Mingze
Jin, Bo
Li, Wenhao
author_facet Li, Wenwu
Song, Yuran
Zhao, Mingze
Jin, Bo
Li, Wenhao
contents While Multi-Agent Systems (MAS) empower Large Language Models to tackle complex reasoning tasks through collaborative interaction, optimizing their dynamics remains a formidable challenge due to the discrete, non-differentiable nature of the computation graph and the sparsity of global supervisory signals. Existing black-box optimizers struggle to attribute trajectory-level failure to specific local components, resulting in inefficient, high-variance exploration. We argue that tractable MAS optimization needs structural inductive biases to disentangle error signals. We propose temporal and structural credit assignment, which decomposes the objective along two axes: (i) temporal credit, using state-space bottlenecks to identify critical rounds, and (ii) structural credit, using stationary role policies to isolate agent contributions. Leveraging these decomposed signals, we introduce a discrete, verbalized block coordinate descent algorithm for iterative refinement. Rather than indiscriminate global updates, it alternates between optimizing role prompts and aggregation protocols, using LLM-generated "proxy gradients" to target only the identified weak links. Across diverse reasoning benchmarks, our approach substantially reduces query complexity while improving performance, providing a principled and interpretable path toward self-improving MAS.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30227
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization
Li, Wenwu
Song, Yuran
Zhao, Mingze
Jin, Bo
Li, Wenhao
Multiagent Systems
Artificial Intelligence
While Multi-Agent Systems (MAS) empower Large Language Models to tackle complex reasoning tasks through collaborative interaction, optimizing their dynamics remains a formidable challenge due to the discrete, non-differentiable nature of the computation graph and the sparsity of global supervisory signals. Existing black-box optimizers struggle to attribute trajectory-level failure to specific local components, resulting in inefficient, high-variance exploration. We argue that tractable MAS optimization needs structural inductive biases to disentangle error signals. We propose temporal and structural credit assignment, which decomposes the objective along two axes: (i) temporal credit, using state-space bottlenecks to identify critical rounds, and (ii) structural credit, using stationary role policies to isolate agent contributions. Leveraging these decomposed signals, we introduce a discrete, verbalized block coordinate descent algorithm for iterative refinement. Rather than indiscriminate global updates, it alternates between optimizing role prompts and aggregation protocols, using LLM-generated "proxy gradients" to target only the identified weak links. Across diverse reasoning benchmarks, our approach substantially reduces query complexity while improving performance, providing a principled and interpretable path toward self-improving MAS.
title Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization
topic Multiagent Systems
Artificial Intelligence
url https://arxiv.org/abs/2605.30227