Policy-Conditioned Policies for Multi-Agent Task Solving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Yue, Zhu, Shuhui, Li, Wenhao, Li, Ang, Qiao, Dan, Poupart, Pascal, Zha, Hongyuan, Wang, Baoxiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918262973923328
author Lin, Yue
Zhu, Shuhui
Li, Wenhao
Li, Ang
Qiao, Dan
Poupart, Pascal
Zha, Hongyuan
Wang, Baoxiang
author_facet Lin, Yue
Zhu, Shuhui
Li, Wenhao
Li, Ang
Qiao, Dan
Poupart, Pascal
Zha, Hongyuan
Wang, Baoxiang
contents In multi-agent tasks, the central challenge lies in the dynamic adaptation of strategies. However, directly conditioning on opponents' strategies is intractable in the prevalent deep reinforcement learning paradigm due to a fundamental ``representational bottleneck'': neural policies are opaque, high-dimensional parameter vectors that are incomprehensible to other agents. In this work, we propose a paradigm shift that bridges this gap by representing policies as human-interpretable source code and utilizing Large Language Models (LLMs) as approximate interpreters. This programmatic representation allows us to operationalize the game-theoretic concept of \textit{Program Equilibrium}. We reformulate the learning problem by utilizing LLMs to perform optimization directly in the space of programmatic policies. The LLM functions as a point-wise best-response operator that iteratively synthesizes and refines the ego agent's policy code to respond to the opponent's strategy. We formalize this process as \textit{Programmatic Iterated Best Response (PIBR)}, an algorithm where the policy code is optimized by textual gradients, using structured feedback derived from game utility and runtime unit tests. We demonstrate that this approach effectively solves several standard coordination matrix games and a cooperative Level-Based Foraging environment.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Policy-Conditioned Policies for Multi-Agent Task Solving
Lin, Yue
Zhu, Shuhui
Li, Wenhao
Li, Ang
Qiao, Dan
Poupart, Pascal
Zha, Hongyuan
Wang, Baoxiang
Computer Science and Game Theory
Artificial Intelligence
In multi-agent tasks, the central challenge lies in the dynamic adaptation of strategies. However, directly conditioning on opponents' strategies is intractable in the prevalent deep reinforcement learning paradigm due to a fundamental ``representational bottleneck'': neural policies are opaque, high-dimensional parameter vectors that are incomprehensible to other agents. In this work, we propose a paradigm shift that bridges this gap by representing policies as human-interpretable source code and utilizing Large Language Models (LLMs) as approximate interpreters. This programmatic representation allows us to operationalize the game-theoretic concept of \textit{Program Equilibrium}. We reformulate the learning problem by utilizing LLMs to perform optimization directly in the space of programmatic policies. The LLM functions as a point-wise best-response operator that iteratively synthesizes and refines the ego agent's policy code to respond to the opponent's strategy. We formalize this process as \textit{Programmatic Iterated Best Response (PIBR)}, an algorithm where the policy code is optimized by textual gradients, using structured feedback derived from game utility and runtime unit tests. We demonstrate that this approach effectively solves several standard coordination matrix games and a cooperative Level-Based Foraging environment.
title Policy-Conditioned Policies for Multi-Agent Task Solving
topic Computer Science and Game Theory
Artificial Intelligence
url https://arxiv.org/abs/2512.21024