Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Jianglin, Tang, Jingcheng, Jing, Gangshan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908389326454784
author Ding, Jianglin
Tang, Jingcheng
Jing, Gangshan
author_facet Ding, Jianglin
Tang, Jingcheng
Jing, Gangshan
contents Action-dependent individual policies, which incorporate both environmental states and the actions of other agents in decision-making, have emerged as a promising paradigm for achieving global optimality in multi-agent reinforcement learning (MARL). However, the existing literature often adopts auto-regressive action-dependent policies, where each agent's policy depends on the actions of all preceding agents. This formulation incurs substantial computational complexity as the number of agents increases, thereby limiting scalability. In this work, we consider a more generalized class of action-dependent policies, which do not necessarily follow the auto-regressive form. We propose to use the `action dependency graph (ADG)' to model the inter-agent action dependencies. Within the context of MARL problems structured by coordination graphs, we prove that an action-dependent policy with a sparse ADG can achieve global optimality, provided the ADG satisfies specific conditions specified by the coordination graph. Building on this theoretical foundation, we develop a tabular policy iteration algorithm with guaranteed global optimality. Furthermore, we integrate our framework into several SOTA algorithms and conduct experiments in complex environments. The empirical results affirm the robustness and applicability of our approach in more general scenarios, underscoring its potential for broader MARL challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning
Ding, Jianglin
Tang, Jingcheng
Jing, Gangshan
Machine Learning
Artificial Intelligence
Systems and Control
Optimization and Control
Action-dependent individual policies, which incorporate both environmental states and the actions of other agents in decision-making, have emerged as a promising paradigm for achieving global optimality in multi-agent reinforcement learning (MARL). However, the existing literature often adopts auto-regressive action-dependent policies, where each agent's policy depends on the actions of all preceding agents. This formulation incurs substantial computational complexity as the number of agents increases, thereby limiting scalability. In this work, we consider a more generalized class of action-dependent policies, which do not necessarily follow the auto-regressive form. We propose to use the `action dependency graph (ADG)' to model the inter-agent action dependencies. Within the context of MARL problems structured by coordination graphs, we prove that an action-dependent policy with a sparse ADG can achieve global optimality, provided the ADG satisfies specific conditions specified by the coordination graph. Building on this theoretical foundation, we develop a tabular policy iteration algorithm with guaranteed global optimality. Furthermore, we integrate our framework into several SOTA algorithms and conduct experiments in complex environments. The empirical results affirm the robustness and applicability of our approach in more general scenarios, underscoring its potential for broader MARL challenges.
title Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2506.00797