Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Atif, Muhammad Ahmed, Haji, Nehal Naeem, Shaikh, Mohammad Shahid, Atif, Muhammad Ebad
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914277650071552
author Atif, Muhammad Ahmed
Haji, Nehal Naeem
Shaikh, Mohammad Shahid
Atif, Muhammad Ebad
author_facet Atif, Muhammad Ahmed
Haji, Nehal Naeem
Shaikh, Mohammad Shahid
Atif, Muhammad Ebad
contents Centralized value learning is often assumed to improve coordination and stability in multi-agent reinforcement learning, yet this assumption is rarely tested under controlled conditions. We directly evaluate it in a fully tabular predator-prey gridworld by comparing independent and centralized Q-learning under explicit embodiment constraints on agent speed and stamina. Across multiple kinematic regimes and asymmetric agent roles, centralized learning fails to provide a consistent advantage and is frequently outperformed by fully independent learning, even under full observability and exact value estimation. Moreover, asymmetric centralized-independent configurations induce persistent coordination breakdowns rather than transient learning instability. By eliminating confounding effects from function approximation and representation learning, our tabular analysis isolates coordination structure as the primary driver of these effects. The results show that increased coordination can become a liability under embodiment constraints, and that the effectiveness of centralized learning is fundamentally regime and role dependent rather than universal.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17454
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning
Atif, Muhammad Ahmed
Haji, Nehal Naeem
Shaikh, Mohammad Shahid
Atif, Muhammad Ebad
Multiagent Systems
Artificial Intelligence
Machine Learning
Centralized value learning is often assumed to improve coordination and stability in multi-agent reinforcement learning, yet this assumption is rarely tested under controlled conditions. We directly evaluate it in a fully tabular predator-prey gridworld by comparing independent and centralized Q-learning under explicit embodiment constraints on agent speed and stamina. Across multiple kinematic regimes and asymmetric agent roles, centralized learning fails to provide a consistent advantage and is frequently outperformed by fully independent learning, even under full observability and exact value estimation. Moreover, asymmetric centralized-independent configurations induce persistent coordination breakdowns rather than transient learning instability. By eliminating confounding effects from function approximation and representation learning, our tabular analysis isolates coordination structure as the primary driver of these effects. The results show that increased coordination can become a liability under embodiment constraints, and that the effectiveness of centralized learning is fundamentally regime and role dependent rather than universal.
title Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning
topic Multiagent Systems
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.17454