Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914277650071552 |
|---|---|
| author | Atif, Muhammad Ahmed Haji, Nehal Naeem Shaikh, Mohammad Shahid Atif, Muhammad Ebad |
| author_facet | Atif, Muhammad Ahmed Haji, Nehal Naeem Shaikh, Mohammad Shahid Atif, Muhammad Ebad |
| contents | Centralized value learning is often assumed to improve coordination and stability in multi-agent reinforcement learning, yet this assumption is rarely tested under controlled conditions. We directly evaluate it in a fully tabular predator-prey gridworld by comparing independent and centralized Q-learning under explicit embodiment constraints on agent speed and stamina. Across multiple kinematic regimes and asymmetric agent roles, centralized learning fails to provide a consistent advantage and is frequently outperformed by fully independent learning, even under full observability and exact value estimation. Moreover, asymmetric centralized-independent configurations induce persistent coordination breakdowns rather than transient learning instability. By eliminating confounding effects from function approximation and representation learning, our tabular analysis isolates coordination structure as the primary driver of these effects. The results show that increased coordination can become a liability under embodiment constraints, and that the effectiveness of centralized learning is fundamentally regime and role dependent rather than universal. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_17454 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning Atif, Muhammad Ahmed Haji, Nehal Naeem Shaikh, Mohammad Shahid Atif, Muhammad Ebad Multiagent Systems Artificial Intelligence Machine Learning Centralized value learning is often assumed to improve coordination and stability in multi-agent reinforcement learning, yet this assumption is rarely tested under controlled conditions. We directly evaluate it in a fully tabular predator-prey gridworld by comparing independent and centralized Q-learning under explicit embodiment constraints on agent speed and stamina. Across multiple kinematic regimes and asymmetric agent roles, centralized learning fails to provide a consistent advantage and is frequently outperformed by fully independent learning, even under full observability and exact value estimation. Moreover, asymmetric centralized-independent configurations induce persistent coordination breakdowns rather than transient learning instability. By eliminating confounding effects from function approximation and representation learning, our tabular analysis isolates coordination structure as the primary driver of these effects. The results show that increased coordination can become a liability under embodiment constraints, and that the effectiveness of centralized learning is fundamentally regime and role dependent rather than universal. |
| title | Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning |
| topic | Multiagent Systems Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2601.17454 |