Sequence Modeling for N-Agent Ad Hoc Teamwork

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Caroline, Shi, Di Yang, Liebman, Elad, Durugkar, Ishan, Rahman, Arrasy, Stone, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914102180315136
author Wang, Caroline
Shi, Di Yang
Liebman, Elad
Durugkar, Ishan
Rahman, Arrasy
Stone, Peter
author_facet Wang, Caroline
Shi, Di Yang
Liebman, Elad
Durugkar, Ishan
Rahman, Arrasy
Stone, Peter
contents N-agent ad hoc teamwork (NAHT) is a newly introduced challenge in multi-agent reinforcement learning, where controlled subteams of varying sizes must dynamically collaborate with varying numbers and types of unknown teammates without pre-coordination. The existing learning algorithm (POAM) considers only independent learning for its flexibility in dealing with a changing number of agents. However, independent learning fails to fully capture the inter-agent dynamics essential for effective collaboration. Based on our observation that transformers deal effectively with sequences with varying lengths and have been shown to be highly effective for a variety of machine learning problems, this work introduces a centralized, transformer-based method for N-agent ad hoc teamwork. Our proposed approach incorporates historical observations and actions of all controlled agents, enabling optimal responses to diverse and unseen teammates in partially observable environments. Empirical evaluation on a StarCraft II task demonstrates that MAT-NAHT outperforms POAM, achieving superior sample efficiency and generalization, without auxiliary agent-modeling objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sequence Modeling for N-Agent Ad Hoc Teamwork
Wang, Caroline
Shi, Di Yang
Liebman, Elad
Durugkar, Ishan
Rahman, Arrasy
Stone, Peter
Multiagent Systems
N-agent ad hoc teamwork (NAHT) is a newly introduced challenge in multi-agent reinforcement learning, where controlled subteams of varying sizes must dynamically collaborate with varying numbers and types of unknown teammates without pre-coordination. The existing learning algorithm (POAM) considers only independent learning for its flexibility in dealing with a changing number of agents. However, independent learning fails to fully capture the inter-agent dynamics essential for effective collaboration. Based on our observation that transformers deal effectively with sequences with varying lengths and have been shown to be highly effective for a variety of machine learning problems, this work introduces a centralized, transformer-based method for N-agent ad hoc teamwork. Our proposed approach incorporates historical observations and actions of all controlled agents, enabling optimal responses to diverse and unseen teammates in partially observable environments. Empirical evaluation on a StarCraft II task demonstrates that MAT-NAHT outperforms POAM, achieving superior sample efficiency and generalization, without auxiliary agent-modeling objectives.
title Sequence Modeling for N-Agent Ad Hoc Teamwork
topic Multiagent Systems
url https://arxiv.org/abs/2506.05527