Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Deihim, Azad, Alonso, Eduardo, Apostolopoulou, Dimitra
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915355317764096
author Deihim, Azad
Alonso, Eduardo
Apostolopoulou, Dimitra
author_facet Deihim, Azad
Alonso, Eduardo
Apostolopoulou, Dimitra
contents We present the Multi-Agent Transformer World Model (MATWM), a novel transformer-based world model designed for multi-agent reinforcement learning in both vector- and image-based environments. MATWM combines a decentralized imagination framework with a semi-centralized critic and a teammate prediction module, enabling agents to model and anticipate the behavior of others under partial observability. To address non-stationarity, we incorporate a prioritized replay mechanism that trains the world model on recent experiences, allowing it to adapt to agents' evolving policies. We evaluated MATWM on a broad suite of benchmarks, including the StarCraft Multi-Agent Challenge, PettingZoo, and MeltingPot. MATWM achieves state-of-the-art performance, outperforming both model-free and prior world model approaches, while demonstrating strong sample efficiency, achieving near-optimal performance in as few as 50K environment interactions. Ablation studies confirm the impact of each component, with substantial gains in coordination-heavy tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18537
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning
Deihim, Azad
Alonso, Eduardo
Apostolopoulou, Dimitra
Machine Learning
Multiagent Systems
We present the Multi-Agent Transformer World Model (MATWM), a novel transformer-based world model designed for multi-agent reinforcement learning in both vector- and image-based environments. MATWM combines a decentralized imagination framework with a semi-centralized critic and a teammate prediction module, enabling agents to model and anticipate the behavior of others under partial observability. To address non-stationarity, we incorporate a prioritized replay mechanism that trains the world model on recent experiences, allowing it to adapt to agents' evolving policies. We evaluated MATWM on a broad suite of benchmarks, including the StarCraft Multi-Agent Challenge, PettingZoo, and MeltingPot. MATWM achieves state-of-the-art performance, outperforming both model-free and prior world model approaches, while demonstrating strong sample efficiency, achieving near-optimal performance in as few as 50K environment interactions. Ablation studies confirm the impact of each component, with substantial gains in coordination-heavy tasks.
title Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2506.18537