From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akin, Sureyya, Srivastava, Kavita, Kapoor, Prateek B., Sethi, Pradeep G., Patel, Sunita Q., Srivastava, Rahu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912700825600000
author Akin, Sureyya
Srivastava, Kavita
Kapoor, Prateek B.
Sethi, Pradeep G.
Patel, Sunita Q.
Srivastava, Rahu
author_facet Akin, Sureyya
Srivastava, Kavita
Kapoor, Prateek B.
Sethi, Pradeep G.
Patel, Sunita Q.
Srivastava, Rahu
contents Learning cooperative multi-agent policies directly from high-dimensional, multimodal sensory inputs like pixels and audio (from pixels) is notoriously sample-inefficient. Model-free Multi-Agent Reinforcement Learning (MARL) algorithms struggle with the joint challenge of representation learning, partial observability, and credit assignment. To address this, we propose a novel framework based on a shared, generative Multimodal World Model (MWM). Our MWM is trained to learn a compressed latent representation of the environment's dynamics by fusing distributed, multimodal observations from all agents using a scalable attention-based mechanism. Subsequently, we leverage this learned MWM as a fast, "imagined" simulator to train cooperative MARL policies (e.g., MAPPO) entirely within its latent space, decoupling representation learning from policy learning. We introduce a new set of challenging multimodal, multi-agent benchmarks built on a 3D physics simulator. Our experiments demonstrate that our MWM-MARL framework achieves orders-of-magnitude greater sample efficiency compared to state-of-the-art model-free MARL baselines. We further show that our proposed multimodal fusion is essential for task success in environments with sensory asymmetry and that our architecture provides superior robustness to sensor-dropout, a critical feature for real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models
Akin, Sureyya
Srivastava, Kavita
Kapoor, Prateek B.
Sethi, Pradeep G.
Patel, Sunita Q.
Srivastava, Rahu
Multiagent Systems
Learning cooperative multi-agent policies directly from high-dimensional, multimodal sensory inputs like pixels and audio (from pixels) is notoriously sample-inefficient. Model-free Multi-Agent Reinforcement Learning (MARL) algorithms struggle with the joint challenge of representation learning, partial observability, and credit assignment. To address this, we propose a novel framework based on a shared, generative Multimodal World Model (MWM). Our MWM is trained to learn a compressed latent representation of the environment's dynamics by fusing distributed, multimodal observations from all agents using a scalable attention-based mechanism. Subsequently, we leverage this learned MWM as a fast, "imagined" simulator to train cooperative MARL policies (e.g., MAPPO) entirely within its latent space, decoupling representation learning from policy learning. We introduce a new set of challenging multimodal, multi-agent benchmarks built on a 3D physics simulator. Our experiments demonstrate that our MWM-MARL framework achieves orders-of-magnitude greater sample efficiency compared to state-of-the-art model-free MARL baselines. We further show that our proposed multimodal fusion is essential for task success in environments with sensory asymmetry and that our architecture provides superior robustness to sensor-dropout, a critical feature for real-world deployment.
title From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models
topic Multiagent Systems
url https://arxiv.org/abs/2511.01310