Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Poupart, Yoann, Beynier, Aurélie, Maudet, Nicolas
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909472775995392
author Poupart, Yoann
Beynier, Aurélie
Maudet, Nicolas
author_facet Poupart, Yoann
Beynier, Aurélie
Maudet, Nicolas
contents Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While learning intrinsically interpretable models remains a prominent approach, its scalability and flexibility are limited in handling complex tasks or multi-agent dynamics. This paper advocates for direct interpretability, generating post hoc explanations directly from trained models, as a versatile and scalable alternative, offering insights into agents' behaviour, emergent phenomena, and biases without altering models' architectures. We explore modern methods, including relevance backpropagation, knowledge edition, model steering, activation patching, sparse autoencoders and circuit discovery, to highlight their applicability to single-agent, multi-agent, and training process challenges. By addressing MADRL interpretability, we propose directions aiming to advance active topics such as team identification, swarm coordination and sample efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
Poupart, Yoann
Beynier, Aurélie
Maudet, Nicolas
Artificial Intelligence
Multi-Agent Deep Reinforcement Learning (MADRL) was proven efficient in solving complex problems in robotics or games, yet most of the trained models are hard to interpret. While learning intrinsically interpretable models remains a prominent approach, its scalability and flexibility are limited in handling complex tasks or multi-agent dynamics. This paper advocates for direct interpretability, generating post hoc explanations directly from trained models, as a versatile and scalable alternative, offering insights into agents' behaviour, emergent phenomena, and biases without altering models' architectures. We explore modern methods, including relevance backpropagation, knowledge edition, model steering, activation patching, sparse autoencoders and circuit discovery, to highlight their applicability to single-agent, multi-agent, and training process challenges. By addressing MADRL interpretability, we propose directions aiming to advance active topics such as team identification, swarm coordination and sample efficiency.
title Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2502.00726