High entropy leads to symmetry equivariant policies in Dec-POMDPs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Forkel, Johannes, Ruhdorfer, Constantin, Beukman, Michael, Bulling, Andreas, Foerster, Jakob
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911656158691328
author Forkel, Johannes
Ruhdorfer, Constantin
Beukman, Michael
Bulling, Andreas
Foerster, Jakob
author_facet Forkel, Johannes
Ruhdorfer, Constantin
Beukman, Michael
Bulling, Andreas
Foerster, Jakob
contents We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.r.t. all symmetries of the Dec-POMDP. In particular, policies coming from different initializations will be fully compatible, in that their cross-play returns are equal to their self-play returns. Through extensive evaluation of independent PPO, arguably the standard baseline deep multi-agent policy gradient algorithm, in the Hanabi, Overcooked and Yokai environments, we find that the entropy coefficient has a massive influence on the cross-play returns between independently trained policies, and that the decrease in self-play returns coming from increased entropy regularization can often be counteracted by greedifying the learned policies after training. In Hanabi in particular we achieve a new SOTA in inter-seed cross-play this way. While we give examples of Dec-POMDPs in which one cannot learn the optimal symmetry equivariant policy this way, both our theoretical and empirical results suggest that one should consider far higher entropy coefficients during hyperparameter sweeps in Dec-POMDPs than is typically done.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle High entropy leads to symmetry equivariant policies in Dec-POMDPs
Forkel, Johannes
Ruhdorfer, Constantin
Beukman, Michael
Bulling, Andreas
Foerster, Jakob
Machine Learning
Multiagent Systems
We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.r.t. all symmetries of the Dec-POMDP. In particular, policies coming from different initializations will be fully compatible, in that their cross-play returns are equal to their self-play returns. Through extensive evaluation of independent PPO, arguably the standard baseline deep multi-agent policy gradient algorithm, in the Hanabi, Overcooked and Yokai environments, we find that the entropy coefficient has a massive influence on the cross-play returns between independently trained policies, and that the decrease in self-play returns coming from increased entropy regularization can often be counteracted by greedifying the learned policies after training. In Hanabi in particular we achieve a new SOTA in inter-seed cross-play this way. While we give examples of Dec-POMDPs in which one cannot learn the optimal symmetry equivariant policy this way, both our theoretical and empirical results suggest that one should consider far higher entropy coefficients during hyperparameter sweeps in Dec-POMDPs than is typically done.
title High entropy leads to symmetry equivariant policies in Dec-POMDPs
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2511.22581