Successor Features for Transfer in Alternating Markov Games

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Amatya, Sunny, Ren, Yi, Xu, Zhe, Zhang, Wenlong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915417506709504
author Amatya, Sunny
Ren, Yi
Xu, Zhe
Zhang, Wenlong
author_facet Amatya, Sunny
Ren, Yi
Xu, Zhe
Zhang, Wenlong
contents This paper explores successor features for knowledge transfer in zero-sum, complete-information, and turn-based games. Prior research in single-agent systems has shown that successor features can provide a ``jump start" for agents when facing new tasks with varying reward structures. However, knowledge transfer in games typically relies on value and equilibrium transfers, which heavily depends on the similarity between tasks. This reliance can lead to failures when the tasks differ significantly. To address this issue, this paper presents an application of successor features to games and presents a novel algorithm called Game Generalized Policy Improvement (GGPI), designed to address Markov games in multi-agent reinforcement learning. The proposed algorithm enables the transfer of learning values and policies across games. An upper bound of the errors for transfer is derived as a function the similarity of the task. Through experiments with a turn-based pursuer-evader game, we demonstrate that the GGPI algorithm can generate high-reward interactions and one-shot policy transfer. When further tested in a wider set of initial conditions, the GGPI algorithm achieves higher success rates with improved path efficiency compared to those of the baseline algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Successor Features for Transfer in Alternating Markov Games
Amatya, Sunny
Ren, Yi
Xu, Zhe
Zhang, Wenlong
Multiagent Systems
Computer Science and Game Theory
This paper explores successor features for knowledge transfer in zero-sum, complete-information, and turn-based games. Prior research in single-agent systems has shown that successor features can provide a ``jump start" for agents when facing new tasks with varying reward structures. However, knowledge transfer in games typically relies on value and equilibrium transfers, which heavily depends on the similarity between tasks. This reliance can lead to failures when the tasks differ significantly. To address this issue, this paper presents an application of successor features to games and presents a novel algorithm called Game Generalized Policy Improvement (GGPI), designed to address Markov games in multi-agent reinforcement learning. The proposed algorithm enables the transfer of learning values and policies across games. An upper bound of the errors for transfer is derived as a function the similarity of the task. Through experiments with a turn-based pursuer-evader game, we demonstrate that the GGPI algorithm can generate high-reward interactions and one-shot policy transfer. When further tested in a wider set of initial conditions, the GGPI algorithm achieves higher success rates with improved path efficiency compared to those of the baseline algorithms.
title Successor Features for Transfer in Alternating Markov Games
topic Multiagent Systems
Computer Science and Game Theory
url https://arxiv.org/abs/2507.22278