Reinforcement Learning for Hanabi

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Nina, France, Kordel K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918041658327040
author Cohen, Nina
France, Kordel K.
author_facet Cohen, Nina
France, Kordel K.
contents Hanabi has become a popular game for research when it comes to reinforcement learning (RL) as it is one of the few cooperative card games where you have incomplete knowledge of the entire environment, thus presenting a challenge for a RL agent. We explored different tabular and deep reinforcement learning algorithms to see which had the best performance both against an agent of the same type and also against other types of agents. We establish that certain agents played their highest scoring games against specific agents while others exhibited higher scores on average by adapting to the opposing agent's behavior. We attempted to quantify the conditions under which each algorithm provides the best advantage and identified the most interesting interactions between agents of different types. In the end, we found that temporal difference (TD) algorithms had better overall performance and balancing of play types compared to tabular agents. Specifically, tabular Expected SARSA and deep Q-Learning agents showed the best performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00458
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning for Hanabi
Cohen, Nina
France, Kordel K.
Machine Learning
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
Hanabi has become a popular game for research when it comes to reinforcement learning (RL) as it is one of the few cooperative card games where you have incomplete knowledge of the entire environment, thus presenting a challenge for a RL agent. We explored different tabular and deep reinforcement learning algorithms to see which had the best performance both against an agent of the same type and also against other types of agents. We establish that certain agents played their highest scoring games against specific agents while others exhibited higher scores on average by adapting to the opposing agent's behavior. We attempted to quantify the conditions under which each algorithm provides the best advantage and identified the most interesting interactions between agents of different types. In the end, we found that temporal difference (TD) algorithms had better overall performance and balancing of play types compared to tabular agents. Specifically, tabular Expected SARSA and deep Q-Learning agents showed the best performance.
title Reinforcement Learning for Hanabi
topic Machine Learning
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
url https://arxiv.org/abs/2506.00458