Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rebstock, Douglas, Solinas, Christopher, Sturtevant, Nathan R., Buro, Michael
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2404.13150
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910415561162752
author Rebstock, Douglas
Solinas, Christopher
Sturtevant, Nathan R.
Buro, Michael
author_facet Rebstock, Douglas
Solinas, Christopher
Sturtevant, Nathan R.
Buro, Michael
contents Traditional search algorithms have issues when applied to games of imperfect information where the number of possible underlying states and trajectories are very large. This challenge is particularly evident in trick-taking card games. While state sampling techniques such as Perfect Information Monte Carlo (PIMC) search has shown success in these contexts, they still have major limitations. We present Generative Observation Monte Carlo Tree Search (GO-MCTS), which utilizes MCTS on observation sequences generated by a game specific model. This method performs the search within the observation space and advances the search using a model that depends solely on the agent's observations. Additionally, we demonstrate that transformers are well-suited as the generative model in this context, and we demonstrate a process for iteratively training the transformer via population-based self-play. The efficacy of GO-MCTS is demonstrated in various games of imperfect information, such as Hearts, Skat, and "The Crew: The Quest for Planet Nine," with promising results.
format Preprint
id arxiv_https___arxiv_org_abs_2404_13150
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transformer Based Planning in the Observation Space with Applications to Trick Taking Card Games
Rebstock, Douglas
Solinas, Christopher
Sturtevant, Nathan R.
Buro, Michael
Artificial Intelligence
Machine Learning
Traditional search algorithms have issues when applied to games of imperfect information where the number of possible underlying states and trajectories are very large. This challenge is particularly evident in trick-taking card games. While state sampling techniques such as Perfect Information Monte Carlo (PIMC) search has shown success in these contexts, they still have major limitations. We present Generative Observation Monte Carlo Tree Search (GO-MCTS), which utilizes MCTS on observation sequences generated by a game specific model. This method performs the search within the observation space and advances the search using a model that depends solely on the agent's observations. Additionally, we demonstrate that transformers are well-suited as the generative model in this context, and we demonstrate a process for iteratively training the transformer via population-based self-play. The efficacy of GO-MCTS is demonstrated in various games of imperfect information, such as Hearts, Skat, and "The Crew: The Quest for Planet Nine," with promising results.
title Transformer Based Planning in the Observation Space with Applications to Trick Taking Card Games
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.13150