Minimax Strikes Back

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cohen-Solal, Quentin, Cazenave, Tristan
Formato: Preprint
Publicado: 2020
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918463324291072
author Cohen-Solal, Quentin
Cazenave, Tristan
author_facet Cohen-Solal, Quentin
Cazenave, Tristan
contents Deep Reinforcement Learning reaches a superhuman level of play in many complete information games. The state of the art algorithm for learning with zero knowledge is AlphaZero. We take another approach, Athénan, which uses a different, Minimax-based, search algorithm called Descent, as well as different learning targets and that does not use a policy. We show that for multiple games it is much more efficient than the reimplementation of AlphaZero: Polygames. It is even competitive with Polygames when Polygames uses 100 times more GPU (at least for some games). One of the keys to the superior performance is that the cost of generating state data for training is approximately 296 times lower with Athénan. With the same reasonable ressources, Athénan without reinforcement heuristic is at least 7 times faster than Polygames and much more than 30 times faster with reinforcement heuristic.
format Preprint
id arxiv_https___arxiv_org_abs_2012_10700
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Minimax Strikes Back
Cohen-Solal, Quentin
Cazenave, Tristan
Artificial Intelligence
Deep Reinforcement Learning reaches a superhuman level of play in many complete information games. The state of the art algorithm for learning with zero knowledge is AlphaZero. We take another approach, Athénan, which uses a different, Minimax-based, search algorithm called Descent, as well as different learning targets and that does not use a policy. We show that for multiple games it is much more efficient than the reimplementation of AlphaZero: Polygames. It is even competitive with Polygames when Polygames uses 100 times more GPU (at least for some games). One of the keys to the superior performance is that the cost of generating state data for training is approximately 296 times lower with Athénan. With the same reasonable ressources, Athénan without reinforcement heuristic is at least 7 times faster than Polygames and much more than 30 times faster with reinforcement heuristic.
title Minimax Strikes Back
topic Artificial Intelligence
url https://arxiv.org/abs/2012.10700