Saved in:
Bibliographic Details
Main Author: Cohen-Solal, Quentin
Format: Preprint
Published: 2020
Subjects:
Online Access:https://arxiv.org/abs/2008.01188
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912363747213312
author Cohen-Solal, Quentin
author_facet Cohen-Solal, Quentin
contents In this paper, several techniques for learning game state evaluation functions by reinforcement are proposed. The first is a generalization of tree bootstrapping (tree learning): it is adapted to the context of reinforcement learning without knowledge based on non-linear functions. With this technique, no information is lost during the reinforcement learning process. The second is a modification of minimax with unbounded depth extending the best sequences of actions to the terminal states. This modified search is intended to be used during the learning process. The third is to replace the classic gain of a game (+1 / -1) with a reinforcement heuristic. We study particular reinforcement heuristics such as: quick wins and slow defeats ; scoring ; mobility or presence. The four is a new action selection distribution. The conducted experiments suggest that these techniques improve the level of play. Finally, we apply these different techniques to design program-players to the game of Hex (size 11 and 13) surpassing the level of Mohex 3HNN with reinforcement learning from self-play without knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2008_01188
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Learning to Play Two-Player Perfect-Information Games without Knowledge
Cohen-Solal, Quentin
Artificial Intelligence
In this paper, several techniques for learning game state evaluation functions by reinforcement are proposed. The first is a generalization of tree bootstrapping (tree learning): it is adapted to the context of reinforcement learning without knowledge based on non-linear functions. With this technique, no information is lost during the reinforcement learning process. The second is a modification of minimax with unbounded depth extending the best sequences of actions to the terminal states. This modified search is intended to be used during the learning process. The third is to replace the classic gain of a game (+1 / -1) with a reinforcement heuristic. We study particular reinforcement heuristics such as: quick wins and slow defeats ; scoring ; mobility or presence. The four is a new action selection distribution. The conducted experiments suggest that these techniques improve the level of play. Finally, we apply these different techniques to design program-players to the game of Hex (size 11 and 13) surpassing the level of Mohex 3HNN with reinforcement learning from self-play without knowledge.
title Learning to Play Two-Player Perfect-Information Games without Knowledge
topic Artificial Intelligence
url https://arxiv.org/abs/2008.01188