Investigating Intra-Abstraction Policies For Non-exact Abstraction Algorithms

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Schmöcker, Robin, Dockhorn, Alexander, Rosenhahn, Bodo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915582713004032
author Schmöcker, Robin
Dockhorn, Alexander
Rosenhahn, Bodo
author_facet Schmöcker, Robin
Dockhorn, Alexander
Rosenhahn, Bodo
contents One weakness of Monte Carlo Tree Search (MCTS) is its sample efficiency which can be addressed by building and using state and/or action abstractions in parallel to the tree search such that information can be shared among nodes of the same layer. The primary usage of abstractions for MCTS is to enhance the Upper Confidence Bound (UCB) value during the tree policy by aggregating visits and returns of an abstract node. However, this direct usage of abstractions does not take the case into account where multiple actions with the same parent might be in the same abstract node, as these would then all have the same UCB value, thus requiring a tiebreak rule. In state-of-the-art abstraction algorithms such as pruned On the Go Abstractions (pruned OGA), this case has not been noticed, and a random tiebreak rule was implicitly chosen. In this paper, we propose and empirically evaluate several alternative intra-abstraction policies, several of which outperform the random policy across a majority of environments and parameter settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Intra-Abstraction Policies For Non-exact Abstraction Algorithms
Schmöcker, Robin
Dockhorn, Alexander
Rosenhahn, Bodo
Artificial Intelligence
One weakness of Monte Carlo Tree Search (MCTS) is its sample efficiency which can be addressed by building and using state and/or action abstractions in parallel to the tree search such that information can be shared among nodes of the same layer. The primary usage of abstractions for MCTS is to enhance the Upper Confidence Bound (UCB) value during the tree policy by aggregating visits and returns of an abstract node. However, this direct usage of abstractions does not take the case into account where multiple actions with the same parent might be in the same abstract node, as these would then all have the same UCB value, thus requiring a tiebreak rule. In state-of-the-art abstraction algorithms such as pruned On the Go Abstractions (pruned OGA), this case has not been noticed, and a random tiebreak rule was implicitly chosen. In this paper, we propose and empirically evaluate several alternative intra-abstraction policies, several of which outperform the random policy across a majority of environments and parameter settings.
title Investigating Intra-Abstraction Policies For Non-exact Abstraction Algorithms
topic Artificial Intelligence
url https://arxiv.org/abs/2510.24297