Learning Generalized Policies for Fully Observable Non-Deterministic Planning Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hofmann, Till, Geffner, Hector
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929340168536064
author Hofmann, Till
Geffner, Hector
author_facet Hofmann, Till
Geffner, Hector
contents General policies represent reactive strategies for solving large families of planning problems like the infinite collection of solvable instances from a given domain. Methods for learning such policies from a collection of small training instances have been developed successfully for classical domains. In this work, we extend the formulations and the resulting combinatorial methods for learning general policies over fully observable, non-deterministic (FOND) domains. We also evaluate the resulting approach experimentally over a number of benchmark domains in FOND planning, present the general policies that result in some of these domains, and prove their correctness. The method for learning general policies for FOND planning can actually be seen as an alternative FOND planning method that searches for solutions, not in the given state space but in an abstract space defined by features that must be learned as well.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02499
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Generalized Policies for Fully Observable Non-Deterministic Planning Domains
Hofmann, Till
Geffner, Hector
Artificial Intelligence
Machine Learning
General policies represent reactive strategies for solving large families of planning problems like the infinite collection of solvable instances from a given domain. Methods for learning such policies from a collection of small training instances have been developed successfully for classical domains. In this work, we extend the formulations and the resulting combinatorial methods for learning general policies over fully observable, non-deterministic (FOND) domains. We also evaluate the resulting approach experimentally over a number of benchmark domains in FOND planning, present the general policies that result in some of these domains, and prove their correctness. The method for learning general policies for FOND planning can actually be seen as an alternative FOND planning method that searches for solutions, not in the given state space but in an abstract space defined by features that must be learned as well.
title Learning Generalized Policies for Fully Observable Non-Deterministic Planning Domains
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.02499