Planning with a Learned Policy Basis to Optimally Solve Complex Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Infante, Guillermo, Kuric, David, Jonsson, Anders, Gómez, Vicenç, van Hoof, Herke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929370862452736
author Infante, Guillermo
Kuric, David
Jonsson, Anders
Gómez, Vicenç
van Hoof, Herke
author_facet Infante, Guillermo
Kuric, David
Jonsson, Anders
Gómez, Vicenç
van Hoof, Herke
contents Conventional reinforcement learning (RL) methods can successfully solve a wide range of sequential decision problems. However, learning policies that can generalize predictably across multiple tasks in a setting with non-Markovian reward specifications is a challenging problem. We propose to use successor features to learn a policy basis so that each (sub)policy in it solves a well-defined subproblem. In a task described by a finite state automaton (FSA) that involves the same set of subproblems, the combination of these (sub)policies can then be used to generate an optimal solution without additional learning. In contrast to other methods that combine (sub)policies via planning, our method asymptotically attains global optimality, even in stochastic environments.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15301
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Planning with a Learned Policy Basis to Optimally Solve Complex Tasks
Infante, Guillermo
Kuric, David
Jonsson, Anders
Gómez, Vicenç
van Hoof, Herke
Machine Learning
Artificial Intelligence
Conventional reinforcement learning (RL) methods can successfully solve a wide range of sequential decision problems. However, learning policies that can generalize predictably across multiple tasks in a setting with non-Markovian reward specifications is a challenging problem. We propose to use successor features to learn a policy basis so that each (sub)policy in it solves a well-defined subproblem. In a task described by a finite state automaton (FSA) that involves the same set of subproblems, the combination of these (sub)policies can then be used to generate an optimal solution without additional learning. In contrast to other methods that combine (sub)policies via planning, our method asymptotically attains global optimality, even in stochastic environments.
title Planning with a Learned Policy Basis to Optimally Solve Complex Tasks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2403.15301