An agent design with goal reaching guarantees for enhancement of learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Osinenko, Pavel, Yaremenko, Grigory, Malaniya, Georgiy, Bolychev, Anton, Gepperth, Alexander
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914920253095936
author Osinenko, Pavel
Yaremenko, Grigory
Malaniya, Georgiy
Bolychev, Anton
Gepperth, Alexander
author_facet Osinenko, Pavel
Yaremenko, Grigory
Malaniya, Georgiy
Bolychev, Anton
Gepperth, Alexander
contents Reinforcement learning is commonly concerned with problems of maximizing accumulated rewards in Markov decision processes. Oftentimes, a certain goal state or a subset of the state space attain maximal reward. In such a case, the environment may be considered solved when the goal is reached. Whereas numerous techniques, learning or non-learning based, exist for solving environments, doing so optimally is the biggest challenge. Say, one may choose a reward rate which penalizes the action effort. Reinforcement learning is currently among the most actively developed frameworks for solving environments optimally by virtue of maximizing accumulated reward, in other words, returns. Yet, tuning agents is a notoriously hard task as reported in a series of works. Our aim here is to help the agent learn a near-optimal policy efficiently while ensuring a goal reaching property of some basis policy that merely solves the environment. We suggest an algorithm, which is fairly flexible, and can be used to augment practically any agent as long as it comprises of a critic. A formal proof of a goal reaching property is provided. Comparative experiments on several problems under popular baseline agents provided an empirical evidence that the learning can indeed be boosted while ensuring goal reaching property.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18118
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An agent design with goal reaching guarantees for enhancement of learning
Osinenko, Pavel
Yaremenko, Grigory
Malaniya, Georgiy
Bolychev, Anton
Gepperth, Alexander
Artificial Intelligence
Systems and Control
Dynamical Systems
Reinforcement learning is commonly concerned with problems of maximizing accumulated rewards in Markov decision processes. Oftentimes, a certain goal state or a subset of the state space attain maximal reward. In such a case, the environment may be considered solved when the goal is reached. Whereas numerous techniques, learning or non-learning based, exist for solving environments, doing so optimally is the biggest challenge. Say, one may choose a reward rate which penalizes the action effort. Reinforcement learning is currently among the most actively developed frameworks for solving environments optimally by virtue of maximizing accumulated reward, in other words, returns. Yet, tuning agents is a notoriously hard task as reported in a series of works. Our aim here is to help the agent learn a near-optimal policy efficiently while ensuring a goal reaching property of some basis policy that merely solves the environment. We suggest an algorithm, which is fairly flexible, and can be used to augment practically any agent as long as it comprises of a critic. A formal proof of a goal reaching property is provided. Comparative experiments on several problems under popular baseline agents provided an empirical evidence that the learning can indeed be boosted while ensuring goal reaching property.
title An agent design with goal reaching guarantees for enhancement of learning
topic Artificial Intelligence
Systems and Control
Dynamical Systems
url https://arxiv.org/abs/2405.18118