Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vamplew, Peter, Foale, Cameron, Hayes, Conor F., Mannion, Patrick, Howley, Enda, Dazeley, Richard, Johnson, Scott, Källström, Johan, Ramos, Gabriel, Rădulescu, Roxana, Röpke, Willem, Roijers, Diederik M.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929234250825728
author Vamplew, Peter
Foale, Cameron
Hayes, Conor F.
Mannion, Patrick
Howley, Enda
Dazeley, Richard
Johnson, Scott
Källström, Johan
Ramos, Gabriel
Rădulescu, Roxana
Röpke, Willem
Roijers, Diederik M.
author_facet Vamplew, Peter
Foale, Cameron
Hayes, Conor F.
Mannion, Patrick
Howley, Enda
Dazeley, Richard
Johnson, Scott
Källström, Johan
Ramos, Gabriel
Rădulescu, Roxana
Röpke, Willem
Roijers, Diederik M.
contents Research in multi-objective reinforcement learning (MORL) has introduced the utility-based paradigm, which makes use of both environmental rewards and a function that defines the utility derived by the user from those rewards. In this paper we extend this paradigm to the context of single-objective reinforcement learning (RL), and outline multiple potential benefits including the ability to perform multi-policy learning across tasks relating to uncertain objectives, risk-aware RL, discounting, and safe RL. We also examine the algorithmic implications of adopting a utility-based approach.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02665
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning
Vamplew, Peter
Foale, Cameron
Hayes, Conor F.
Mannion, Patrick
Howley, Enda
Dazeley, Richard
Johnson, Scott
Källström, Johan
Ramos, Gabriel
Rădulescu, Roxana
Röpke, Willem
Roijers, Diederik M.
Machine Learning
Research in multi-objective reinforcement learning (MORL) has introduced the utility-based paradigm, which makes use of both environmental rewards and a function that defines the utility derived by the user from those rewards. In this paper we extend this paradigm to the context of single-objective reinforcement learning (RL), and outline multiple potential benefits including the ability to perform multi-policy learning across tasks relating to uncertain objectives, risk-aware RL, discounting, and safe RL. We also examine the algorithmic implications of adopting a utility-based approach.
title Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2402.02665