Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929234250825728 |
|---|---|
| author | Vamplew, Peter Foale, Cameron Hayes, Conor F. Mannion, Patrick Howley, Enda Dazeley, Richard Johnson, Scott Källström, Johan Ramos, Gabriel Rădulescu, Roxana Röpke, Willem Roijers, Diederik M. |
| author_facet | Vamplew, Peter Foale, Cameron Hayes, Conor F. Mannion, Patrick Howley, Enda Dazeley, Richard Johnson, Scott Källström, Johan Ramos, Gabriel Rădulescu, Roxana Röpke, Willem Roijers, Diederik M. |
| contents | Research in multi-objective reinforcement learning (MORL) has introduced the utility-based paradigm, which makes use of both environmental rewards and a function that defines the utility derived by the user from those rewards. In this paper we extend this paradigm to the context of single-objective reinforcement learning (RL), and outline multiple potential benefits including the ability to perform multi-policy learning across tasks relating to uncertain objectives, risk-aware RL, discounting, and safe RL. We also examine the algorithmic implications of adopting a utility-based approach. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_02665 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning Vamplew, Peter Foale, Cameron Hayes, Conor F. Mannion, Patrick Howley, Enda Dazeley, Richard Johnson, Scott Källström, Johan Ramos, Gabriel Rădulescu, Roxana Röpke, Willem Roijers, Diederik M. Machine Learning Research in multi-objective reinforcement learning (MORL) has introduced the utility-based paradigm, which makes use of both environmental rewards and a function that defines the utility derived by the user from those rewards. In this paper we extend this paradigm to the context of single-objective reinforcement learning (RL), and outline multiple potential benefits including the ability to perform multi-policy learning across tasks relating to uncertain objectives, risk-aware RL, discounting, and safe RL. We also examine the algorithmic implications of adopting a utility-based approach. |
| title | Utility-Based Reinforcement Learning: Unifying Single-objective and Multi-objective Reinforcement Learning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2402.02665 |