An Empirical Investigation of Value-Based Multi-objective Reinforcement Learning for Stochastic Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Kewen, Vamplew, Peter, Foale, Cameron, Dazeley, Richard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913187750739968
author Ding, Kewen
Vamplew, Peter
Foale, Cameron
Dazeley, Richard
author_facet Ding, Kewen
Vamplew, Peter
Foale, Cameron
Dazeley, Richard
contents One common approach to solve multi-objective reinforcement learning (MORL) problems is to extend conventional Q-learning by using vector Q-values in combination with a utility function. However issues can arise with this approach in the context of stochastic environments, particularly when optimising for the Scalarised Expected Reward (SER) criterion. This paper extends prior research, providing a detailed examination of the factors influencing the frequency with which value-based MORL Q-learning algorithms learn the SER-optimal policy for an environment with stochastic state transitions. We empirically examine several variations of the core multi-objective Q-learning algorithm as well as reward engineering approaches, and demonstrate the limitations of these methods. In particular, we highlight the critical impact of the noisy Q-value estimates issue on the stability and convergence of these algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2401_03163
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Investigation of Value-Based Multi-objective Reinforcement Learning for Stochastic Environments
Ding, Kewen
Vamplew, Peter
Foale, Cameron
Dazeley, Richard
Machine Learning
One common approach to solve multi-objective reinforcement learning (MORL) problems is to extend conventional Q-learning by using vector Q-values in combination with a utility function. However issues can arise with this approach in the context of stochastic environments, particularly when optimising for the Scalarised Expected Reward (SER) criterion. This paper extends prior research, providing a detailed examination of the factors influencing the frequency with which value-based MORL Q-learning algorithms learn the SER-optimal policy for an environment with stochastic state transitions. We empirically examine several variations of the core multi-objective Q-learning algorithm as well as reward engineering approaches, and demonstrate the limitations of these methods. In particular, we highlight the critical impact of the noisy Q-value estimates issue on the stability and convergence of these algorithms.
title An Empirical Investigation of Value-Based Multi-objective Reinforcement Learning for Stochastic Environments
topic Machine Learning
url https://arxiv.org/abs/2401.03163