Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seyeon, Lee, Joonhun, Cho, Namhoon, Han, Sungjun, Hwang, Wooseop
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929694403723264
author Kim, Seyeon
Lee, Joonhun
Cho, Namhoon
Han, Sungjun
Hwang, Wooseop
author_facet Kim, Seyeon
Lee, Joonhun
Cho, Namhoon
Han, Sungjun
Hwang, Wooseop
contents Conventional uncertainty-aware temporal difference (TD) learning often assumes a zero-mean Gaussian distribution for TD errors, leading to inaccurate error representations and compromised uncertainty estimation. We introduce a novel framework for generalized Gaussian error modeling in deep reinforcement learning to enhance the flexibility of error distribution modeling by incorporating additional higher-order moment, particularly kurtosis, thereby improving the estimation and mitigation of data-dependent aleatoric uncertainty. We examine the influence of the shape parameter of the generalized Gaussian distribution (GGD) on aleatoric uncertainty and provide a closed-form expression that demonstrates an inverse relationship between uncertainty and the shape parameter. Additionally, we propose a theoretically grounded weighting scheme to address epistemic uncertainty by fully leveraging the GGD. We refine batch inverse variance weighting with bias reduction and kurtosis considerations, enhancing robustness. Experiments with policy gradient algorithms demonstrate significant performance gains.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02295
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
Kim, Seyeon
Lee, Joonhun
Cho, Namhoon
Han, Sungjun
Hwang, Wooseop
Machine Learning
Artificial Intelligence
Probability
Conventional uncertainty-aware temporal difference (TD) learning often assumes a zero-mean Gaussian distribution for TD errors, leading to inaccurate error representations and compromised uncertainty estimation. We introduce a novel framework for generalized Gaussian error modeling in deep reinforcement learning to enhance the flexibility of error distribution modeling by incorporating additional higher-order moment, particularly kurtosis, thereby improving the estimation and mitigation of data-dependent aleatoric uncertainty. We examine the influence of the shape parameter of the generalized Gaussian distribution (GGD) on aleatoric uncertainty and provide a closed-form expression that demonstrates an inverse relationship between uncertainty and the shape parameter. Additionally, we propose a theoretically grounded weighting scheme to address epistemic uncertainty by fully leveraging the GGD. We refine batch inverse variance weighting with bias reduction and kurtosis considerations, enhancing robustness. Experiments with policy gradient algorithms demonstrate significant performance gains.
title Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Probability
url https://arxiv.org/abs/2408.02295