Adaptive Exploration for Data-Efficient General Value Function Evaluations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jain, Arushi, Hanna, Josiah P., Precup, Doina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910647559651328
author Jain, Arushi
Hanna, Josiah P.
Precup, Doina
author_facet Jain, Arushi
Hanna, Josiah P.
Precup, Doina
contents General Value Functions (GVFs) (Sutton et al., 2011) represent predictive knowledge in reinforcement learning. Each GVF computes the expected return for a given policy, based on a unique reward. Existing methods relying on fixed behavior policies or pre-collected data often face data efficiency issues when learning multiple GVFs in parallel using off-policy methods. To address this, we introduce GVFExplorer, which adaptively learns a single behavior policy that efficiently collects data for evaluating multiple GVFs in parallel. Our method optimizes the behavior policy by minimizing the total variance in return across GVFs, thereby reducing the required environmental interactions. We use an existing temporal-difference-style variance estimator to approximate the return variance. We prove that each behavior policy update decreases the overall mean squared error in GVF predictions. We empirically show our method's performance in tabular and nonlinear function approximation settings, including Mujoco environments, with stationary and non-stationary reward signals, optimizing data usage and reducing prediction errors across multiple GVFs.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07838
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Exploration for Data-Efficient General Value Function Evaluations
Jain, Arushi
Hanna, Josiah P.
Precup, Doina
Machine Learning
Artificial Intelligence
General Value Functions (GVFs) (Sutton et al., 2011) represent predictive knowledge in reinforcement learning. Each GVF computes the expected return for a given policy, based on a unique reward. Existing methods relying on fixed behavior policies or pre-collected data often face data efficiency issues when learning multiple GVFs in parallel using off-policy methods. To address this, we introduce GVFExplorer, which adaptively learns a single behavior policy that efficiently collects data for evaluating multiple GVFs in parallel. Our method optimizes the behavior policy by minimizing the total variance in return across GVFs, thereby reducing the required environmental interactions. We use an existing temporal-difference-style variance estimator to approximate the return variance. We prove that each behavior policy update decreases the overall mean squared error in GVF predictions. We empirically show our method's performance in tabular and nonlinear function approximation settings, including Mujoco environments, with stationary and non-stationary reward signals, optimizing data usage and reducing prediction errors across multiple GVFs.
title Adaptive Exploration for Data-Efficient General Value Function Evaluations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.07838