Doubly Optimal Policy Evaluation for Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shuze Daniel, Chen, Claire, Zhang, Shangtong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915205947064320
author Liu, Shuze Daniel
Chen, Claire
Zhang, Shangtong
author_facet Liu, Shuze Daniel
Chen, Claire
Zhang, Shangtong
contents Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any improper data-collecting policy or data-processing method substantially deteriorates the variance of evaluation results over long time steps. Thus, policy evaluation often suffers from large variance and requires massive data to achieve the desired accuracy. In this work, we design an optimal combination of data-collecting policy and data-processing baseline. Theoretically, we prove our doubly optimal policy evaluation method is unbiased and guaranteed to have lower variance than previously best-performing methods. Empirically, compared with previous works, we show our method reduces variance substantially and achieves superior empirical performance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02226
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Doubly Optimal Policy Evaluation for Reinforcement Learning
Liu, Shuze Daniel
Chen, Claire
Zhang, Shangtong
Machine Learning
Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any improper data-collecting policy or data-processing method substantially deteriorates the variance of evaluation results over long time steps. Thus, policy evaluation often suffers from large variance and requires massive data to achieve the desired accuracy. In this work, we design an optimal combination of data-collecting policy and data-processing baseline. Theoretically, we prove our doubly optimal policy evaluation method is unbiased and guaranteed to have lower variance than previously best-performing methods. Empirically, compared with previous works, we show our method reduces variance substantially and achieves superior empirical performance.
title Doubly Optimal Policy Evaluation for Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2410.02226