Pausing Policy Learning in Non-stationary Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Hyunin, Jin, Ming, Lavaei, Javad, Sojoudi, Somayeh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911887242821632
author Lee, Hyunin
Jin, Ming
Lavaei, Javad
Sojoudi, Somayeh
author_facet Lee, Hyunin
Jin, Ming
Lavaei, Javad
Sojoudi, Somayeh
contents Real-time inference is a challenge of real-world reinforcement learning due to temporal differences in time-varying environments: the system collects data from the past, updates the decision model in the present, and deploys it in the future. We tackle a common belief that continually updating the decision is optimal to minimize the temporal gap. We propose forecasting an online reinforcement learning framework and show that strategically pausing decision updates yields better overall performance by effectively managing aleatoric uncertainty. Theoretically, we compute an optimal ratio between policy update and hold duration, and show that a non-zero policy hold duration provides a sharper upper bound on the dynamic regret. Our experimental evaluations on three different environments also reveal that a non-zero policy hold duration yields higher rewards compared to continuous decision updates.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16053
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pausing Policy Learning in Non-stationary Reinforcement Learning
Lee, Hyunin
Jin, Ming
Lavaei, Javad
Sojoudi, Somayeh
Machine Learning
Real-time inference is a challenge of real-world reinforcement learning due to temporal differences in time-varying environments: the system collects data from the past, updates the decision model in the present, and deploys it in the future. We tackle a common belief that continually updating the decision is optimal to minimize the temporal gap. We propose forecasting an online reinforcement learning framework and show that strategically pausing decision updates yields better overall performance by effectively managing aleatoric uncertainty. Theoretically, we compute an optimal ratio between policy update and hold duration, and show that a non-zero policy hold duration provides a sharper upper bound on the dynamic regret. Our experimental evaluations on three different environments also reveal that a non-zero policy hold duration yields higher rewards compared to continuous decision updates.
title Pausing Policy Learning in Non-stationary Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2405.16053