Saved in:
Bibliographic Details
Main Authors: Nonaka, Hiroshi, Ambrozak, Simon, Miskala-Dinc, Sofia R., Ercole, Amedeo, Prins, Aviva
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.11933
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908591195160576
author Nonaka, Hiroshi
Ambrozak, Simon
Miskala-Dinc, Sofia R.
Ercole, Amedeo
Prins, Aviva
author_facet Nonaka, Hiroshi
Ambrozak, Simon
Miskala-Dinc, Sofia R.
Ercole, Amedeo
Prins, Aviva
contents In this work, we propose three efficient restart paradigms for model-free non-stationary reinforcement learning (RL). We identify two core issues with the restart design of Mao et al. (2022)'s RestartQ-UCB algorithm: (1) complete forgetting, where all the information learned about an environment is lost after a restart, and (2) scheduled restarts, in which restarts occur only at predefined timings, regardless of the incompatibility of the policy with the current environment dynamics. We introduce three approaches, which we call partial, adaptive, and selective restarts to modify the algorithms RestartQ-UCB and RANDOMIZEDQ (Wang et al., 2025). We find near-optimal empirical performance in multiple different environments, decreasing dynamic regret by up to $91$% relative to RestartQ-UCB.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Restarts in Non-Stationary Model-Free Reinforcement Learning
Nonaka, Hiroshi
Ambrozak, Simon
Miskala-Dinc, Sofia R.
Ercole, Amedeo
Prins, Aviva
Machine Learning
In this work, we propose three efficient restart paradigms for model-free non-stationary reinforcement learning (RL). We identify two core issues with the restart design of Mao et al. (2022)'s RestartQ-UCB algorithm: (1) complete forgetting, where all the information learned about an environment is lost after a restart, and (2) scheduled restarts, in which restarts occur only at predefined timings, regardless of the incompatibility of the policy with the current environment dynamics. We introduce three approaches, which we call partial, adaptive, and selective restarts to modify the algorithms RestartQ-UCB and RANDOMIZEDQ (Wang et al., 2025). We find near-optimal empirical performance in multiple different environments, decreasing dynamic regret by up to $91$% relative to RestartQ-UCB.
title Efficient Restarts in Non-Stationary Model-Free Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2510.11933