Sample Efficient Experience Replay in Non-stationary Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Tianyang, Zhang, Zongyuan, Guo, Songxiao, Zhao, Yuanye, Lin, Zheng, Fang, Zihan, Liu, Yi, Luan, Dianxin, Huang, Dong, Cui, Heming, Cui, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908545843200000
author Duan, Tianyang
Zhang, Zongyuan
Guo, Songxiao
Zhao, Yuanye
Lin, Zheng
Fang, Zihan
Liu, Yi
Luan, Dianxin
Huang, Dong
Cui, Heming
Cui, Yong
author_facet Duan, Tianyang
Zhang, Zongyuan
Guo, Songxiao
Zhao, Yuanye
Lin, Zheng
Fang, Zihan
Liu, Yi
Luan, Dianxin
Huang, Dong
Cui, Heming
Cui, Yong
contents Reinforcement learning (RL) in non-stationary environments is challenging, as changing dynamics and rewards quickly make past experiences outdated. Traditional experience replay (ER) methods, especially those using TD-error prioritization, struggle to distinguish between changes caused by the agent's policy and those from the environment, resulting in inefficient learning under dynamic conditions. To address this challenge, we propose the Discrepancy of Environment Dynamics (DoE), a metric that isolates the effects of environment shifts on value functions. Building on this, we introduce Discrepancy of Environment Prioritized Experience Replay (DEER), an adaptive ER framework that prioritizes transitions based on both policy updates and environmental changes. DEER uses a binary classifier to detect environment changes and applies distinct prioritization strategies before and after each shift, enabling more sample-efficient learning. Experiments on four non-stationary benchmarks demonstrate that DEER further improves the performance of off-policy algorithms by 11.54 percent compared to the best-performing state-of-the-art ER methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15032
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sample Efficient Experience Replay in Non-stationary Environments
Duan, Tianyang
Zhang, Zongyuan
Guo, Songxiao
Zhao, Yuanye
Lin, Zheng
Fang, Zihan
Liu, Yi
Luan, Dianxin
Huang, Dong
Cui, Heming
Cui, Yong
Machine Learning
Artificial Intelligence
Networking and Internet Architecture
Reinforcement learning (RL) in non-stationary environments is challenging, as changing dynamics and rewards quickly make past experiences outdated. Traditional experience replay (ER) methods, especially those using TD-error prioritization, struggle to distinguish between changes caused by the agent's policy and those from the environment, resulting in inefficient learning under dynamic conditions. To address this challenge, we propose the Discrepancy of Environment Dynamics (DoE), a metric that isolates the effects of environment shifts on value functions. Building on this, we introduce Discrepancy of Environment Prioritized Experience Replay (DEER), an adaptive ER framework that prioritizes transitions based on both policy updates and environmental changes. DEER uses a binary classifier to detect environment changes and applies distinct prioritization strategies before and after each shift, enabling more sample-efficient learning. Experiments on four non-stationary benchmarks demonstrate that DEER further improves the performance of off-policy algorithms by 11.54 percent compared to the best-performing state-of-the-art ER methods.
title Sample Efficient Experience Replay in Non-stationary Environments
topic Machine Learning
Artificial Intelligence
Networking and Internet Architecture
url https://arxiv.org/abs/2509.15032